AI/ML arXiv cs.AI

Mean-to-Score Discrete Diffusion: Posterior-Mean Denoisers for Score Entropy

This research introduces Mean-to-Score (M2S), a method to improve discrete diffusion models by ensuring score vectors are Bayes realizable through a posterior-mean prediction mechanism.

AI/ML arXiv cs.AI

VoLN: Vision-Only Long-Horizon Navigation---Paradigm, Benchmark, and Method

The authors propose Vision-Only Long-Horizon Navigation (VoLN), a new paradigm for autonomous agents to navigate using only local in-scene cues rather than external instructions.

AI/ML arXiv cs.AI

When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation

ResponseGuard is a high-speed, single-pass vision-language guardrail that outperforms reasoning-based models in detecting harmfulness with significantly lower latency.

Other arXiv cs.AI

Cycle-Consistent and Uncertainty-Aware Neural Surrogates for Tokamak Edge Plasmas

This work presents cycle-consistent neural surrogates for tokamak edge plasma simulations, enabling real-time control and parameter recovery with high accuracy.

AI/ML arXiv cs.AI

Token Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought Models

The paper explores how the convergence of chain-of-thought reasoning in LLMs can be detected early using linear probes on internal model representations.

AI/ML arXiv cs.AI

Adaptive Identity Anchoring: Closed-Loop Keyframe Placement for Synthetic Paired Supervision in Video Face Swapping

Adaptive Identity Anchoring (AIA) is proposed to improve video face swapping by dynamically placing keyframes to maintain identity stability and texture quality.

AI/ML arXiv cs.AI

RUMBA: Russian User Memory Benchmark

RUMBA is a new benchmark for testing the long-term conversational memory of LLMs, focusing on temporal reasoning and multi-session retrieval in Russian.

Other arXiv cs.AI

Thinkink: 2D Spatial Ink-native Interaction with LLMs

Thinkink is a 2D spatial ink-native interface designed for collaborative ideation between humans and LLMs via handwritten sketches and text.

AI/ML TechCrunch

I tried out OpenAI’s new AI keypad — which will be fun for some coders and slightly mystifying to everyone else

OpenAI has introduced a new AI-powered keypad, designed to assist users with coding and other tasks, though its utility may be limited to a specific subset of users.

Other Ars Technica

Wildfire forces evacuation of NASA's Deep Space Network complex in Spain

A wildfire in Spain has forced the evacuation of NASA's Deep Space Network complex, with damage assessments pending.

Tech Business/VC Ars Technica

Paramount/WBD merger delayed for months as states' lawsuit moves toward trial

The merger between Paramount and Warner Bros. Discovery (WBD) has been delayed by several months due to an ongoing lawsuit from state attorneys general.

AI/ML arXiv cs.AI

Scaling Up Formal Representation of Clinical Trial Protocols in Ensemble Logic Using LLMs: A Preliminary Study

Researchers introduce the CT-TEL workflow, which uses LLMs to translate unstructured clinical trial protocols into formal Temporal Ensemble Logic (TEL) for better automated reasoning.

AI/ML arXiv cs.AI

PC-Edit: Prompt-Contrastive Region Discovery and Region-Guided Editing

PC-Edit is a prompt-contrastive framework for training-free image editing using MM-DiT, enabling high-quality object replacement and background preservation without manual masks.

AI/ML arXiv cs.AI

GRADRAG: Cross-Component Prompt Adaptation for Coordinated Multi-Agent RAG

GRADRAG is a framework for coordinated multi-agent RAG systems that uses a computational graph to propagate evaluation feedback and iteratively optimize prompts.

Cybersecurity arXiv cs.AI

Toward cryptographically verifiable authorization for autonomous AI agents: A security hypothesis, preliminary formal model, and proof-of-concept implementation

This paper proposes a cryptographically verifiable authorization framework for autonomous AI agents using zk-SNARKs (Groth16) to prove policy compliance without revealing private attributes.

AI/ML arXiv cs.AI

From Static Bibliometrics to Dynamic Knowledge Graphs: An LLM-Powered Framework for Modernizing Science, Technology, and Innovation (STI) Analytics

A hybrid, symbolic-first framework is proposed to modernize science and technology analytics by integrating dynamic knowledge graphs with LLM-assisted semantic augmentation.

AI/ML arXiv cs.AI

Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin

Researchers developed a bootstrapping pipeline using GMM-HMM and audio encoders to create phonetic forced aligners for low-resource language varieties like Chengdu Mandarin.

AI/ML arXiv cs.AI

M$^3$-Gen: Interpretable Multimodal Generation of Gene Expression Profiles Using Clinical and Imaging Data

M3-Gen is a multimodal GAN framework that generates interpretable gene expression profiles by conditioning on histopathology images and clinical metadata.

AI/ML Hacker News

AIs don't do what you want. This is bad

A discussion on the misalignment between AI's outputs and user intent, highlighting the risks of AI not doing what the user actually wants.

Other Hacker News

The road to epsilon-zero: Nim always ends, even with infinite ordinals

A mathematical exploration of the Nim game and its properties regarding infinite ordinals and the road to epsilon-zero.