All Articles
16423 articles total
Mean-to-Score Discrete Diffusion: Posterior-Mean Denoisers for Score Entropy
This research introduces Mean-to-Score (M2S), a method to improve discrete diffusion models by ensuring score vectors are Bayes realizable through a posterior-mean prediction mechanism.
VoLN: Vision-Only Long-Horizon Navigation---Paradigm, Benchmark, and Method
The authors propose Vision-Only Long-Horizon Navigation (VoLN), a new paradigm for autonomous agents to navigate using only local in-scene cues rather than external instructions.
When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation
ResponseGuard is a high-speed, single-pass vision-language guardrail that outperforms reasoning-based models in detecting harmfulness with significantly lower latency.
Cycle-Consistent and Uncertainty-Aware Neural Surrogates for Tokamak Edge Plasmas
This work presents cycle-consistent neural surrogates for tokamak edge plasma simulations, enabling real-time control and parameter recovery with high accuracy.
Token Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought Models
The paper explores how the convergence of chain-of-thought reasoning in LLMs can be detected early using linear probes on internal model representations.
Adaptive Identity Anchoring: Closed-Loop Keyframe Placement for Synthetic Paired Supervision in Video Face Swapping
Adaptive Identity Anchoring (AIA) is proposed to improve video face swapping by dynamically placing keyframes to maintain identity stability and texture quality.
RUMBA: Russian User Memory Benchmark
RUMBA is a new benchmark for testing the long-term conversational memory of LLMs, focusing on temporal reasoning and multi-session retrieval in Russian.
Thinkink: 2D Spatial Ink-native Interaction with LLMs
Thinkink is a 2D spatial ink-native interface designed for collaborative ideation between humans and LLMs via handwritten sketches and text.
I tried out OpenAI’s new AI keypad — which will be fun for some coders and slightly mystifying to everyone else
OpenAI has introduced a new AI-powered keypad, designed to assist users with coding and other tasks, though its utility may be limited to a specific subset of users.
Wildfire forces evacuation of NASA's Deep Space Network complex in Spain
A wildfire in Spain has forced the evacuation of NASA's Deep Space Network complex, with damage assessments pending.
Paramount/WBD merger delayed for months as states' lawsuit moves toward trial
The merger between Paramount and Warner Bros. Discovery (WBD) has been delayed by several months due to an ongoing lawsuit from state attorneys general.
Scaling Up Formal Representation of Clinical Trial Protocols in Ensemble Logic Using LLMs: A Preliminary Study
Researchers introduce the CT-TEL workflow, which uses LLMs to translate unstructured clinical trial protocols into formal Temporal Ensemble Logic (TEL) for better automated reasoning.
PC-Edit: Prompt-Contrastive Region Discovery and Region-Guided Editing
PC-Edit is a prompt-contrastive framework for training-free image editing using MM-DiT, enabling high-quality object replacement and background preservation without manual masks.
GRADRAG: Cross-Component Prompt Adaptation for Coordinated Multi-Agent RAG
GRADRAG is a framework for coordinated multi-agent RAG systems that uses a computational graph to propagate evaluation feedback and iteratively optimize prompts.
Toward cryptographically verifiable authorization for autonomous AI agents: A security hypothesis, preliminary formal model, and proof-of-concept implementation
This paper proposes a cryptographically verifiable authorization framework for autonomous AI agents using zk-SNARKs (Groth16) to prove policy compliance without revealing private attributes.
From Static Bibliometrics to Dynamic Knowledge Graphs: An LLM-Powered Framework for Modernizing Science, Technology, and Innovation (STI) Analytics
A hybrid, symbolic-first framework is proposed to modernize science and technology analytics by integrating dynamic knowledge graphs with LLM-assisted semantic augmentation.
Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin
Researchers developed a bootstrapping pipeline using GMM-HMM and audio encoders to create phonetic forced aligners for low-resource language varieties like Chengdu Mandarin.
M$^3$-Gen: Interpretable Multimodal Generation of Gene Expression Profiles Using Clinical and Imaging Data
M3-Gen is a multimodal GAN framework that generates interpretable gene expression profiles by conditioning on histopathology images and clinical metadata.
AIs don't do what you want. This is bad
A discussion on the misalignment between AI's outputs and user intent, highlighting the risks of AI not doing what the user actually wants.
The road to epsilon-zero: Nim always ends, even with infinite ordinals
A mathematical exploration of the Nim game and its properties regarding infinite ordinals and the road to epsilon-zero.