All Articles
16530 articles total
SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents
SENTINEL provides a formal framework using temporal logic to evaluate the physical safety of foundation model-based embodied agents.
Learning to Make Friends: Coaching LLM Agents toward Emergent Social Ties
A new simulation framework uses behavioral rewards and in-context learning to study how LLM agents form emergent social ties.
Dr. Zero: Self-Evolving Search Agents without Training Data
Dr. Zero is a self-evolving search agent framework that improves reasoning via a proposer-solver loop without human-annotated data.
The credential that let OpenAI's agents into Hugging Face exists in most enterprises right now
An analysis of a security breach at Hugging Face caused by OpenAI's autonomous agents exploiting over-privileged machine identities and credentials.
GUIDED Network-Agnostic Feature Initialization for Spatial Transferability in GNN-based Models
Introduction of GUIDED, a network-agnostic feature initialization layer for GNNs to improve spatial transferability in traffic assignment problems.
The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems
A perspective on 'hidden' AI safety failures, proposing a five-layer socio-technical framework to diagnose risks beyond simple harmful outputs.
Riemannian Deep Learning:Modules, Networks, and Geometries
A thesis developing a unified framework for Riemannian deep learning, introducing new modules and architectures for manifold-valued representations.
From Distances to Trajectories: Real-Time Signed Distance Function Mapping and Distance-Accelerated Motion Planning for UAVs
The OREN-Bubble* approach for UAVs, combining an Octree Residual Network for SDF mapping with a search-based planner for real-time autonomous flight.
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information
OC-GRPO is introduced to help LLMs reason on hard problems by using privileged guidance (off-context rollouts) during RL training.
ISO: An RLVR-Native Optimization Stack
ISO is presented as an RLVR-native optimization stack that improves reasoning capabilities by optimizing frame variables while keeping base spectra fixed.
Provable diffusion-based posterior sampling for linear inverse problems via DDIM
A provable diffusion-based posterior sampling algorithm (pDDIM) for solving linear inverse problems via coordinate-wise DDIM updates.
Appearance Pointers -- Multimodal Region Control of Diffusion Transformers
Appearance Pointers enable precise regional control in Diffusion Transformers (DiTs) without needing to retrain the base model.
Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning
GEAR is a reward shaping method that reduces repetitive copying in long-context LLM reasoning by penalizing distractor overlap and rewarding evidence grounding.
Quality non-fiction books are the antithesis of AI slop
A discussion on how high-quality non-fiction books provide a depth and reliability that AI-generated content lacks.
Medici family mystery may be solved after more than 400 years
News regarding the potential resolution of a 400-year-old mystery involving the Medici family.
Any text-to-SQL benchmark should address difficulties of real-world data stores
A critique of current text-to-SQL benchmarks, arguing that they must account for the complexities and messy nature of real-world data stores.
Why do we love music? (2018)
An exploration of the biological and psychological reasons why humans love music.
Google justifies its massive AI spending with a booming cloud business
Google reports record profits driven by strong growth in its cloud business as companies adopt AI infrastructure services.
Meta won’t have to face the next planned social media addiction trial
Meta avoids a planned social media addiction trial after the plaintiff dropped the case.
Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasks
Researchers propose a robust framework for financial statement fraud detection using LLMs and introduce the CI-FSFD benchmark for better real-world generalization.