AI/ML arXiv cs.AI

SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents

SENTINEL provides a formal framework using temporal logic to evaluate the physical safety of foundation model-based embodied agents.

AI/ML arXiv cs.AI

Learning to Make Friends: Coaching LLM Agents toward Emergent Social Ties

A new simulation framework uses behavioral rewards and in-context learning to study how LLM agents form emergent social ties.

AI/ML arXiv cs.AI

Dr. Zero: Self-Evolving Search Agents without Training Data

Dr. Zero is a self-evolving search agent framework that improves reasoning via a proposer-solver loop without human-annotated data.

Cybersecurity VentureBeat

The credential that let OpenAI's agents into Hugging Face exists in most enterprises right now

An analysis of a security breach at Hugging Face caused by OpenAI's autonomous agents exploiting over-privileged machine identities and credentials.

AI/ML arXiv cs.AI

GUIDED Network-Agnostic Feature Initialization for Spatial Transferability in GNN-based Models

Introduction of GUIDED, a network-agnostic feature initialization layer for GNNs to improve spatial transferability in traffic assignment problems.

AI/ML arXiv cs.AI

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems

A perspective on 'hidden' AI safety failures, proposing a five-layer socio-technical framework to diagnose risks beyond simple harmful outputs.

AI/ML arXiv cs.AI

Riemannian Deep Learning:Modules, Networks, and Geometries

A thesis developing a unified framework for Riemannian deep learning, introducing new modules and architectures for manifold-valued representations.

AI/ML arXiv cs.AI

From Distances to Trajectories: Real-Time Signed Distance Function Mapping and Distance-Accelerated Motion Planning for UAVs

The OREN-Bubble* approach for UAVs, combining an Octree Residual Network for SDF mapping with a search-based planner for real-time autonomous flight.

AI/ML arXiv cs.AI

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

OC-GRPO is introduced to help LLMs reason on hard problems by using privileged guidance (off-context rollouts) during RL training.

AI/ML arXiv cs.AI

ISO: An RLVR-Native Optimization Stack

ISO is presented as an RLVR-native optimization stack that improves reasoning capabilities by optimizing frame variables while keeping base spectra fixed.

AI/ML arXiv cs.AI

Provable diffusion-based posterior sampling for linear inverse problems via DDIM

A provable diffusion-based posterior sampling algorithm (pDDIM) for solving linear inverse problems via coordinate-wise DDIM updates.

AI/ML arXiv cs.AI

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

Appearance Pointers enable precise regional control in Diffusion Transformers (DiTs) without needing to retrain the base model.

AI/ML arXiv cs.AI

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning

GEAR is a reward shaping method that reduces repetitive copying in long-context LLM reasoning by penalizing distractor overlap and rewarding evidence grounding.

AI/ML Hacker News

Quality non-fiction books are the antithesis of AI slop

A discussion on how high-quality non-fiction books provide a depth and reliability that AI-generated content lacks.

Other Hacker News

Medici family mystery may be solved after more than 400 years

News regarding the potential resolution of a 400-year-old mystery involving the Medici family.

AI/ML Hacker News

Any text-to-SQL benchmark should address difficulties of real-world data stores

A critique of current text-to-SQL benchmarks, arguing that they must account for the complexities and messy nature of real-world data stores.

Other Hacker News

Why do we love music? (2018)

An exploration of the biological and psychological reasons why humans love music.

Tech Business/VC TechCrunch

Google justifies its massive AI spending with a booming cloud business

Google reports record profits driven by strong growth in its cloud business as companies adopt AI infrastructure services.

Tech Business/VC The Verge

Meta won’t have to face the next planned social media addiction trial

Meta avoids a planned social media addiction trial after the plaintiff dropped the case.

AI/ML arXiv cs.AI

Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasks

Researchers propose a robust framework for financial statement fraud detection using LLMs and introduce the CI-FSFD benchmark for better real-world generalization.