All Articles
15966 articles total
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL
Proposes PRISM, a multi-reward RL framework that optimizes standalone positive and negative policies to reduce conflict and alignment tax in LLM post-training.
Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents
Identifies safety risks in AI agent tool specifications and introduces SafeKeep, an inference-time safeguard that decouples safety judgment from tool execution.
MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation
Introduces MAGA, a method for multi-platform self-fusion of GUI agents via structured action distillation to create a single cross-environment policy.
Beyond Component Testing: Validating Agentic AI Systems
A survey of 257 papers on validating agentic AI systems, emphasizing the need to validate trajectories in context rather than isolated components.
ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models
Presents ModelEquivBench, a multi-relational evaluation system for certifying the equivalence of LLM-generated optimization models.
Beyond Retrieval: Analytic Memory for Multimodal Agents
Introduces AdaMM, a framework that combines retrieval memory with analytic memory for multimodal agents to support complex queries over observations.
Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember
Presents SESA, a self-evolving search agent that co-evolves task generation and procedural skill memory through a self-play loop.
Fragility of Value under Imperfect Alignment
This paper examines the risk of AI systems optimizing for imperfect proxies of human values, potentially leading to catastrophic outcomes if optimization pressure is too high.
Identifying Informative Environments for Cognition Parameter Inference via Bayesian Experimental Design
The authors propose a Bayesian Experimental Design (BED) framework to identify the most informative environments for inferring cognitive parameters in computational modeling.
NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability
NeSyFS is a neuro-symbolic framework that uses knowledge graphs and a twisted sequential Monte Carlo algorithm to help LLM agents handle partial observability.
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
MerchantBench is a new 365-day simulation benchmark for testing the long-term coherence of LLM agents in complex e-commerce operations.
Scaling Scientific Discovery Environments for Turn-Level Agentic RL
SciDisco is a scalable framework for training scientific discovery agents using process-verifiable environments and turn-level credit assignment.
MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping Agents
MMShopBench introduces a real-log benchmark and an offline shopping sandbox for evaluating multimodal, multi-turn shopping agents.
Evidence-Grounded Constraint Checking in Construction Documents
This research investigates the trade-offs between resolution and breadth when using RAG-based pipelines for evidence-grounded constraint checking in construction documents.
On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness
The study explores how activation steering vectors can improve the faithfulness of Chain-of-Thought reasoning across different LLMs.
A Generalized-Bayes Perspective on Counterfactual Explanations: Posterior-Based Decision-Making and Evaluation
The authors present a Generalized-Bayes perspective on counterfactual explanations, introducing new decision rules for more interpretable ML model outputs.
Harnessing the Wisdom of LLM Crowds through Complementarity-Driven Iterative Collaboration
WILC is a framework for coordinating multiple LLMs through iterative collaboration and complementarity-driven model selection to achieve higher collective intelligence.
OpenAI's super PAC is funding AI-generated news site attacking industry critics
OpenAI's super PAC is allegedly funding an AI-generated news site used to attack critics of the AI industry.
Cro – elegant reactive services in Raku
Introduction to Cro, a framework for building elegant reactive services using the Raku programming language.
AI migrated legacy COBOL programs to Java, bugs included
A report on the failure of AI to migrate legacy COBOL programs to Java without introducing bugs.