All Articles
17661 articles total
Database Traffic Control
A discussion on managing and controlling database traffic to ensure system stability and performance.
Learn Vim motions with an ice-cream van
A creative and interactive tool designed to help users learn Vim keyboard motions using an ice-cream van theme.
Mnemosyne: Agentic Transaction Processing for Validating and Repairing AI-generated Workflows
Introduces Mnemosyne, a runtime using Agentic Transaction Processing to validate and repair AI-generated workflows via executable constraints.
Managed Autonomy at Runtime: Gear-Based Safety and Governance for Single- and Multi-Agent Cyber-Physical Systems
Proposes a gear-based safety and governance system for single and multi-agent cyber-physical systems to prevent safety violations and instability.
Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising
Presents SPIRE, a framework that treats personalized slide generation as an inverse planning problem using structural denoising and RL.
PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents
Introduces PHREEQC-MCQ-200, a diagnostic benchmark for evaluating tool-augmented agents in aqueous-geochemistry simulations.
Agri-SAGE: Simulation-Grounded Multi-Agent LLM for Context-Aware Agricultural Advisory Generation
Describes Agri-SAGE, a multi-agent LLM framework that integrates biophysical simulation for context-aware agricultural advisory generation.
Multi-scale Mixture of World Models for Embodied Agents in Evolving Environments
Introduces MuSix, a framework for embodied agents using a scale-aware mixture of world models to handle evolving environments.
AI Native Games: A Survey and Roadmap
A survey and roadmap for AI-native games, defining games where generative AI is constitutive of the core gameplay loop.
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment
Introduces HARC, a fine-tuning method that couples harmfulness and refusal directions to improve LLM safety alignment and robustness against jailbreaks.
Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction
Introduces 'Constructive Alignment', a framework treating AI alignment as a control problem over evolving human preferences rather than static goals.
Bounded Morality: Defining the Space of Moral Computation
Proposes 'Bounded Morality', a formal framework that analyzes the computational limits of moral reasoning for finite agents.
The MMM Data Model -- A Normative Specification for Knowledge Interoperability in a Decentralisable Knowledge Commons
Presents the MMM data model, a normative specification designed for knowledge interoperability in decentralisable knowledge commons.
Making Failure Safe: A Constrained, Verifiable Agent Framework for Open-Web Data Collection
Introduces a constrained, verifiable agent framework that uses typed JSON configurations instead of free-form code for reliable open-web data collection.
Solution space path planning for supporting en-route air traffic control
Develops a conflict-free path-planning algorithm for air traffic control designed for human interpretability and computational efficiency.
RareDxR1: Autonomous Medical Reasoning for Rare Disease Diagnosis Beyond Human Annotation
Presents RareDxR1, an end-to-end reasoning LLM for rare disease diagnosis that uses autonomous evolutionary learning to bypass human annotation.
A Contextual-Bandit Oversight Game with Two-Sided Informational Asymmetry
Analyzes a contextual-bandit oversight game to study human-AI interaction under two-sided informational asymmetry.
Constructing Epistemic AI Literacy: Detecting Epistemic Aims and Processes in Student-AI Co-Programming
Introduces 'Epistemic AI Literacy' (EAIL) to analyze how students use GenAI for co-programming and evaluate their epistemic processes.
From Signals to Structure: How Memory Architecture Drives Language Emergence in LLM Agents
Explores how memory architecture in LLM agents affects the emergence of shared languages through a Lewis signaling game.
Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity
Presents the Seed2.0 model card, highlighting improvements in long-tail knowledge and complex instruction following for real-world tasks.