All Articles
17040 articles total
A Multi-Agent System for Autonomous, Fine-Tuning-Free Clinical Symptom Detection: Development and Validation Study
The Pythia multi-agent system uses autonomous prompt optimization to extract clinical symptoms from notes without manual fine-tuning, maintaining high specificity.
MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon Conversations
The MemOps benchmark reformulates conversational memory evaluation for LLM agents as a sequence of lifecycle operations rather than simple question-answering tasks.
Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes
Researchers propose KGRL, a neuro-symbolic reinforcement learning algorithm that uses Datalog knowledge bases to improve sample efficiency and provide explainable decision pruning.
The bread paradox: why convenience always wins, and why SaaS isn't doomed
A discussion on why convenience continues to drive market success, arguing that SaaS is not doomed despite current industry skepticism.
The Model Knows Your Project, Not You: Measuring Recognition in LLMs with NameRank
Introduces NameRank, a metric to measure how well LLMs recognize specific entities based on their weights rather than external retrieval.
Evidence-Grounded AI for Musculoskeletal Care
Presents OrthoPilot, an AI system for longitudinal musculoskeletal care that outperforms expert physicians in diagnostic reasoning and management.
Vertical Standardisation for High-Risk AI Systems under the EU AI Act: A Domain-Specific Framework for Algorithmic Hiring
Proposes a domain-specific framework for standardizing algorithmic hiring systems to comply with the EU AI Act.
Agentic Service-Oriented Computing: A Manifesto for the Next Frontier of Service-Oriented Computing
Introduces Agentic Service-Oriented Computing (ASOC), a framework for engineering LLM agents as dependable enterprise services.
Atomic Units of X: The Compression Layer of Intelligence
Proposes a theoretical 'Compression Calculus' to understand intelligence as a process of atomic compression and compositional reuse.
A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism
Investigates the failure of GRPO in small web agents, finding it doesn't improve success rates unless the sampled policy already succeeds.
Internet of Agentic Things: Networked AI Agents for Closed-Loop IoT Orchestration
Introduces the Internet of Agentic Things (IoAT), a framework linking AI agents with IoT and edge computing for closed-loop orchestration.
MaxSAT-Based Feedback for Guiding Vision-Language Models in Sudoku
A neuro-symbolic approach using a MaxSAT oracle to provide logical feedback to VLMs for solving Sudoku puzzles.
LLMs Can See the Smoke but not the Fire: Evaluating Abductive Reasoning with Elenchos
Introduces Elenchos, a framework to evaluate the abductive reasoning capabilities of LLMs by detecting mutations in formal systems.
How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks
A replay analysis of LLM agent benchmarks investigating the minimum task fraction required to maintain the same pairwise performance conclusions as full runs.
PM-Bench: Evaluating Prospective Memory in LLM Agents
Introduction of PM-Bench, a text-based benchmark designed to evaluate prospective memory—the ability to execute delayed intentions—in LLM agents.
Critic Experience Bank: Self-Evolving Step-Level Confidence Estimation for LLM Agents
Proposes the Critic Experience Bank (CEB), a training-free framework that uses a memory bank of past judgments to improve step-level confidence estimation for LLM agents.
Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions
A survey paper proposing isolation as a first-class principle for LLM-agent system safety, providing a taxonomy of five critical boundaries to prevent failure propagation.
Accepted Prefixes Are Not All You Need: A Negative Result on PEFT-Based Block-Diffusion Drafting
A negative result study showing that PEFT-based block-diffusion drafting for speculative decoding fails to provide practical speedups because it is not compute-efficient.
EVOQUANT: Self-Evolving Verifier-Guided Strategy Optimization for Robust Quantitative Trading
Introduces EVOQUANT, an LLM-guided framework for automated quantitative trading strategy optimization that uses a verifier-guided loop to improve Sharpe ratios.
Do We Really Need Transformers for Global Spatial Information Extraction in Traffic Forecasting?
Research suggesting that simple global aggregation operators can perform as well as complex Transformers for spatial information extraction in traffic forecasting with lower complexity.