AI/ML arXiv cs.AI

A Multi-Agent System for Autonomous, Fine-Tuning-Free Clinical Symptom Detection: Development and Validation Study

The Pythia multi-agent system uses autonomous prompt optimization to extract clinical symptoms from notes without manual fine-tuning, maintaining high specificity.

AI/ML arXiv cs.AI

MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon Conversations

The MemOps benchmark reformulates conversational memory evaluation for LLM agents as a sequence of lifecycle operations rather than simple question-answering tasks.

AI/ML arXiv cs.AI

Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes

Researchers propose KGRL, a neuro-symbolic reinforcement learning algorithm that uses Datalog knowledge bases to improve sample efficiency and provide explainable decision pruning.

Tech Business/VC Hacker News

The bread paradox: why convenience always wins, and why SaaS isn't doomed

A discussion on why convenience continues to drive market success, arguing that SaaS is not doomed despite current industry skepticism.

AI/ML arXiv cs.AI

The Model Knows Your Project, Not You: Measuring Recognition in LLMs with NameRank

Introduces NameRank, a metric to measure how well LLMs recognize specific entities based on their weights rather than external retrieval.

AI/ML arXiv cs.AI

Evidence-Grounded AI for Musculoskeletal Care

Presents OrthoPilot, an AI system for longitudinal musculoskeletal care that outperforms expert physicians in diagnostic reasoning and management.

Other arXiv cs.AI

Vertical Standardisation for High-Risk AI Systems under the EU AI Act: A Domain-Specific Framework for Algorithmic Hiring

Proposes a domain-specific framework for standardizing algorithmic hiring systems to comply with the EU AI Act.

Software Engineering arXiv cs.AI

Agentic Service-Oriented Computing: A Manifesto for the Next Frontier of Service-Oriented Computing

Introduces Agentic Service-Oriented Computing (ASOC), a framework for engineering LLM agents as dependable enterprise services.

AI/ML arXiv cs.AI

Atomic Units of X: The Compression Layer of Intelligence

Proposes a theoretical 'Compression Calculus' to understand intelligence as a process of atomic compression and compositional reuse.

AI/ML arXiv cs.AI

A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism

Investigates the failure of GRPO in small web agents, finding it doesn't improve success rates unless the sampled policy already succeeds.

AI/ML arXiv cs.AI

Internet of Agentic Things: Networked AI Agents for Closed-Loop IoT Orchestration

Introduces the Internet of Agentic Things (IoAT), a framework linking AI agents with IoT and edge computing for closed-loop orchestration.

AI/ML arXiv cs.AI

MaxSAT-Based Feedback for Guiding Vision-Language Models in Sudoku

A neuro-symbolic approach using a MaxSAT oracle to provide logical feedback to VLMs for solving Sudoku puzzles.

AI/ML arXiv cs.AI

LLMs Can See the Smoke but not the Fire: Evaluating Abductive Reasoning with Elenchos

Introduces Elenchos, a framework to evaluate the abductive reasoning capabilities of LLMs by detecting mutations in formal systems.

AI/ML arXiv cs.AI

How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks

A replay analysis of LLM agent benchmarks investigating the minimum task fraction required to maintain the same pairwise performance conclusions as full runs.

AI/ML arXiv cs.AI

PM-Bench: Evaluating Prospective Memory in LLM Agents

Introduction of PM-Bench, a text-based benchmark designed to evaluate prospective memory—the ability to execute delayed intentions—in LLM agents.

AI/ML arXiv cs.AI

Critic Experience Bank: Self-Evolving Step-Level Confidence Estimation for LLM Agents

Proposes the Critic Experience Bank (CEB), a training-free framework that uses a memory bank of past judgments to improve step-level confidence estimation for LLM agents.

Cybersecurity arXiv cs.AI

Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions

A survey paper proposing isolation as a first-class principle for LLM-agent system safety, providing a taxonomy of five critical boundaries to prevent failure propagation.

AI/ML arXiv cs.AI

Accepted Prefixes Are Not All You Need: A Negative Result on PEFT-Based Block-Diffusion Drafting

A negative result study showing that PEFT-based block-diffusion drafting for speculative decoding fails to provide practical speedups because it is not compute-efficient.

AI/ML arXiv cs.AI

EVOQUANT: Self-Evolving Verifier-Guided Strategy Optimization for Robust Quantitative Trading

Introduces EVOQUANT, an LLM-guided framework for automated quantitative trading strategy optimization that uses a verifier-guided loop to improve Sharpe ratios.

AI/ML arXiv cs.AI

Do We Really Need Transformers for Global Spatial Information Extraction in Traffic Forecasting?

Research suggesting that simple global aggregation operators can perform as well as complex Transformers for spatial information extraction in traffic forecasting with lower complexity.