All Articles
19339 articles total
Curl will not accept vulnerability reports during July 2026
The curl project will pause the acceptance of vulnerability reports in July 2026.
The Last Surviving Japanese Porsche 912 Police Car
An article discussing the last surviving Japanese Porsche 912 Police Car.
Apple Foundation Models
A discussion regarding the foundation models used by Apple.
StreamMemBench: Streaming Evaluation of Agent Memory for Future-Oriented Assistance
Introduction of StreamMemBench, a streaming benchmark to evaluate agent memory's ability to provide future-oriented assistance based on streaming observations.
VISTA: View-Consistent Self-Verified Training for GUI Grounding
VISTA is a GRPO-based training framework that improves GUI grounding accuracy by using multiple target-preserving views for self-verification.
A Temporal Planning Framework for Disruption Aware Dynamic Route Optimization in Heterogeneous Railway Systems
A temporal planning framework using PDDL 2.1 for dynamic route optimization and disruption management in heterogeneous railway systems.
Abstracting Cross-Domain Action Sequences into Interpretable Workflows
WorkflowView uses LLMs to abstract low-level interaction logs into interpretable high-level activities across various digital domains.
Towards Direct Latent-Space Synthesis for Parallel Branches in LLM-Agent Workflows
Parallel-Synthesis is a framework that allows synthesizers to directly consume KV caches from parallel agent branches, significantly reducing time-to-first-token.
GAGPO: Generalized Advantage Grouped Policy Optimization
GAGPO is a critic-free reinforcement learning method designed for precise step-aligned temporal credit assignment in multi-turn LLM agent environments.
Simplex-Constrained Sparse Bagging: Transitioning from Uniform Priors to Sparse Posteriors in Ensemble Learning
SCSB is a mathematically rigorous framework for post-training compression and probability calibration of bootstrap-based bagging ensembles.
AFFORDANCE20Q: Evaluating Affordance Reasoning from Physical Properties
Researchers introduce Affordance20Q, a benchmark that tests if LLMs can reason about an object's physical properties to determine its uses without knowing the object's identity.
HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
HarnessX is proposed as a foundry for creating adaptive agent harnesses that evolve based on execution traces to improve AI agent performance.
Communication Policy Evolution for Proactive LLM Agents
This paper explores communication policies for proactive LLM agents and proposes a self-evolution framework called CPE to refine these policies via prompt refinement.
CSPO: Constraint-Sensitive Policy Optimization for Safe Reinforcement Learning
The authors propose CSPO, a first-order primal-dual method for Safe Reinforcement Learning that reduces oscillations and accelerates recovery to safety boundaries.
Causal Object-Centric Models for Planning with Monte Carlo Tree Search
COMET is a model-based RL algorithm that uses a slot-structured latent space and Monte Carlo Tree Search to improve planning in object-centric environments.
GitOfThoughts: Version-Controlled Reasoning and Agent Memory You Can Replay, Diff, and Merge
GitOfThoughts introduces a version-control system for LLM reasoning trees using Git, though the study finds that memory formats don't significantly improve accuracy for novel problems.
When the Tool Decides: LLM Agents Defer Blindly to Graph Neural Network Tools, and Stronger Backbones Defer More
Research shows that LLM agents tend to blindly follow the outputs of Graph Neural Network tools rather than exercising judgment, even as model capability increases.
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI
The paper discusses a paradigm shift from simple chatbots to persistent autonomous AI 'Digital Colleagues' using a 'Workspace + Skill' architecture.
Dense Coordinate-List Fine-Tuning Induces a Controllable Interference Surface in Vision-Language Models
Study on Vision-Language Models finds that fine-tuning for dense coordinate lists creates a controllable interference surface that affects output serialization.
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
Every Eval Ever introduces a unifying schema and crowdsourced repository on Hugging Face to standardize AI evaluation results across diverse benchmarks.