All Articles
17740 articles total
When Summaries Distort Decisions: Information Fidelity in LLM-Compressed Financial Analysis
Researchers investigate how LLM-based context compression in financial analysis can distort decision-making by removing critical caveats and qualifiers.
The Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth Scaling
The Complexity Ceiling Benchmark (CCB) evaluates how LLM reasoning performance decays as the number of sequential steps increases across multiple domains.
Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners
Introduces PASS, a middleware for process-supervised reinforcement learning that addresses structural pathologies in Group Relative Policy Optimization (GRPO) for LLM reasoners.
Hierarchical Experimentalist Agents
HExA is a training-free framework that enables agents to learn from active experimentation and build reusable skill libraries to solve complex long-horizon tasks.
PHF: Privileged Hidden Flow for On-Policy Self-Distillation
Proposes Privileged Hidden Flow (PHF), a method for on-policy self-distillation that aligns teacher-student hidden state trajectories rather than just output distributions.
When LLMs Develop Languages: Symbolic Communication for Efficient Multi-Agent Reasoning
Introduces Communicative Language Symbolism Routing (CLSR), allowing multi-agent LLMs to invent compact symbolic languages to reduce token costs and latency.
Diagnosing and Repairing Factual Errors in RAG under Budget Constraints
D2R-RAG is a resource-aware framework that diagnoses and repairs factual errors in RAG systems under specific latency and VRAM constraints.
LLM-Guided Planning for Multi-hop Reasoning over Multimodal Nuclear Regulatory Documents
A planning-based LLM agent system for multi-hop reasoning over multimodal nuclear regulatory documents, outperforming standard RAG approaches.
Mixture of Debaters: Learn to Debate at Architectural Level in Multi-Agent Reasoning
Introduces Mixture of Debaters (MoD), using an MoE-style architecture to enable a single model to perform dynamic self-debate with lower latency and token use.
FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models
FADE is a training-free method to mitigate hallucinations in Large Vision-Language Models by attenuating FFN outputs to reduce language-prior dominance.
Open Source Low Tech
A discussion on Hacker News regarding the concept of Open Source Low Tech.
Pooled Leaderboards Hide System-Specific Winners: A Reporting-Protocol Audit of Offline Root-Cause Analysis Benchmarks
An audit of offline root-cause analysis benchmarks showing that pooled leaderboards often hide system-specific winners, potentially leading to poor recommendations.
Direct Causation in International Humanitarian Law and the Challenge of AI-Mediated Civilian Cyber Operations
An analysis of how AI-mediated civilian cyber operations challenge the direct causation element of International Humanitarian Law.
Selective Memory Retention for Long-Horizon LLM Agents
Introduction of TraceRetain, a lightweight framework for bounded external memory in LLM agents to reduce memory pollution and improve efficiency.
Measuring Graph-to-Graph Semantic Similarity in Knowledge Graphs: An Empirical Evaluation of Knowledge Graph Embeddings
An empirical evaluation of knowledge graph embeddings for measuring graph-to-graph semantic similarity, introducing the EmbPairSim and AvgEmbSim scoring functions.
Evidence-Informed LLM Beliefs for Continual Scientific Discovery
A proposal for evidence-informed LLM beliefs to improve continual scientific discovery by updating priors based on evidence from previous hypotheses.
AI Trading's Alpha Singularity: Emergent Market Reasoning through Agent-to-Agent Self-Evolution
The SJS framework and Agora system demonstrate emergent market reasoning in AI trading through agent-to-agent self-evolution.
A Cognition-Emotion-Personality Framework for Modeling Human-Like Awareness and Behavior in Emergency Evacuations
A cognition-emotion-personality framework for modeling more realistic human behavior and awareness in emergency evacuation simulations.
PolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM Agents
Introduction of PolicyGuard, a dialogue-grounded sub-agent verifier that improves policy adherence in LLM agents by reasoning over conversation context.
SurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical Robotics
SurgVLA-Bench is introduced as the first comprehensive benchmark for evaluating Vision-Language-Action models in laparoscopic surgical robotics.