All Articles
15986 articles total
Multi-Agent Planning with Spatio-Temporal and Topological Constraints using STL-GO
A research paper presenting STL-GO, a formalism for multi-agent planning with spatio-temporal and topological constraints using MIP and SMT encodings.
Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic Discovery
An analysis of LSR-Synth, a benchmark designed to measure symbolic discovery in scientific equations while preventing model memorization.
Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks
A validity audit of current agent-safety benchmarks, arguing that many scores correlate more with general capability than actual safety.
SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition
Introduction of SciToolAgent-Evo, an agent that self-evolves to acquire scientific tools, and OpenSciToolBench, a new evaluation benchmark.
EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported ED-Encounter Diagnoses
EarlyDx is introduced as a benchmark for open-ended generation of evidence-supported emergency department encounter diagnoses.
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
A proposal for an interaction-centric taxonomy to better localize and repair failures in AI agents across different components.
Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions
A study using the ANCHOR audit to evaluate persona collapse and behavioral drift in long-horizon AI companions.
OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems
Proposes a layered architecture for Agentic AI using Ollama for inference and OpenClaw for orchestration to enable autonomous, goal-driven AI agents.
Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review
Introduces a benchmarking protocol using multi-model LLM peer-review to evaluate autonomous AI research generation systems.
LLM Framework for Discovering Major Mathematical Conjectures: AI's Quest for the Next Riemann Hypothesis
Presents a three-stage pipeline for discovering mathematical conjectures, utilizing Lean 4 for formal validation of the findings.
ThinkReset: Learnable Intermediate Interface Construction for Bounded-Context Long-Horizon Reasoning
Introduces ThinkReset, a method to handle bounded context windows in long-horizon reasoning by constructing reusable intermediate interfaces.
TAPR: Enhancing LLM Performance with a Task-Aware Prompt Rewriter
Presents TAPR, a task-aware prompt rewriter trained via GRPO to optimize user prompts for better LLM performance.
Empowering Cross-Domain Sequential Recommendation with Hybrid Tokenization and Serial-Parallel Decoding
Proposes GenCDSR, a generative framework for cross-domain sequential recommendation using hybrid tokenization and serial-parallel decoding.
An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous Documents
Describes a production extraction layer for constructing knowledge graphs from heterogeneous documents using ontology-guided extraction and a local Qwen3.5-9B model.
How Hard Does It Think? Analyzing Step-Aware Reasoning Energy in LLM Chain-of-Thought Trajectories
Introduces Step-Aware Reasoning Energy (SARE) to quantify computational effort at individual chain-of-thought reasoning steps in LLMs.
Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support
Argues that current LLMs are unsafe for autonomous clinical triage due to failures in gathering information under uncertainty and asymmetric cost risks.
ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding
Proposes ViSAGE, a multimodal memory framework that uses self-correcting, entity-centric memories to improve long-form video understanding.
Why we write our own C and C++ inference engines
A technical discussion on Hacker News regarding the motivations and technical challenges associated with developing custom C and C++ inference engines for AI models.
Apple engineer says he was fired after refusing to send cust. device IDs to AT&T
An Apple engineer claims he was terminated after refusing to share customer device IDs with AT&T, raising privacy concerns.
Qwen3.8-Max: A New Bar for Coding and Cowork
Alibaba introduces Qwen3.8-Max, a new large language model focused on enhancing coding capabilities and collaborative workflows.