All Articles
16960 articles total
When Bots Join the Team: Bot Adoption and the Institutional Fabric of Open-Source Software Projects
A study examining how the introduction of AI bots into open-source software projects affects group coordination and social infrastructure.
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
AgentCompass is introduced as an open-source, extensible infrastructure for the unified evaluation of LLM-based agents.
CAVA: Canonical Action Verification and Attestation for Runtime Governance of Agentic AI Systems
A framework called CAVA is proposed to provide canonical runtime action verification and attestation for agentic AI systems.
Experience Memory Graph: One-Shot Error Correction for Agents
The Experience Memory Graph (EMG) framework aims to improve LLM agent recovery from failures by using graph matching to extract successful workflows.
AIMO Interpretability Challenge
The AIMO Interpretability Challenge is an upcoming competition for identifying robust reasoning versus spurious shortcuts in mathematical language models.
A Self-Evolving Agent for Longitudinal Personal Health Management
The HealthClaw architecture is an open-source agent for longitudinal personal health management that uses self-evolving memory.
Teardown: A Generic 7-Port USB 3.0 Hub That Wasn't
A teardown of a generic 7-port USB 3.0 hub reveals significant hardware quality issues.
Reynard: A real Firefox web browser for iOS 13 or later
Reynard is a real Firefox web browser implementation for iOS 13 and later.
EZSMT Version 3, Matured
The paper introduces EZSMTV3, an extensible SMT-based CASP framework for complex combinatorial search problems.
Set-shifting Behavioral Test for Harnessed Agents
Researchers study how LLM agents adapt to hidden reliability shifts in their toolsets using a set-shifting behavioral test.
LAPO: Leave-One-Turn Attribution for Self-Generated Process Rewards in Multi-Turn Search Reasoning
The paper proposes LAPO, a self-generated process-supervision method to improve multi-turn reasoning in LLM agents without requiring a teacher model.
How Far Can Root Cause Analysis Go on Real-World Telemetry Data?
The study explores the effectiveness of root cause analysis on real-world telemetry data, finding that agentic reasoning is the primary bottleneck.
Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region Profiling
UrbanAgent is a multi-agent collaborative framework that uses tool-augmented reasoning for urban region profiling.
AI advice suppresses people's willingness to say "I don't know", even when the advice is wrong and accuracy is incentivized
A study shows that AI advice can suppress a person's willingness to say "I don't know", leading to decreased accuracy and increased confidence.
SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing
Safety Sentry is a lightweight guard model that uses a three-way routing decision (EXECUTE, ASK, REFUSE) to manage LLM agent tool calls safely.
Automatic Ordinary Differential Equations Discovery For Biological Systems Using Large Language Model Powered Agentic System
The MEDA system is an LLM and symbolic regression-powered agentic framework for discovering ODE models of biological systems.
The lost joy of music piracy
A discussion on the cultural and practical shift from the era of music piracy to the current streaming model.
Can LLMs Perform Deep Technical Comprehension of Computer Architecture Papers
An exploration of whether Large Language Models can truly comprehend deep technical details of computer architecture papers.
DJB Netstrings (1997)
A look back at DJB Netstrings, a simple and efficient string representation format from 1997.
Stop saying that AI is just a tool and it only matters how it is used
A critique of the common narrative that AI is merely a tool, arguing instead for its more complex nature.