Open Source arXiv cs.AI

When Bots Join the Team: Bot Adoption and the Institutional Fabric of Open-Source Software Projects

A study examining how the introduction of AI bots into open-source software projects affects group coordination and social infrastructure.

AI/ML arXiv cs.AI

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

AgentCompass is introduced as an open-source, extensible infrastructure for the unified evaluation of LLM-based agents.

Cybersecurity arXiv cs.AI

CAVA: Canonical Action Verification and Attestation for Runtime Governance of Agentic AI Systems

A framework called CAVA is proposed to provide canonical runtime action verification and attestation for agentic AI systems.

AI/ML arXiv cs.AI

Experience Memory Graph: One-Shot Error Correction for Agents

The Experience Memory Graph (EMG) framework aims to improve LLM agent recovery from failures by using graph matching to extract successful workflows.

AI/ML arXiv cs.AI

AIMO Interpretability Challenge

The AIMO Interpretability Challenge is an upcoming competition for identifying robust reasoning versus spurious shortcuts in mathematical language models.

AI/ML arXiv cs.AI

A Self-Evolving Agent for Longitudinal Personal Health Management

The HealthClaw architecture is an open-source agent for longitudinal personal health management that uses self-evolving memory.

Hardware/Chips Hacker News

Teardown: A Generic 7-Port USB 3.0 Hub That Wasn't

A teardown of a generic 7-port USB 3.0 hub reveals significant hardware quality issues.

Other Hacker News

Reynard: A real Firefox web browser for iOS 13 or later

Reynard is a real Firefox web browser implementation for iOS 13 and later.

AI/ML arXiv cs.AI

EZSMT Version 3, Matured

The paper introduces EZSMTV3, an extensible SMT-based CASP framework for complex combinatorial search problems.

AI/ML arXiv cs.AI

Set-shifting Behavioral Test for Harnessed Agents

Researchers study how LLM agents adapt to hidden reliability shifts in their toolsets using a set-shifting behavioral test.

AI/ML arXiv cs.AI

LAPO: Leave-One-Turn Attribution for Self-Generated Process Rewards in Multi-Turn Search Reasoning

The paper proposes LAPO, a self-generated process-supervision method to improve multi-turn reasoning in LLM agents without requiring a teacher model.

AI/ML arXiv cs.AI

How Far Can Root Cause Analysis Go on Real-World Telemetry Data?

The study explores the effectiveness of root cause analysis on real-world telemetry data, finding that agentic reasoning is the primary bottleneck.

AI/ML arXiv cs.AI

Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region Profiling

UrbanAgent is a multi-agent collaborative framework that uses tool-augmented reasoning for urban region profiling.

Other arXiv cs.AI

AI advice suppresses people's willingness to say "I don't know", even when the advice is wrong and accuracy is incentivized

A study shows that AI advice can suppress a person's willingness to say "I don't know", leading to decreased accuracy and increased confidence.

AI/ML arXiv cs.AI

SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing

Safety Sentry is a lightweight guard model that uses a three-way routing decision (EXECUTE, ASK, REFUSE) to manage LLM agent tool calls safely.

AI/ML arXiv cs.AI

Automatic Ordinary Differential Equations Discovery For Biological Systems Using Large Language Model Powered Agentic System

The MEDA system is an LLM and symbolic regression-powered agentic framework for discovering ODE models of biological systems.

Other Hacker News

The lost joy of music piracy

A discussion on the cultural and practical shift from the era of music piracy to the current streaming model.

AI/ML Hacker News

Can LLMs Perform Deep Technical Comprehension of Computer Architecture Papers

An exploration of whether Large Language Models can truly comprehend deep technical details of computer architecture papers.

Software Engineering Hacker News

DJB Netstrings (1997)

A look back at DJB Netstrings, a simple and efficient string representation format from 1997.

AI/ML Hacker News

Stop saying that AI is just a tool and it only matters how it is used

A critique of the common narrative that AI is merely a tool, arguing instead for its more complex nature.