All Articles
16116 articles total
Show HN: What should the GUI for AI agents look like?
A community discussion and showcase regarding the design and user interface for AI agents.
The mean means nothing: data visualization to debug a latency problem
A technical guide on using data visualization to debug latency problems, arguing that the mean value is an insufficient metric.
AgentMap: Joint Equivalence and Subsumption Discovery for Ontology Matching
Introduction of AgentMap, a multi-agent LLM framework for Hybrid Ontology Matching that discovers both equivalence and subsumption mappings.
Linguistic Monoculture in LLM-Assisted Language Use
A research paper analyzing 'linguistic monoculture,' where widespread LLM use potentially reduces diversity in human language and communication.
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
OmegaUse-OfficeVal is a new benchmark for evaluating LLM agents on long-horizon office-suite tasks with economic grounding.
Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork
Proposes CE-CM, a Bayesian method for autonomous agents to estimate the capabilities of novel partners in ad-hoc teamwork settings.
Can AI agents conduct open-ended AI research? Early evidence from two case studies
A study revealing that while AI agents can handle the engineering side of AI research, they currently fail at open-ended research and critical thinking.
A Methodology for Designing Knowledge-Driven Missions for Robots
Presents a methodology for integrating knowledge graphs into ROS 2 systems to enhance autonomous robotic mission planning.
Where USB Memory Sticks Are Born
An exploration into the manufacturing process of USB memory sticks.
AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution
Introduces AgenticCANN, a framework for automated Ascend C operator synthesis to optimize NPU inference performance.
UrbanDS: A Graph-Guided LLM Multi-Agent System for Data-Intensive Urban Tasks
Presents UrbanDS, a graph-guided multi-agent system designed to automate data-intensive urban data science tasks.
Do Latent Channels Actually Communicate? A Causal Audit of Latent Multi-Agent LLM
A causal audit of latent communication in multi-agent LLM systems to determine if receivers actually use task-relevant information.
Property-driven Causal Abstractions for Markov Decision Processes
Introduces a property-driven causal abstraction technique to reduce state space blowup in Markov Decision Processes (MDPs).
From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence
Introduces Pegasus, a framework that converts human manipulation videos into robot-learnable data via structured knowledge transfer.
What Does It Take to Detect an AI Agent? Minimal Feature Sets for Behavioral Detection under Browser Automation
Researches the detection of AI agents using browser automation, identifying key behavioral features that distinguish them from humans.
Belief-Guided Decision Making with Uncertainty Gating in the Game of Go
Proposes a Belief-Guided architecture for Computer Go to reduce reliance on costly MCTS search on consumer-grade hardware.
Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data
Introduces Setoka, a benchmark for evaluating hierarchical user understanding in personalized agents over heterogeneous data.
On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment
Proposes Routing-based On-Policy Distillation (ROPD) for LLM safety realignment to mitigate template-mismatch risks and jailbreaking.
Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?
Researchers introduce StatMechBench-v0 to test if LLM-based agents can discover statistical mechanical mappings in theoretical physics, finding that while numerical feedback helps, agents still struggle with structural discovery.
CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games
CaM-Wolf is a multimodal AI agent for social deduction games like Werewolf that uses video perception and causal-aware reasoning to better simulate human social interaction.