AI/ML Hacker News

Show HN: What should the GUI for AI agents look like?

A community discussion and showcase regarding the design and user interface for AI agents.

Software Engineering Hacker News

The mean means nothing: data visualization to debug a latency problem

A technical guide on using data visualization to debug latency problems, arguing that the mean value is an insufficient metric.

AI/ML arXiv cs.AI

AgentMap: Joint Equivalence and Subsumption Discovery for Ontology Matching

Introduction of AgentMap, a multi-agent LLM framework for Hybrid Ontology Matching that discovers both equivalence and subsumption mappings.

AI/ML arXiv cs.AI

Linguistic Monoculture in LLM-Assisted Language Use

A research paper analyzing 'linguistic monoculture,' where widespread LLM use potentially reduces diversity in human language and communication.

AI/ML arXiv cs.AI

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

OmegaUse-OfficeVal is a new benchmark for evaluating LLM agents on long-horizon office-suite tasks with economic grounding.

AI/ML arXiv cs.AI

Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork

Proposes CE-CM, a Bayesian method for autonomous agents to estimate the capabilities of novel partners in ad-hoc teamwork settings.

AI/ML arXiv cs.AI

Can AI agents conduct open-ended AI research? Early evidence from two case studies

A study revealing that while AI agents can handle the engineering side of AI research, they currently fail at open-ended research and critical thinking.

AI/ML arXiv cs.AI

A Methodology for Designing Knowledge-Driven Missions for Robots

Presents a methodology for integrating knowledge graphs into ROS 2 systems to enhance autonomous robotic mission planning.

Hardware/Chips Hacker News

Where USB Memory Sticks Are Born

An exploration into the manufacturing process of USB memory sticks.

AI/ML arXiv cs.AI

AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution

Introduces AgenticCANN, a framework for automated Ascend C operator synthesis to optimize NPU inference performance.

AI/ML arXiv cs.AI

UrbanDS: A Graph-Guided LLM Multi-Agent System for Data-Intensive Urban Tasks

Presents UrbanDS, a graph-guided multi-agent system designed to automate data-intensive urban data science tasks.

AI/ML arXiv cs.AI

Do Latent Channels Actually Communicate? A Causal Audit of Latent Multi-Agent LLM

A causal audit of latent communication in multi-agent LLM systems to determine if receivers actually use task-relevant information.

AI/ML arXiv cs.AI

Property-driven Causal Abstractions for Markov Decision Processes

Introduces a property-driven causal abstraction technique to reduce state space blowup in Markov Decision Processes (MDPs).

AI/ML arXiv cs.AI

From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence

Introduces Pegasus, a framework that converts human manipulation videos into robot-learnable data via structured knowledge transfer.

Cybersecurity arXiv cs.AI

What Does It Take to Detect an AI Agent? Minimal Feature Sets for Behavioral Detection under Browser Automation

Researches the detection of AI agents using browser automation, identifying key behavioral features that distinguish them from humans.

AI/ML arXiv cs.AI

Belief-Guided Decision Making with Uncertainty Gating in the Game of Go

Proposes a Belief-Guided architecture for Computer Go to reduce reliance on costly MCTS search on consumer-grade hardware.

AI/ML arXiv cs.AI

Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data

Introduces Setoka, a benchmark for evaluating hierarchical user understanding in personalized agents over heterogeneous data.

AI/ML arXiv cs.AI

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

Proposes Routing-based On-Policy Distillation (ROPD) for LLM safety realignment to mitigate template-mismatch risks and jailbreaking.

AI/ML arXiv cs.AI

Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?

Researchers introduce StatMechBench-v0 to test if LLM-based agents can discover statistical mechanical mappings in theoretical physics, finding that while numerical feedback helps, agents still struggle with structural discovery.

AI/ML arXiv cs.AI

CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games

CaM-Wolf is a multimodal AI agent for social deduction games like Werewolf that uses video perception and causal-aware reasoning to better simulate human social interaction.