AI/ML arXiv cs.AI

Voice Memory for Agentic Speech Recognition

Voice Memory introduces a listener-thinker architecture for speech recognition that uses an auditable memory file instead of weight updates for correction.

AI/ML arXiv cs.AI

FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing

FAS-R1 is a multi-task MLLM for face anti-spoofing that combines a long-CoT dataset with GRPO post-training for better reasoning.

Software Engineering Hacker News

Google fixed more Chrome bugs in June than over the past two years, thanks to AI

Google reports a significant increase in Chrome bug fixes in June, attributing the surge to the use of AI.

AI/ML arXiv cs.AI

Entity Resolution in Practice: Lessons from a Self-Serve Pipeline

Researchers share practical lessons from building a self-serve entity resolution system, emphasizing the need for multiple algorithms and separate precision/recall fixes.

AI/ML arXiv cs.AI

AgentGUI: An Interface for Observing and Steering Long-Running AI Agents

AgentGUI is introduced as a locally hosted interface for observing and steering long-running AI agents with rich trajectory visualizations.

AI/ML arXiv cs.AI

SARC-DQ: Runtime Data-Quality Gating for Agentic AI: Silent Evidence Defects, the Incompetence Shield, and Downstream-Only Remediation

The SARC-DQ framework addresses data-quality gating for agentic AI to prevent costly actions based on stale or incorrect metadata.

Cybersecurity arXiv cs.AI

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents

StealthBench is a new benchmark for measuring operational stealth and OPSEC tradecraft in autonomous offensive-security agents.

AI/ML arXiv cs.AI

Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach

Cognitive Diagnostic Profiling (CDP) is proposed to align LLM-simulated examinees with real human psychometric data for better test calibration.

AI/ML arXiv cs.AI

Automorphism-Induced Non-Canonicity in Top-k Explanations of Graph Neural Networks

Researchers identify a structural obstruction in Graph Neural Network explainers where automorphisms lead to arbitrary top-k edge reports.

AI/ML arXiv cs.AI

When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses

A cross-domain benchmark reveals that LLM-simulated synthetic users fail to accurately replicate human survey responses, especially across cultures.

Software Engineering arXiv cs.AI

Pramana: A Composable, Domain-Specific Backend for Empirical Networking Research

Pramana is introduced as a composable backend designed to accelerate empirical networking research by bridging the gap between hypothesis and data generation.

AI/ML arXiv cs.AI

High-Order Markov Blanket Discovery via a k-Order Relaxation of the Faithfulness Assumption

The k-order Markov blanket (kOMB) algorithm is proposed to discover graphical Markov blankets while relaxing the faithfulness assumption to handle parity-type relations.

Other Hacker News

"the very foundation of modern academia has been blown to bits"

A Hacker News discussion regarding the perceived collapse of the foundations of modern academia due to external pressures or AI.

AI/ML arXiv cs.AI

Try Again, Don't Look Back: Blind Resampling Outperforms Self-Repair in Small Code Models

Research showing that 'blind resampling' (retrying without looking at previous failures) is more efficient and often more effective than self-repair in small code LLMs.

AI/ML arXiv cs.AI

Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels

A proposed systems framework and graded trustworthiness levels for embodied AI to ensure safe and reliable physical interaction.

AI/ML arXiv cs.AI

A Picture Says Thousands of Words - Harnessing Dermal Exposure Data from Images through Hybrid Deep Learning for Enhanced Safety Assessment

A hybrid deep learning approach using Mask R-CNN and color-based segmentation to quantify skin exposure for safety assessments.

AI/ML arXiv cs.AI

Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition

An analysis of the structural and cognitive convergence between Large Language Models and human cognitive organization.

Cybersecurity arXiv cs.AI

(EC)2: Event-Centric Explainability for Cybersecurity Through Multi-Agent LLM Investigations

Introduction of (EC)2, a multi-agent LLM framework for providing event-centric, verifiable explanations for cybersecurity alerts.

AI/ML arXiv cs.AI

Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges

A systematic survey and taxonomy of Multi-Agent Debate (MAD) strategies to improve LLM robustness and accuracy.

Software Engineering arXiv cs.AI

Model-Driven Requirements Configuration with Three-Valued Uncertainty Scoring

A neuro-symbolic multi-agent architecture that combines LLMs with a deterministic symbolic validator to ensure structural integrity in requirements engineering.