All Articles
16106 articles total
Voice Memory for Agentic Speech Recognition
Voice Memory introduces a listener-thinker architecture for speech recognition that uses an auditable memory file instead of weight updates for correction.
FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing
FAS-R1 is a multi-task MLLM for face anti-spoofing that combines a long-CoT dataset with GRPO post-training for better reasoning.
Google fixed more Chrome bugs in June than over the past two years, thanks to AI
Google reports a significant increase in Chrome bug fixes in June, attributing the surge to the use of AI.
Entity Resolution in Practice: Lessons from a Self-Serve Pipeline
Researchers share practical lessons from building a self-serve entity resolution system, emphasizing the need for multiple algorithms and separate precision/recall fixes.
AgentGUI: An Interface for Observing and Steering Long-Running AI Agents
AgentGUI is introduced as a locally hosted interface for observing and steering long-running AI agents with rich trajectory visualizations.
SARC-DQ: Runtime Data-Quality Gating for Agentic AI: Silent Evidence Defects, the Incompetence Shield, and Downstream-Only Remediation
The SARC-DQ framework addresses data-quality gating for agentic AI to prevent costly actions based on stale or incorrect metadata.
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents
StealthBench is a new benchmark for measuring operational stealth and OPSEC tradecraft in autonomous offensive-security agents.
Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach
Cognitive Diagnostic Profiling (CDP) is proposed to align LLM-simulated examinees with real human psychometric data for better test calibration.
Automorphism-Induced Non-Canonicity in Top-k Explanations of Graph Neural Networks
Researchers identify a structural obstruction in Graph Neural Network explainers where automorphisms lead to arbitrary top-k edge reports.
When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses
A cross-domain benchmark reveals that LLM-simulated synthetic users fail to accurately replicate human survey responses, especially across cultures.
Pramana: A Composable, Domain-Specific Backend for Empirical Networking Research
Pramana is introduced as a composable backend designed to accelerate empirical networking research by bridging the gap between hypothesis and data generation.
High-Order Markov Blanket Discovery via a k-Order Relaxation of the Faithfulness Assumption
The k-order Markov blanket (kOMB) algorithm is proposed to discover graphical Markov blankets while relaxing the faithfulness assumption to handle parity-type relations.
"the very foundation of modern academia has been blown to bits"
A Hacker News discussion regarding the perceived collapse of the foundations of modern academia due to external pressures or AI.
Try Again, Don't Look Back: Blind Resampling Outperforms Self-Repair in Small Code Models
Research showing that 'blind resampling' (retrying without looking at previous failures) is more efficient and often more effective than self-repair in small code LLMs.
Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels
A proposed systems framework and graded trustworthiness levels for embodied AI to ensure safe and reliable physical interaction.
A Picture Says Thousands of Words - Harnessing Dermal Exposure Data from Images through Hybrid Deep Learning for Enhanced Safety Assessment
A hybrid deep learning approach using Mask R-CNN and color-based segmentation to quantify skin exposure for safety assessments.
Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition
An analysis of the structural and cognitive convergence between Large Language Models and human cognitive organization.
(EC)2: Event-Centric Explainability for Cybersecurity Through Multi-Agent LLM Investigations
Introduction of (EC)2, a multi-agent LLM framework for providing event-centric, verifiable explanations for cybersecurity alerts.
Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges
A systematic survey and taxonomy of Multi-Agent Debate (MAD) strategies to improve LLM robustness and accuracy.
Model-Driven Requirements Configuration with Three-Valued Uncertainty Scoring
A neuro-symbolic multi-agent architecture that combines LLMs with a deterministic symbolic validator to ensure structural integrity in requirements engineering.