AI/ML arXiv cs.AI

Beyond Output-Space Calibration: Spectral Evidence Bundling for Selective Reliability Estimation in Time-Series Classification

Researchers introduce a spectral evidence bundling method to improve the reliability estimation of time-series classifiers beyond standard output-space calibration.

AI/ML arXiv cs.AI

Beyond Single-Dimensional Compression: The Compound Sparsity Frontier of Large Language Models

A new compound sparsity framework combines low-rank approximation, channel pruning, and dynamic token skipping to delay performance decay in compressed LLMs.

Hardware/Chips Hacker News

Atomically Thin Materials Significantly Shrink Qubits

Research explores the use of atomically thin materials to significantly reduce the size of qubits, potentially advancing quantum computing scalability.

Other Hacker News

Tesla Balance Bike

A discussion or showcase regarding a Tesla-branded balance bike for children.

AI/ML arXiv cs.AI

Sequential Learner Modeling Using Multi-Relational Graph Convolutional Networks

Introduces MR-ConceptGCN, an unsupervised approach for sequential learner modeling using multi-relational graph convolutional networks and SBERT.

AI/ML arXiv cs.AI

BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance

Presents BioSecBench-Surveillance, a verifiable benchmark for testing AI agents' ability to handle pathogen genomic surveillance pipelines.

AI/ML arXiv cs.AI

Graph-Based Agentic AI with LangGraph: Workflow Pathways for Long-Running Stateful Business Processes

A practitioner's guide to using LangGraph for building long-running, stateful AI agent workflows in business processes.

AI/ML arXiv cs.AI

LLM Detection as an Intervention: Downstream Impact under Strategic User Behavior

Study examines how LLM detection tools can paradoxically lead users to increase LLM usage and potentially decrease output quality.

AI/ML arXiv cs.AI

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

Introduces ResearchArena, a framework for evaluating sabotage and monitoring in automated AI R&D to ensure safe deployment of untrusted agents.

AI/ML arXiv cs.AI

Associative Emotional Learning in Convolutional Neural Networks

Proposed a deep neural network model of visual valence processing to simulate human associative emotional learning.

AI/ML arXiv cs.AI

Agents in the Wild: Where Research Meets Deployment

A tutorial on transitioning agentic LLM systems from research prototypes to reliable production-scale deployments.

AI/ML arXiv cs.AI

CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents

Introduces CodeRescue, a budget-calibrated recovery routing system for coding agents to optimize the cost-efficiency of error recovery.

Software Engineering Hacker News

I graded 36 popular MCP servers on agent usability. A third got a D or F

An evaluation of 36 popular Model Context Protocol (MCP) servers, finding that a significant portion suffer from poor agent usability.

AI/ML arXiv cs.AI

Mi-Memory: A Lifecycle Memory Framework for Personal AI

Introduction of Mi-Memory, a lifecycle memory framework for Personal AI designed to maintain durable user state and auditable evidence across devices.

AI/ML arXiv cs.AI

Fishing Out Free Riders: Shapley-Based Reward Attribution for Parallel Reasoning via Reinforcement Learning

Parallel Shapley is a new RL framework that uses Shapley values to attribute rewards to specific reasoning paths in multi-path LLM reasoning.

AI/ML arXiv cs.AI

Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interactio

Technical report on Athena-Brain-8B, a compact LLM optimized for on-device embodied intelligence and robot interaction.

AI/ML arXiv cs.AI

Vector-Bench: Can Models Surgically Edit SVG Code?

Introduction of Vector-Bench, a benchmark for testing the ability of LLMs to surgically edit SVG code without introducing unintended changes.

AI/ML arXiv cs.AI

Quality Action Assurance: Multimodal Verification of Examiner Claims in VR OSCEs

Quality Action Assurance (QAA) is a multimodal framework that verifies examiner claims in VR clinical examinations against actual event logs.

AI/ML arXiv cs.AI

On the Effectiveness of Pretraining for Graph Combinatorial Optimization

Research on using self-supervised pretraining with geometric augmentations to improve neural solvers for graph combinatorial optimization problems like TSP.

AI/ML arXiv cs.AI

Supra Cognitive Modes: A Routed Architecture for Agent Memory

Supra Cognitive Modes (SCM) is a routed architecture for agent memory that maps queries to specific retrieval and synthesis modes.