All Articles
16570 articles total
Beyond Output-Space Calibration: Spectral Evidence Bundling for Selective Reliability Estimation in Time-Series Classification
Researchers introduce a spectral evidence bundling method to improve the reliability estimation of time-series classifiers beyond standard output-space calibration.
Beyond Single-Dimensional Compression: The Compound Sparsity Frontier of Large Language Models
A new compound sparsity framework combines low-rank approximation, channel pruning, and dynamic token skipping to delay performance decay in compressed LLMs.
Atomically Thin Materials Significantly Shrink Qubits
Research explores the use of atomically thin materials to significantly reduce the size of qubits, potentially advancing quantum computing scalability.
Tesla Balance Bike
A discussion or showcase regarding a Tesla-branded balance bike for children.
Sequential Learner Modeling Using Multi-Relational Graph Convolutional Networks
Introduces MR-ConceptGCN, an unsupervised approach for sequential learner modeling using multi-relational graph convolutional networks and SBERT.
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance
Presents BioSecBench-Surveillance, a verifiable benchmark for testing AI agents' ability to handle pathogen genomic surveillance pipelines.
Graph-Based Agentic AI with LangGraph: Workflow Pathways for Long-Running Stateful Business Processes
A practitioner's guide to using LangGraph for building long-running, stateful AI agent workflows in business processes.
LLM Detection as an Intervention: Downstream Impact under Strategic User Behavior
Study examines how LLM detection tools can paradoxically lead users to increase LLM usage and potentially decrease output quality.
ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
Introduces ResearchArena, a framework for evaluating sabotage and monitoring in automated AI R&D to ensure safe deployment of untrusted agents.
Associative Emotional Learning in Convolutional Neural Networks
Proposed a deep neural network model of visual valence processing to simulate human associative emotional learning.
Agents in the Wild: Where Research Meets Deployment
A tutorial on transitioning agentic LLM systems from research prototypes to reliable production-scale deployments.
CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents
Introduces CodeRescue, a budget-calibrated recovery routing system for coding agents to optimize the cost-efficiency of error recovery.
I graded 36 popular MCP servers on agent usability. A third got a D or F
An evaluation of 36 popular Model Context Protocol (MCP) servers, finding that a significant portion suffer from poor agent usability.
Mi-Memory: A Lifecycle Memory Framework for Personal AI
Introduction of Mi-Memory, a lifecycle memory framework for Personal AI designed to maintain durable user state and auditable evidence across devices.
Fishing Out Free Riders: Shapley-Based Reward Attribution for Parallel Reasoning via Reinforcement Learning
Parallel Shapley is a new RL framework that uses Shapley values to attribute rewards to specific reasoning paths in multi-path LLM reasoning.
Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interactio
Technical report on Athena-Brain-8B, a compact LLM optimized for on-device embodied intelligence and robot interaction.
Vector-Bench: Can Models Surgically Edit SVG Code?
Introduction of Vector-Bench, a benchmark for testing the ability of LLMs to surgically edit SVG code without introducing unintended changes.
Quality Action Assurance: Multimodal Verification of Examiner Claims in VR OSCEs
Quality Action Assurance (QAA) is a multimodal framework that verifies examiner claims in VR clinical examinations against actual event logs.
On the Effectiveness of Pretraining for Graph Combinatorial Optimization
Research on using self-supervised pretraining with geometric augmentations to improve neural solvers for graph combinatorial optimization problems like TSP.
Supra Cognitive Modes: A Routed Architecture for Agent Memory
Supra Cognitive Modes (SCM) is a routed architecture for agent memory that maps queries to specific retrieval and synthesis modes.