All Articles
17740 articles total
Popping the GPU Bubble
A community discussion on Hacker News regarding the potential economic bubble surrounding GPU demand and infrastructure.
Study suggests most Americans would be healthier without daylight saving time
A study suggesting that Americans would be healthier if daylight saving time were abolished.
Meituan open sources LongCat-2.0, the 1.6T, near-frontier agentic coding model that's been leading OpenRouter — trained entirely on Chinese chips
Meituan open sources LongCat-2.0, a 1.6T parameter MoE coding model trained on Chinese ASICs, featuring a 1M token context window and MIT license.
Preventing Error Propagation in Multi-Agent AI through Runtime Monitoring
Research on preventing error propagation in multi-agent AI systems through runtime monitoring and reasoning exchange.
Memory as an Attack Surface in LLM Agents: A Study on Multiple-Choice Question Answering
A study exploring how memory in LLM agents can be used as an attack surface to manipulate model outputs.
Low-cost concept-based localized explanations: How far can we get with training-free approaches?
An evaluation of training-free, zero-shot approaches for localized concept naming in Explainable AI using MLLMs.
Managing the Human Fallback: Skill Investment Under Improving AI and Worker Mobility
An economic model analyzing the balance of human skill investment and AI deployment within firms.
Characterizing Large Language Model Agentic Workflows: A Study on N8n Ecosystem
An empirical study of over 6,000 n8n workflows to characterize the design and reliability of LLM agentic workflows in low-code platforms.
HiComm: Hierarchical Communication for Multi-agent Reinforcement Learning
Introduction of HiComm, a hierarchical communication module for Multi-agent Reinforcement Learning that reduces communication volume.
Flow Reasoning Models: Scaling Reasoning Through Iterative Self-Refinement
Introduction of Flow Reasoning Models (FRMs) that use iterative self-refinement and test-time scaling to solve structured reasoning tasks.
A Fake Shell for Pangenomics
A discussion or tool related to a fake shell environment designed for pangenomics research.
Agentic Abstention: Do Agents Know When to Stop Instead of Act?
Researchers introduce 'Agentic Abstention' and the CONVOLVE context engineering method to help LLM agents recognize when to stop acting to avoid unnecessary tool calls.
Agent Safety Is Action Alignment
The paper argues that agent safety should be enforced as 'least privilege' outside the model at the action boundary rather than relying on in-weight refusal training.
Self-Supervised Theorem Discovery in a Formal Axiomatic System
An agent is developed that can autonomously discover useful mathematical theorems from axioms alone via a self-supervised algorithm without human priors.
Mechanistic Personality Analysis of LLMs Steering Personality via Latent Feature Interventions
A mechanistic interpretability approach uses sparse autoencoders to identify and steer personality traits in LLMs via latent feature interventions.
HyphaeDB: A Living Knowledge Topology for Agent-First Memory
HyphaeDB is introduced as an agent-native memory infrastructure that uses HNSW graph topology and gossip protocols for multi-agent knowledge propagation.
Primary ICD Category Prediction using LLM-based Probing
A study demonstrates that frozen medical LLM representations can effectively unify structured and narrative EHR data for primary ICD category prediction.
MedEvoEval: Evaluating Continual Evolution of Doctor Agents through Simulated Clinical Episodes
MedEvoEval is a new longitudinal evaluation framework for assessing how 'doctor agents' evolve and improve through simulated clinical episodes.
Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries
Physician-led evaluation shows that specialized clinical AI tools outperform general-purpose frontier models on real-world point-of-care queries.
Customized Generative AI Agent for Transportation Engineering Practice: A Development and Continued Pre-training Guideline
A guideline is proposed for developing transportation engineering AI agents using continued pre-training with LoRA on domain-specific manuals.