All Articles
17869 articles total
The Spec Growth Engine: Spec-Anchored, Code-Coupled, Drift-Enforced Architecture for AI-Assisted Software Development
The Spec Growth Engine is a framework for AI-assisted software development that uses a machine-readable spec graph to prevent context explosion and spec-code drift.
NuclearQAv2: A Structured Benchmark for Evaluating Domain-Science Competence in Large Language Models
NuclearQAv2 is a new structured benchmark designed to evaluate the quantitative reasoning and domain-science competence of LLMs in nuclear engineering.
What Is a Nomogram and Why Would It Interest Me?
A discussion on Hacker News exploring the concept of nomograms and their utility in calculation and data visualization.
Scaling Multi-Reference Image Generation with Dynamic Reward Optimization
Introduces OmniRef-Bench for evaluating multi-reference image generation and DyRef, a training framework to improve model performance in complex scenarios.
XMSE-Aware Adaptive Empirical Bayes Estimation
Proposes an XMSE-aware mixed estimator that interpolates between maximum likelihood and Empirical Bayes shrinkage to improve estimation under kernel misspecification.
In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis from Language Models to Physics
Presents In-Context Model Predictive Generation (ICMPG), a framework integrating LLM planning with physics simulation for realistic human motion synthesis.
Auditing Framing-Sensitive Behavioral Instability in Large Language Models for Mental Health Interactions
Analyzes how different contextual framings impact the behavioral stability and internal representations of LLMs in mental health interaction scenarios.
ReaORE: Reasoning-Guided Progressive Open Relation Extraction Empowered by Large Reasoning Models
Introduces ReaORE, a reasoning-guided progressive framework that uses coarse-to-fine reasoning to improve Open Relation Extraction.
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
Investigates emotion vectors in open-weight LLMs like Apertus and Gemma, discovering how valence representations emerge differently across model depth.
Decision-Aligned Evaluation of Uncertainty Quantification
Proposes a decision-aligned evaluation framework and prior-weighted utility metrics to better assess uncertainty quantification in machine learning.
Event-Aware Instructed Assistant for Referring Video Segmentation
Introduces EVIS, an event-aware video instructed segmentation assistant that decomposes videos into simple events for better target tracking.
Inverse Design of Compact and Wideband Inverted Doherty Power Amplifiers Using Deep Learning
Utilizes CNNs and genetic algorithms for the inverse design of compact, wideband inverted Doherty power amplifiers using GaN HEMT technology.
Modern GPU Programming for MLSys
A community discussion on modern GPU programming techniques specifically tailored for Machine Learning Systems (MLSys).
U.S. government will decide who gets to use latest upgrade to ChatGPT
The U.S. government is implementing a process to decide who receives access to the latest ChatGPT updates.
OpenAI limits GPT-5.6 rollout after government request, says restrictions shouldn’t be the norm
OpenAI restricts the rollout of GPT-5.6 following government requests, arguing that such restrictions hinder users and developers.
OpenAI poaches Uber India chief to lead its biggest market outside the U.S.
OpenAI hires the former Uber India chief to expand its presence and lead its operations in the Indian market.
Antibiotic "megacluster" discovery provides new strategy to fight superbugs
Researchers have discovered an antibiotic 'megacluster' providing a new strategy to fight antibiotic-resistant superbugs.
Confidence-Aware Tool Orchestration for Robust Video Understanding
Introduction of Robust-TO, an agentic video understanding framework that integrates per-frame trustworthiness to improve reliability in corrupted video inputs.
GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning
GEOALIGN is proposed as a lightweight plug-in for LLM reinforcement learning to reduce training instability caused by noisy rewards.
Risk-Aware Selective Multimodal Driver Monitoring with Driver-State World Modeling
A new cost-aware selective inference framework for multimodal driver monitoring that balances low-latency inference with safety interventions.