Software Engineering arXiv cs.AI

The Spec Growth Engine: Spec-Anchored, Code-Coupled, Drift-Enforced Architecture for AI-Assisted Software Development

The Spec Growth Engine is a framework for AI-assisted software development that uses a machine-readable spec graph to prevent context explosion and spec-code drift.

AI/ML arXiv cs.AI

NuclearQAv2: A Structured Benchmark for Evaluating Domain-Science Competence in Large Language Models

NuclearQAv2 is a new structured benchmark designed to evaluate the quantitative reasoning and domain-science competence of LLMs in nuclear engineering.

Other Hacker News

What Is a Nomogram and Why Would It Interest Me?

A discussion on Hacker News exploring the concept of nomograms and their utility in calculation and data visualization.

AI/ML arXiv cs.AI

Scaling Multi-Reference Image Generation with Dynamic Reward Optimization

Introduces OmniRef-Bench for evaluating multi-reference image generation and DyRef, a training framework to improve model performance in complex scenarios.

AI/ML arXiv cs.AI

XMSE-Aware Adaptive Empirical Bayes Estimation

Proposes an XMSE-aware mixed estimator that interpolates between maximum likelihood and Empirical Bayes shrinkage to improve estimation under kernel misspecification.

AI/ML arXiv cs.AI

In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis from Language Models to Physics

Presents In-Context Model Predictive Generation (ICMPG), a framework integrating LLM planning with physics simulation for realistic human motion synthesis.

AI/ML arXiv cs.AI

Auditing Framing-Sensitive Behavioral Instability in Large Language Models for Mental Health Interactions

Analyzes how different contextual framings impact the behavioral stability and internal representations of LLMs in mental health interaction scenarios.

AI/ML arXiv cs.AI

ReaORE: Reasoning-Guided Progressive Open Relation Extraction Empowered by Large Reasoning Models

Introduces ReaORE, a reasoning-guided progressive framework that uses coarse-to-fine reasoning to improve Open Relation Extraction.

AI/ML arXiv cs.AI

Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs

Investigates emotion vectors in open-weight LLMs like Apertus and Gemma, discovering how valence representations emerge differently across model depth.

AI/ML arXiv cs.AI

Decision-Aligned Evaluation of Uncertainty Quantification

Proposes a decision-aligned evaluation framework and prior-weighted utility metrics to better assess uncertainty quantification in machine learning.

AI/ML arXiv cs.AI

Event-Aware Instructed Assistant for Referring Video Segmentation

Introduces EVIS, an event-aware video instructed segmentation assistant that decomposes videos into simple events for better target tracking.

Hardware/Chips arXiv cs.AI

Inverse Design of Compact and Wideband Inverted Doherty Power Amplifiers Using Deep Learning

Utilizes CNNs and genetic algorithms for the inverse design of compact, wideband inverted Doherty power amplifiers using GaN HEMT technology.

AI/ML Hacker News

Modern GPU Programming for MLSys

A community discussion on modern GPU programming techniques specifically tailored for Machine Learning Systems (MLSys).

Tech Business/VC Hacker News

U.S. government will decide who gets to use latest upgrade to ChatGPT

The U.S. government is implementing a process to decide who receives access to the latest ChatGPT updates.

Tech Business/VC TechCrunch

OpenAI limits GPT-5.6 rollout after government request, says restrictions shouldn’t be the norm

OpenAI restricts the rollout of GPT-5.6 following government requests, arguing that such restrictions hinder users and developers.

Tech Business/VC TechCrunch

OpenAI poaches Uber India chief to lead its biggest market outside the U.S.

OpenAI hires the former Uber India chief to expand its presence and lead its operations in the Indian market.

Other Ars Technica

Antibiotic "megacluster" discovery provides new strategy to fight superbugs

Researchers have discovered an antibiotic 'megacluster' providing a new strategy to fight antibiotic-resistant superbugs.

AI/ML arXiv cs.AI

Confidence-Aware Tool Orchestration for Robust Video Understanding

Introduction of Robust-TO, an agentic video understanding framework that integrates per-frame trustworthiness to improve reliability in corrupted video inputs.

AI/ML arXiv cs.AI

GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning

GEOALIGN is proposed as a lightweight plug-in for LLM reinforcement learning to reduce training instability caused by noisy rewards.

AI/ML arXiv cs.AI

Risk-Aware Selective Multimodal Driver Monitoring with Driver-State World Modeling

A new cost-aware selective inference framework for multimodal driver monitoring that balances low-latency inference with safety interventions.