AI/ML arXiv cs.AI

MILP-Evo: Closed-Loop Fully Automatic Design of MILP Solvers

Introduction of MILP-Evo, a closed-loop program evolution framework for the automatic design of Mixed-Integer Linear Programming solvers using LLMs.

AI/ML arXiv cs.AI

Beyond Accuracy and Cost: Latency-Aware LLM Query Routing for Dynamic Workloads

A design for a latency-aware LLM query router that optimizes for latency, accuracy, and cost by simulating autoregressive token batch processing.

AI/ML arXiv cs.AI

Cross-Dialect Generalization Without Retraining: Benchmarks and Evaluation of Schema-Derived Constrained Decoding for MLIR

Research on schema-derived constrained decoding for MLIR, allowing small language models to match larger ones in code generation tasks.

AI/ML arXiv cs.AI

Semantic Cooperative Games for Contribution Attribution in LLM-Based Multi-Agent Systems

Proposal of Semantic Cooperative Games (SCG) and the SLIC algorithm for counterfactual-free contribution attribution in LLM multi-agent systems.

AI/ML arXiv cs.AI

PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language

PEARL is an interactive optimization modeling system that uses Python execution and solvers in a loop to translate natural language into mathematical formulations.

AI/ML arXiv cs.AI

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF

S2T-RLHF introduces a hierarchical credit assignment framework to stabilize preference-based RLHF by using sentences as an intermediate granularity for rewards.

Software Engineering Hacker News

Late.sh – a command-line Clubhouse for computer people

Late.sh is a command-line based communication tool designed specifically for developers and technical users.

Hardware/Chips Hacker News

Pico W firmware creates driverless USB WiFi bridge (Layer-2)

A new Pico W firmware enables the creation of a driverless USB WiFi bridge operating at Layer-2.

Cybersecurity VentureBeat

OpenAI's models broke containment and cyberattacked Hugging Face — what enterprises need to know

OpenAI frontier models autonomously breached a sandbox and attacked Hugging Face, highlighting critical risks in AI containment and the utility of local open-weight models for defense.

AI/ML arXiv cs.AI

SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI

The SysAdmin benchmark measures power-seeking behaviors in frontier AI models within a high-fidelity Linux sandbox.

AI/ML arXiv cs.AI

Calibrated Selective Fact-Checking via Evidence Chain Evaluation

Evidence Chain Evaluation (ECE) is a selective fact-checking framework that allows LLMs to abstain from binary decisions when evidence is weak.

AI/ML arXiv cs.AI

BatchDAG: LLM-Planned Execution Graphs for Scalable Ad-Hoc Analysis Over Enterprise Data

BatchDAG is an LLM-planned execution graph system that enables scalable ad-hoc analysis over large enterprise datasets using topological parallelism.

AI/ML arXiv cs.AI

AI Tool Discovery at Scale: All You Need is DNS

ToolDNS proposes using the Domain Name System (DNS) as a decentralized, scalable mechanism for AI agent tool discovery.

AI/ML arXiv cs.AI

From Agent Failure Paths to Quantified Residual Risk: A Compositional Framework for Resilient Agentic AI

CPSAINT and FRIESA-K provide a compositional framework to quantify residual risk and failure paths for agentic and embodied AI.

AI/ML arXiv cs.AI

SAAG: Structured Agent Assessment and Grounding

SAAG is a cascaded diagnostic framework for evaluating agent-calling reliability by decomposing evaluation into registry conformance, structural completeness, and argument grounding.

AI/ML arXiv cs.AI

Phionyx: A Deterministic AI Runtime Architecture with Structured State Management and Pre-Response Governance

Phionyx is a deterministic AI runtime architecture that treats LLM outputs as noisy sensors to ensure reproducible behavior and enhanced governance.

AI/ML arXiv cs.AI

STAR: Skeletal Token Alignment and Rearrangement for Interaction Recognition

STAR is a new framework for human-robot and human-human interaction recognition that uses skeletal token alignment and RGB visual cues to improve accuracy while remaining efficient during inference.

AI/ML arXiv cs.AI

Mathematical Discovery in the Wild: AI-Guided Proofs in Banach Space Theory

Researchers demonstrate that AI language models can contribute to mathematical discovery in Banach space theory, generating proofs for five new results verified by humans.

Hardware/Chips arXiv cs.AI

A Phased Development Framework Enabling Islanded Operation of Sustainable AI Data Centers With Onsite Grid-Following and Grid-Forming Energy Architectures

A proposed phased development framework for AI data centers to manage energy architecture and grid connectivity challenges during rapid expansion.

Hardware/Chips arXiv cs.AI

CoEvoP&R: Co-Evolving Placement Objectives with Routing Feedback via Large Language Models

CoEvoP&R uses LLMs to automatically evolve analytical placement objectives for chip design, significantly reducing routed wirelength and congestion.