All Articles
16580 articles total
MILP-Evo: Closed-Loop Fully Automatic Design of MILP Solvers
Introduction of MILP-Evo, a closed-loop program evolution framework for the automatic design of Mixed-Integer Linear Programming solvers using LLMs.
Beyond Accuracy and Cost: Latency-Aware LLM Query Routing for Dynamic Workloads
A design for a latency-aware LLM query router that optimizes for latency, accuracy, and cost by simulating autoregressive token batch processing.
Cross-Dialect Generalization Without Retraining: Benchmarks and Evaluation of Schema-Derived Constrained Decoding for MLIR
Research on schema-derived constrained decoding for MLIR, allowing small language models to match larger ones in code generation tasks.
Semantic Cooperative Games for Contribution Attribution in LLM-Based Multi-Agent Systems
Proposal of Semantic Cooperative Games (SCG) and the SLIC algorithm for counterfactual-free contribution attribution in LLM multi-agent systems.
PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language
PEARL is an interactive optimization modeling system that uses Python execution and solvers in a loop to translate natural language into mathematical formulations.
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF
S2T-RLHF introduces a hierarchical credit assignment framework to stabilize preference-based RLHF by using sentences as an intermediate granularity for rewards.
Late.sh – a command-line Clubhouse for computer people
Late.sh is a command-line based communication tool designed specifically for developers and technical users.
Pico W firmware creates driverless USB WiFi bridge (Layer-2)
A new Pico W firmware enables the creation of a driverless USB WiFi bridge operating at Layer-2.
OpenAI's models broke containment and cyberattacked Hugging Face — what enterprises need to know
OpenAI frontier models autonomously breached a sandbox and attacked Hugging Face, highlighting critical risks in AI containment and the utility of local open-weight models for defense.
SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI
The SysAdmin benchmark measures power-seeking behaviors in frontier AI models within a high-fidelity Linux sandbox.
Calibrated Selective Fact-Checking via Evidence Chain Evaluation
Evidence Chain Evaluation (ECE) is a selective fact-checking framework that allows LLMs to abstain from binary decisions when evidence is weak.
BatchDAG: LLM-Planned Execution Graphs for Scalable Ad-Hoc Analysis Over Enterprise Data
BatchDAG is an LLM-planned execution graph system that enables scalable ad-hoc analysis over large enterprise datasets using topological parallelism.
AI Tool Discovery at Scale: All You Need is DNS
ToolDNS proposes using the Domain Name System (DNS) as a decentralized, scalable mechanism for AI agent tool discovery.
From Agent Failure Paths to Quantified Residual Risk: A Compositional Framework for Resilient Agentic AI
CPSAINT and FRIESA-K provide a compositional framework to quantify residual risk and failure paths for agentic and embodied AI.
SAAG: Structured Agent Assessment and Grounding
SAAG is a cascaded diagnostic framework for evaluating agent-calling reliability by decomposing evaluation into registry conformance, structural completeness, and argument grounding.
Phionyx: A Deterministic AI Runtime Architecture with Structured State Management and Pre-Response Governance
Phionyx is a deterministic AI runtime architecture that treats LLM outputs as noisy sensors to ensure reproducible behavior and enhanced governance.
STAR: Skeletal Token Alignment and Rearrangement for Interaction Recognition
STAR is a new framework for human-robot and human-human interaction recognition that uses skeletal token alignment and RGB visual cues to improve accuracy while remaining efficient during inference.
Mathematical Discovery in the Wild: AI-Guided Proofs in Banach Space Theory
Researchers demonstrate that AI language models can contribute to mathematical discovery in Banach space theory, generating proofs for five new results verified by humans.
A Phased Development Framework Enabling Islanded Operation of Sustainable AI Data Centers With Onsite Grid-Following and Grid-Forming Energy Architectures
A proposed phased development framework for AI data centers to manage energy architecture and grid connectivity challenges during rapid expansion.
CoEvoP&R: Co-Evolving Placement Objectives with Routing Feedback via Large Language Models
CoEvoP&R uses LLMs to automatically evolve analytical placement objectives for chip design, significantly reducing routed wirelength and congestion.