All Articles
17040 articles total
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models
Introduces function-aware fill-in-the-middle (FIM) mid-training to improve the ability of coding agents to integrate tool returns, showing gains on SWE-bench.
From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery
Proposes Mechanistic World Models as a new paradigm to move AI from predictive forecasting to autonomous scientific discovery by prioritizing explanatory mechanisms.
TRACE: An Operational Reasoning Schema for Auditable Agentic Commitments
Defines TRACE, an operational reasoning schema for recording auditable agentic commitments to separate associative computation from formal reasoning.
Andon (manufacturing)
A discussion about Andon, a manufacturing system for signaling problems on the production line.
Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking
Presents GenAI Evaluation, a governed pipeline for large-scale evaluation of retail conversational agents focusing on intent alignment and factuality.
Representing and Generating Levels Over Time through Playtrace Reconstructive Partitioning
Introduces a 'cake' representation for game levels and a generation approach called Playtrace Reconstructive Partitioning (PRP) for dynamic level generation.
Connected by Construction: Learning Tractable Near-Tour Marginals for Traveling Salesman Problems
Proposes C2TSP, an unsupervised learning pipeline for the Traveling Salesman Problem that learns structural Hamiltonian information directly.
The Emerging Paradigm of Geospatial Foundation Models: From Pre-Training to Agentic Reasoning
Explores Geospatial Foundation Models (GeoFMs) and their transition from pre-training to agentic reasoning for satellite and aerial imagery analysis.
Cost-Governed RAG: Unified Per-Tenant Cost Attribution Across Retrieval and Generation in Multi-Tenant LLM Systems
Describes Cost-Governed RAG, an architecture using TurboVec for precise per-tenant cost attribution in multi-tenant LLM systems.
A Threshold Exceedance Framework for CBRN Uplift Evaluation in Frontier Language Models
Introduces the Threshold Exceedance Criteria (TEC) framework to evaluate if frontier LLMs increase the ability of non-experts to plan CBRN misuse.
Good Benchmarks
A brief conceptual piece on the characteristics of 'good' benchmarks: correct, solvable, verifiable, and hard for interesting reasons.
Rethinking the Evaluation of Harness Evolution for Agents
Challenges the effectiveness of automatic harness evolution for LLM agents, suggesting it may overfit and not outperform simple test-time scaling.
On-Device Deep Research at 4B: Exposure Bounds Faithfulness, Retrieval Bounds Coverage
Analyzes the faithfulness and coverage of a 4B on-device research agent, finding that source exposure is the primary lever for faithfulness.
Google and Epic give up fighting — third-party Android app stores are coming next week
Google and Epic Games have withdrawn a settlement attempt, leading Google to begin hosting third-party Android app stores within its own Play store starting July 22nd.
Optimal Adaptive Market Making: A Theoretical Framework for High-Yield Liquidity Provision in Perpetual Futures Markets
A new theoretical framework for optimal market making in perpetual futures markets, focusing on high-yield liquidity provision and stochastic optimal control.
In-Context Reinforcement Learning under Non-Stationarity: A Survey
A survey on in-context reinforcement learning (ICRL) specifically focusing on the challenges and mechanisms of non-stationary environments.
Ontology-Amplified Distillation and Contextuality Auditing for Sovereign Enterprise Language Models: A Combined Proof-of-Mechanism and Negative-Results Method Study
A study on ontology-amplified distillation for enterprise language models and a method for auditing contextuality in agent routing, reporting mixed results.
GRID: Grammar-Railed Decoding for Enterprise SQL Generation
Introduction of GRID, a grammar-railed decoding engine for enterprise SQL generation that ensures syntactic validity and role-based access control via Rust kernels.
Calibration-First Reward-Component Auditing for Reinforcement Learning Control in Smart Greenhouses
A reward-component auditing framework for RL control in smart greenhouses, utilizing the GreenLight-Gym simulator.
Optimization Is Not All You Need
A critical philosophical analysis of 'optimization culture' in AI, arguing that measurable improvement does not equate to value or judgment.