AI/ML arXiv cs.AI

Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models

Introduces function-aware fill-in-the-middle (FIM) mid-training to improve the ability of coding agents to integrate tool returns, showing gains on SWE-bench.

AI/ML arXiv cs.AI

From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery

Proposes Mechanistic World Models as a new paradigm to move AI from predictive forecasting to autonomous scientific discovery by prioritizing explanatory mechanisms.

AI/ML arXiv cs.AI

TRACE: An Operational Reasoning Schema for Auditable Agentic Commitments

Defines TRACE, an operational reasoning schema for recording auditable agentic commitments to separate associative computation from formal reasoning.

Other Hacker News

Andon (manufacturing)

A discussion about Andon, a manufacturing system for signaling problems on the production line.

AI/ML arXiv cs.AI

Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking

Presents GenAI Evaluation, a governed pipeline for large-scale evaluation of retail conversational agents focusing on intent alignment and factuality.

AI/ML arXiv cs.AI

Representing and Generating Levels Over Time through Playtrace Reconstructive Partitioning

Introduces a 'cake' representation for game levels and a generation approach called Playtrace Reconstructive Partitioning (PRP) for dynamic level generation.

AI/ML arXiv cs.AI

Connected by Construction: Learning Tractable Near-Tour Marginals for Traveling Salesman Problems

Proposes C2TSP, an unsupervised learning pipeline for the Traveling Salesman Problem that learns structural Hamiltonian information directly.

AI/ML arXiv cs.AI

The Emerging Paradigm of Geospatial Foundation Models: From Pre-Training to Agentic Reasoning

Explores Geospatial Foundation Models (GeoFMs) and their transition from pre-training to agentic reasoning for satellite and aerial imagery analysis.

AI/ML arXiv cs.AI

Cost-Governed RAG: Unified Per-Tenant Cost Attribution Across Retrieval and Generation in Multi-Tenant LLM Systems

Describes Cost-Governed RAG, an architecture using TurboVec for precise per-tenant cost attribution in multi-tenant LLM systems.

Cybersecurity arXiv cs.AI

A Threshold Exceedance Framework for CBRN Uplift Evaluation in Frontier Language Models

Introduces the Threshold Exceedance Criteria (TEC) framework to evaluate if frontier LLMs increase the ability of non-experts to plan CBRN misuse.

AI/ML arXiv cs.AI

Good Benchmarks

A brief conceptual piece on the characteristics of 'good' benchmarks: correct, solvable, verifiable, and hard for interesting reasons.

AI/ML arXiv cs.AI

Rethinking the Evaluation of Harness Evolution for Agents

Challenges the effectiveness of automatic harness evolution for LLM agents, suggesting it may overfit and not outperform simple test-time scaling.

AI/ML arXiv cs.AI

On-Device Deep Research at 4B: Exposure Bounds Faithfulness, Retrieval Bounds Coverage

Analyzes the faithfulness and coverage of a 4B on-device research agent, finding that source exposure is the primary lever for faithfulness.

Tech Business/VC The Verge

Google and Epic give up fighting — third-party Android app stores are coming next week

Google and Epic Games have withdrawn a settlement attempt, leading Google to begin hosting third-party Android app stores within its own Play store starting July 22nd.

AI/ML arXiv cs.AI

Optimal Adaptive Market Making: A Theoretical Framework for High-Yield Liquidity Provision in Perpetual Futures Markets

A new theoretical framework for optimal market making in perpetual futures markets, focusing on high-yield liquidity provision and stochastic optimal control.

AI/ML arXiv cs.AI

In-Context Reinforcement Learning under Non-Stationarity: A Survey

A survey on in-context reinforcement learning (ICRL) specifically focusing on the challenges and mechanisms of non-stationary environments.

AI/ML arXiv cs.AI

Ontology-Amplified Distillation and Contextuality Auditing for Sovereign Enterprise Language Models: A Combined Proof-of-Mechanism and Negative-Results Method Study

A study on ontology-amplified distillation for enterprise language models and a method for auditing contextuality in agent routing, reporting mixed results.

AI/ML arXiv cs.AI

GRID: Grammar-Railed Decoding for Enterprise SQL Generation

Introduction of GRID, a grammar-railed decoding engine for enterprise SQL generation that ensures syntactic validity and role-based access control via Rust kernels.

AI/ML arXiv cs.AI

Calibration-First Reward-Component Auditing for Reinforcement Learning Control in Smart Greenhouses

A reward-component auditing framework for RL control in smart greenhouses, utilizing the GreenLight-Gym simulator.

AI/ML arXiv cs.AI

Optimization Is Not All You Need

A critical philosophical analysis of 'optimization culture' in AI, arguing that measurable improvement does not equate to value or judgment.