AI/ML arXiv cs.AI

Good Benchmarks

A brief conceptual piece on the characteristics of 'good' benchmarks: correct, solvable, verifiable, and hard for interesting reasons.

AI/ML arXiv cs.AI

Rethinking the Evaluation of Harness Evolution for Agents

Challenges the effectiveness of automatic harness evolution for LLM agents, suggesting it may overfit and not outperform simple test-time scaling.

AI/ML arXiv cs.AI

On-Device Deep Research at 4B: Exposure Bounds Faithfulness, Retrieval Bounds Coverage

Analyzes the faithfulness and coverage of a 4B on-device research agent, finding that source exposure is the primary lever for faithfulness.

Tech Business/VC The Verge

Google and Epic give up fighting — third-party Android app stores are coming next week

Google and Epic Games have withdrawn a settlement attempt, leading Google to begin hosting third-party Android app stores within its own Play store starting July 22nd.

AI/ML arXiv cs.AI

Optimal Adaptive Market Making: A Theoretical Framework for High-Yield Liquidity Provision in Perpetual Futures Markets

A new theoretical framework for optimal market making in perpetual futures markets, focusing on high-yield liquidity provision and stochastic optimal control.

AI/ML arXiv cs.AI

In-Context Reinforcement Learning under Non-Stationarity: A Survey

A survey on in-context reinforcement learning (ICRL) specifically focusing on the challenges and mechanisms of non-stationary environments.

AI/ML arXiv cs.AI

Ontology-Amplified Distillation and Contextuality Auditing for Sovereign Enterprise Language Models: A Combined Proof-of-Mechanism and Negative-Results Method Study

A study on ontology-amplified distillation for enterprise language models and a method for auditing contextuality in agent routing, reporting mixed results.

AI/ML arXiv cs.AI

GRID: Grammar-Railed Decoding for Enterprise SQL Generation

Introduction of GRID, a grammar-railed decoding engine for enterprise SQL generation that ensures syntactic validity and role-based access control via Rust kernels.

AI/ML arXiv cs.AI

Calibration-First Reward-Component Auditing for Reinforcement Learning Control in Smart Greenhouses

A reward-component auditing framework for RL control in smart greenhouses, utilizing the GreenLight-Gym simulator.

AI/ML arXiv cs.AI

Optimization Is Not All You Need

A critical philosophical analysis of 'optimization culture' in AI, arguing that measurable improvement does not equate to value or judgment.

AI/ML arXiv cs.AI

LP Mining with LP2Graph: A Use Case for Railway Rescheduling

Presentation of LP2Graph, a method to mine the structure of Linear Programming formulations into reproducible datasets for railway rescheduling.

AI/ML arXiv cs.AI

Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability

A framework for 'agent-ready' website design to improve the reliability and efficiency of AI browser agents in e-commerce.

AI/ML arXiv cs.AI

Graph Feedback Controls Consensus and Clique Formation in Open-Weight Language-Model Populations

Research on how interaction graph feedback controls consensus and clique formation in populations of open-weight language models.

Other Hacker News

Jurassic Park computers in excruciating detail

An exploration of the computer systems depicted in the movie Jurassic Park in extreme detail.

Other Hacker News

The Trade in Looted Antiquities Endures for One Reason: Demand

An article discussing the persistent demand that drives the trade of looted antiquities.

AI/ML arXiv cs.AI

A Multimodal Dataset for Large Language Model Applications in the Energy Domain

Introduction of mAIEnergy, an open-access multimodal dataset designed for LLM applications in the energy sector.

AI/ML arXiv cs.AI

LightMem-Ego: Your AI Memory for Everyday Life

LightMem-Ego is a lightweight streaming multimodal memory system for personal AI assistants on mobile and wearable devices.

AI/ML arXiv cs.AI

Agentic Skill Optimization over Lie Algebroids

LASKO introduces a framework for agentic skill optimization using Lie Algebroids to speed up the refinement of prompts and tool contracts.

AI/ML arXiv cs.AI

IG-GAN: A Generative Adversarial Network for Aerodynamic Data Generation Based on Intrinsic Geometry

IG-GAN proposes a generative adversarial network based on intrinsic geometry for more accurate aerodynamic data generation.

AI/ML arXiv cs.AI

See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models

Robot-centric pointmaps are introduced to solve frame mismatch between visual observations and robot actions in VLA models.