All Articles
17050 articles total
Good Benchmarks
A brief conceptual piece on the characteristics of 'good' benchmarks: correct, solvable, verifiable, and hard for interesting reasons.
Rethinking the Evaluation of Harness Evolution for Agents
Challenges the effectiveness of automatic harness evolution for LLM agents, suggesting it may overfit and not outperform simple test-time scaling.
On-Device Deep Research at 4B: Exposure Bounds Faithfulness, Retrieval Bounds Coverage
Analyzes the faithfulness and coverage of a 4B on-device research agent, finding that source exposure is the primary lever for faithfulness.
Google and Epic give up fighting — third-party Android app stores are coming next week
Google and Epic Games have withdrawn a settlement attempt, leading Google to begin hosting third-party Android app stores within its own Play store starting July 22nd.
Optimal Adaptive Market Making: A Theoretical Framework for High-Yield Liquidity Provision in Perpetual Futures Markets
A new theoretical framework for optimal market making in perpetual futures markets, focusing on high-yield liquidity provision and stochastic optimal control.
In-Context Reinforcement Learning under Non-Stationarity: A Survey
A survey on in-context reinforcement learning (ICRL) specifically focusing on the challenges and mechanisms of non-stationary environments.
Ontology-Amplified Distillation and Contextuality Auditing for Sovereign Enterprise Language Models: A Combined Proof-of-Mechanism and Negative-Results Method Study
A study on ontology-amplified distillation for enterprise language models and a method for auditing contextuality in agent routing, reporting mixed results.
GRID: Grammar-Railed Decoding for Enterprise SQL Generation
Introduction of GRID, a grammar-railed decoding engine for enterprise SQL generation that ensures syntactic validity and role-based access control via Rust kernels.
Calibration-First Reward-Component Auditing for Reinforcement Learning Control in Smart Greenhouses
A reward-component auditing framework for RL control in smart greenhouses, utilizing the GreenLight-Gym simulator.
Optimization Is Not All You Need
A critical philosophical analysis of 'optimization culture' in AI, arguing that measurable improvement does not equate to value or judgment.
LP Mining with LP2Graph: A Use Case for Railway Rescheduling
Presentation of LP2Graph, a method to mine the structure of Linear Programming formulations into reproducible datasets for railway rescheduling.
Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability
A framework for 'agent-ready' website design to improve the reliability and efficiency of AI browser agents in e-commerce.
Graph Feedback Controls Consensus and Clique Formation in Open-Weight Language-Model Populations
Research on how interaction graph feedback controls consensus and clique formation in populations of open-weight language models.
Jurassic Park computers in excruciating detail
An exploration of the computer systems depicted in the movie Jurassic Park in extreme detail.
The Trade in Looted Antiquities Endures for One Reason: Demand
An article discussing the persistent demand that drives the trade of looted antiquities.
A Multimodal Dataset for Large Language Model Applications in the Energy Domain
Introduction of mAIEnergy, an open-access multimodal dataset designed for LLM applications in the energy sector.
LightMem-Ego: Your AI Memory for Everyday Life
LightMem-Ego is a lightweight streaming multimodal memory system for personal AI assistants on mobile and wearable devices.
Agentic Skill Optimization over Lie Algebroids
LASKO introduces a framework for agentic skill optimization using Lie Algebroids to speed up the refinement of prompts and tool contracts.
IG-GAN: A Generative Adversarial Network for Aerodynamic Data Generation Based on Intrinsic Geometry
IG-GAN proposes a generative adversarial network based on intrinsic geometry for more accurate aerodynamic data generation.
See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models
Robot-centric pointmaps are introduced to solve frame mismatch between visual observations and robot actions in VLA models.