All Articles
17547 articles total
Demonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at Scale
Demonstration of TOFFEE, a system that uses MCTS and adaptive model selection to synthesize high-quality data agent trajectories for finetuning and in-context learning.
From Application-Layer Simulation to Native Meta-Architecture: Structural Tension as an Endogenous Driver for Heterogeneous AI Evolution
A theoretical framework proposing a native meta-architecture for AI to move beyond stateless LLMs by introducing structural tension, recurrent loops, and inference-time plasticity.
GitLost: We Tricked GitHub's AI Agent into Leaking Private Repos
Researchers successfully tricked GitHub's AI agent into leaking private repository information, highlighting critical security vulnerabilities in AI-integrated developer tools.
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
TurnOPD introduces a turn-level budgeting strategy for on-policy distillation to improve the efficiency and accuracy of training long-horizon AI agents.
Onnes: A Physics-Grounded Multi-Agent LLM Simulator for Cryogenic Fault Diagnosis in Quantum Computing Infrastructure
Onnes is a physics-grounded digital-twin simulator for cryogenic fault diagnosis in quantum computing, using multi-agent LLMs to match supervised ML classifier performance.
StateFuse: Deterministic Conflict-Preserving Memory for Multi-Agent Systems
StateFuse is a conflict-aware replicated memory contract using CRDTs to preserve contradictions in multi-agent systems for better auditability and correction.
Uncovering Latent Depression Severity for Binary Depression Detection via Advantage-weighting Ranking
A new multimodal framework using Binary Advantage-weighting Ranking Loss improves binary depression detection by better disentangling feature distributions in audio-visual data.
PCBWorld: A Benchmark Environment for Engine-Grounded PCB Design Automation
PCBWorld is an open-source, engine-grounded environment built on KiCad for automating PCB routing using RL and LLM agents.
SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation
SearchEyes utilizes a simulated search world based on typed knowledge graphs to improve multi-hop reasoning and performance in multimodal search agents.
Integrating knowledge graphs and multilingual scholarly corpora for domain-adaptive LLMs in SSH
Research into adapting foundation models for the Social Sciences and Humanities (SSH) through knowledge graphs and multilingual corpora to ensure epistemic responsibility.
Auto-DSM Under the Lens: A Black-Box Evaluation Framework for LLM-Based DSM Generation
A black-box evaluation framework for assessing the ability of LLMs to generate Design Structure Matrices (DSMs) from technical documentation.
AgoraSim: A Hybrid Agent-Based Modeling Framework
AgoraSim is a hybrid agent-based modeling framework that combines LLM agents with classical ABM to analyze social reaction scenarios.
How to Build a Minimal ZFS NAS Without Synology, QNAP, TrueNAS (2024)
A guide on building a minimal ZFS-based Network Attached Storage (NAS) without relying on commercial hardware or heavy OS distributions.
Copy That Floppy – Cambridge guide for preserving data from fragile floppy disks
A technical guide from Cambridge University for the preservation and data recovery of fragile floppy disks.
GPT-5.6 Sol, along with Terra and Luna, will launch publicly this Thursday
Announcement of the public launch of GPT-5.6 Sol, Terra, and Luna models.
The difference between "today's task" and "accretive work"
A discussion on the productivity difference between completing daily tasks and performing work that builds compounding value over time.
Out of the Armchair
An article titled 'Out of the Armchair', likely a commentary or opinion piece on transitioning from theory to practice.
Synthetic Consumer Insight Generation with Large Language Models
Research exploring the use of LLMs to generate synthetic consumer data for projective marketing techniques and comparing it to human responses.
Beyond Static Evaluation: Building Simulation Environments for Scalable Agentic Reinforcement Learning
Introduction of AgenticAI-Supervisor, an RL Gym environment designed for scalable execution and evaluation of LLM-based autonomous agents.
Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents
A comprehensive synthesis and taxonomy of common failure modes in LLM agents across tool-use, planning, and reasoning.