All Articles
18316 articles total
MetaResearcher: Scaling Deep Research via Self-Reflective Reinforcement Learning in Adversarial Virtual Environments
MetaResearcher is a framework for scaling deep research agents using self-reflective reinforcement learning and adversarial virtual environments.
Multi-Agent Transactive Memory
Multi-Agent Transactive Memory (MATM) is a framework for storing and retrieving agent-generated trajectories to enable experience sharing across agent populations.
Barret Zoph is out at OpenAI again after just five months
Barret Zoph, the head of enterprise AI sales at OpenAI, has departed the company after a brief five-month tenure.
Exit-and-Join Dynamics for Decentralized Coalition Formation
A research paper proposes a decentralized dynamical process for coalition formation using the Aumann-Dreze value to evaluate local moves.
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents
This study argues that aggregate-score leaderboards for LLM agents are unstable and proposes ranking configurations based on predictive validity.
GLARE: A Natural Language Interface for Querying Global Explanations
GLARE is introduced as an LLM-based interactive interface that translates natural language questions into SQL queries to explore global explanations for image classifiers.
Interpreting Neural Combinatorial Optimization via Evolving Programmatic Bottlenecks
The Evolving Programmatic Bottlenecks (EPB) framework is introduced to interpret Neural Combinatorial Optimization policies by distilling them into human-readable program portfolios.
A Comparative Study of Pretrained Transformer Models for Quranic ASR: Speech Representations, Label Formats, and Dataset Composition
A comparative study examines the use of pretrained Transformer models like Wav2Vec2.0 and HuBERT for Quranic Automatic Speech Recognition (ASR).
Benchmarking Agentic Review Systems
This paper benchmarks agentic review systems for research papers, finding that systems like OpenAIReview can track human quality judgments reasonably well.
Grounded Inference: Principles for Deterministically Encapsulated Generative Models
The paper establishes principles for 'Grounded Inference' to deterministically encapsulate probabilistic generative models within traditional computational systems.
Optimal Scheduling in a Question-Answering Forum of Knowledge Workers
A study explores optimal scheduling in question-answering forums employing knowledge workers to maintain system stability and capacity.
Beyond Entropy: Learning from Token-Level Distributional Deviations for LLM Reasoning
The Independent Combinatorial Tokens (ICT) framework is proposed to stabilize LLM reasoning training by focusing updates on token-level distributional deviations.
Let's Encrypt has been down most of today
Let's Encrypt experienced significant downtime throughout the day, impacting SSL/TLS certificate issuance and renewal.
Ice Water Drowning Survival After 147-Minute Submersion and Hypothermic Arrest
A medical case report detailing a rare survival instance after 147 minutes of submersion in ice water and hypothermic arrest.
DuckDB Internals: Why Is DuckDB Fast? (Part 1)
A deep dive into the internal architecture of DuckDB and the reasons behind its high performance.
Analyzing the Narration Gap in LLM-Solver Loops
Research on the 'narration gap' in LLM-solver loops, showing how sound formal results can be inverted or corrupted during the natural language narration phase.
Configurable Clinical Information Extraction with Agentic RAG: What Works, What Breaks, and Why
Introduction of ACIE, an on-premise agentic RAG pipeline designed for clinical information extraction from complex patient contexts.
Which Pairs to Compare for LLM Post-Training?
A study on optimizing preference-based post-training for LLMs by identifying and labeling only the most informative comparison pairs for DPO.
Toten: Knowledge-Based Ontological Tokenization Of Physical Quantities And Technical Notation In Brazilian Portuguese
TOTEN is a knowledge-based ontological tokenization framework for technical notation in Brazilian Portuguese, outperforming statistical BPE.
AI4SE and SE4AI Exploration: A Decade Looking Back and Forward
A retrospective look at the intersection of AI and Systems Engineering (AI4SE and SE4AI) over the last decade, identifying research gaps.