AI/ML arXiv cs.AI

MetaResearcher: Scaling Deep Research via Self-Reflective Reinforcement Learning in Adversarial Virtual Environments

MetaResearcher is a framework for scaling deep research agents using self-reflective reinforcement learning and adversarial virtual environments.

AI/ML arXiv cs.AI

Multi-Agent Transactive Memory

Multi-Agent Transactive Memory (MATM) is a framework for storing and retrieving agent-generated trajectories to enable experience sharing across agent populations.

Tech Business/VC The Verge

Barret Zoph is out at OpenAI again after just five months

Barret Zoph, the head of enterprise AI sales at OpenAI, has departed the company after a brief five-month tenure.

AI/ML arXiv cs.AI

Exit-and-Join Dynamics for Decentralized Coalition Formation

A research paper proposes a decentralized dynamical process for coalition formation using the Aumann-Dreze value to evaluate local moves.

AI/ML arXiv cs.AI

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents

This study argues that aggregate-score leaderboards for LLM agents are unstable and proposes ranking configurations based on predictive validity.

AI/ML arXiv cs.AI

GLARE: A Natural Language Interface for Querying Global Explanations

GLARE is introduced as an LLM-based interactive interface that translates natural language questions into SQL queries to explore global explanations for image classifiers.

AI/ML arXiv cs.AI

Interpreting Neural Combinatorial Optimization via Evolving Programmatic Bottlenecks

The Evolving Programmatic Bottlenecks (EPB) framework is introduced to interpret Neural Combinatorial Optimization policies by distilling them into human-readable program portfolios.

AI/ML arXiv cs.AI

A Comparative Study of Pretrained Transformer Models for Quranic ASR: Speech Representations, Label Formats, and Dataset Composition

A comparative study examines the use of pretrained Transformer models like Wav2Vec2.0 and HuBERT for Quranic Automatic Speech Recognition (ASR).

AI/ML arXiv cs.AI

Benchmarking Agentic Review Systems

This paper benchmarks agentic review systems for research papers, finding that systems like OpenAIReview can track human quality judgments reasonably well.

AI/ML arXiv cs.AI

Grounded Inference: Principles for Deterministically Encapsulated Generative Models

The paper establishes principles for 'Grounded Inference' to deterministically encapsulate probabilistic generative models within traditional computational systems.

AI/ML arXiv cs.AI

Optimal Scheduling in a Question-Answering Forum of Knowledge Workers

A study explores optimal scheduling in question-answering forums employing knowledge workers to maintain system stability and capacity.

AI/ML arXiv cs.AI

Beyond Entropy: Learning from Token-Level Distributional Deviations for LLM Reasoning

The Independent Combinatorial Tokens (ICT) framework is proposed to stabilize LLM reasoning training by focusing updates on token-level distributional deviations.

Cybersecurity Hacker News

Let's Encrypt has been down most of today

Let's Encrypt experienced significant downtime throughout the day, impacting SSL/TLS certificate issuance and renewal.

Other Hacker News

Ice Water Drowning Survival After 147-Minute Submersion and Hypothermic Arrest

A medical case report detailing a rare survival instance after 147 minutes of submersion in ice water and hypothermic arrest.

Software Engineering Hacker News

DuckDB Internals: Why Is DuckDB Fast? (Part 1)

A deep dive into the internal architecture of DuckDB and the reasons behind its high performance.

AI/ML arXiv cs.AI

Analyzing the Narration Gap in LLM-Solver Loops

Research on the 'narration gap' in LLM-solver loops, showing how sound formal results can be inverted or corrupted during the natural language narration phase.

AI/ML arXiv cs.AI

Configurable Clinical Information Extraction with Agentic RAG: What Works, What Breaks, and Why

Introduction of ACIE, an on-premise agentic RAG pipeline designed for clinical information extraction from complex patient contexts.

AI/ML arXiv cs.AI

Which Pairs to Compare for LLM Post-Training?

A study on optimizing preference-based post-training for LLMs by identifying and labeling only the most informative comparison pairs for DPO.

AI/ML arXiv cs.AI

Toten: Knowledge-Based Ontological Tokenization Of Physical Quantities And Technical Notation In Brazilian Portuguese

TOTEN is a knowledge-based ontological tokenization framework for technical notation in Brazilian Portuguese, outperforming statistical BPE.

AI/ML arXiv cs.AI

AI4SE and SE4AI Exploration: A Decade Looking Back and Forward

A retrospective look at the intersection of AI and Systems Engineering (AI4SE and SE4AI) over the last decade, identifying research gaps.