AI/ML arXiv cs.AI

CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks?

CODA-BENCH is a new benchmark evaluating the ability of code agents to handle data-intensive tasks within a Linux sandbox.

Cybersecurity arXiv cs.AI

Forced Deferral: Manipulating Routing Decisions in Multimodal LLM Cascades

Researchers introduce the Forced Deferral Attack (FDA), showing how multimodal LLM cascades can be manipulated to force expensive model usage.

AI/ML arXiv cs.AI

ChatPlanner: A Large Language Model Framework for Personalized Public Transit Routing

ChatPlanner uses fine-tuned LLMs and RAG to create personalized public transit routing based on natural language user preferences.

AI/ML arXiv cs.AI

APEX: Adaptive Principle EXtraction A Three-Layer Self-Evolution Framework for Production AI Agents

APEX is a three-layer self-evolution framework that allows AI agents to autonomously evolve their harness, principles, and workflow topology.

AI/ML arXiv cs.AI

S1-DeepResearch: Beyond Search, Toward Real-World Long-Horizon Research Agents

S1-DeepResearch-32B is a new open-source model trained on a unified trajectory paradigm for long-horizon research tasks and knowledge synthesis.

Software Engineering Hacker News

The time the x86 emulator team found code so bad they fixed it during emulation

A discussion on a case where an x86 emulator team identified and fixed bugs in the original software being emulated due to the poor quality of the original code.

AI/ML arXiv cs.AI

Fusion is not one-size-fits-all: Cross-Modal Representation Alignment for Time-to-Event Modeling

A new framework for cross-modal representation alignment between CT imaging and EHR data to improve time-to-event prediction in clinical settings.

AI/ML arXiv cs.AI

Risk-Aware LLM Agents for Geospatial Data Retrieval: Design and Preliminary Adversarial Evaluation

An LLM-driven framework designed for retrieving remote sensing data from geospatial catalogues using natural language queries via a multi-agent architecture.

AI/ML arXiv cs.AI

Cognitive Debt: AI as Intellectual Leverage and the Dynamics of Systemic Fragility

A formal theory on 'cognitive debt,' exploring how substituting first-principles reasoning with AI leads to systemic fragility and intellectual erosion.

AI/ML arXiv cs.AI

VGPT-RSI for RH-Adjacent Formal Progress: Boundary Certificates, Verified Finite Lagarias Inequalities, and Explicit Failure Localization

The use of VGPT-RSI, an AI-assisted reasoning system, to produce formally verified partial progress on the Riemann Hypothesis using Coq and interval arithmetic.

AI/ML arXiv cs.AI

Towards Verifiable Agentic Data Science: Solving Irregular TSQA Via Tool-Grounded Reasoning

Introduction of IRTS-ToolBench, a benchmark for evaluating LLM-based AI agents on irregular time series question answering (TSQA).

AI/ML arXiv cs.AI

CONCORD: Asynchronous Sparse Aggregation for Device-Cloud RAG under Document Isolation

CONCORD is an asynchronous sparse aggregation framework for device-cloud RAG that maintains document isolation for privacy while improving throughput.

AI/ML arXiv cs.AI

CogGuard: Cognitive and Operational Profiling for Proactive Warning in Edge Intelligent Services

CogGuard is a proactive-warning framework for edge intelligent services that decouples profile construction from online score prediction for efficiency.

Cybersecurity arXiv cs.AI

Attribute Inference from Interactive Targeted Ads

Research on attribute inference from interactive targeted ads, providing a model and benchmark to evaluate privacy risks and defense mechanisms.

AI/ML arXiv cs.AI

Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning

Visual-Seeker is a visual-native multimodal deep search agent that uses active visual reasoning to improve factual grounding in open-world scenarios.

Other Hacker News

Laser Phase Plate Cryo-Electron Microscopy

Discussion regarding Laser Phase Plate Cryo-Electron Microscopy.

AI/ML arXiv cs.AI

A Definition of Good Explanations and the Challenges Explaining LLM Outputs

A philosophical and technical exploration of what constitutes a 'good explanation' and the challenges of providing such explanations for LLM outputs.

AI/ML arXiv cs.AI

Dr-DCI: Scaling Direct Corpus Interaction via Dynamic Workspace Expansion

Introduces DR-DCI, a framework that combines retriever-steered direct corpus interaction to scale agentic search over large document collections.

AI/ML arXiv cs.AI

Relational Structural Causal Models

Proposes Relational Structural Causal Models to allow AI to reason about interventions and counterfactuals in environments with varying objects and relations.

AI/ML arXiv cs.AI

Trust Between AI Agents: Measuring Formation, Breakage, and Recovery, with Implications for Governing Multi-Agent Systems

Develops a behavioral measure for trust between AI agents based on costly verification in cooperative survival games.