AI/ML arXiv cs.AI

Towards Verifiable Agentic Data Science: Solving Irregular TSQA Via Tool-Grounded Reasoning

Introduction of IRTS-ToolBench, a benchmark for evaluating LLM-based AI agents on irregular time series question answering (TSQA).

AI/ML arXiv cs.AI

CONCORD: Asynchronous Sparse Aggregation for Device-Cloud RAG under Document Isolation

CONCORD is an asynchronous sparse aggregation framework for device-cloud RAG that maintains document isolation for privacy while improving throughput.

AI/ML arXiv cs.AI

CogGuard: Cognitive and Operational Profiling for Proactive Warning in Edge Intelligent Services

CogGuard is a proactive-warning framework for edge intelligent services that decouples profile construction from online score prediction for efficiency.

Cybersecurity arXiv cs.AI

Attribute Inference from Interactive Targeted Ads

Research on attribute inference from interactive targeted ads, providing a model and benchmark to evaluate privacy risks and defense mechanisms.

AI/ML arXiv cs.AI

Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning

Visual-Seeker is a visual-native multimodal deep search agent that uses active visual reasoning to improve factual grounding in open-world scenarios.

Other Hacker News

Laser Phase Plate Cryo-Electron Microscopy

Discussion regarding Laser Phase Plate Cryo-Electron Microscopy.

AI/ML arXiv cs.AI

A Definition of Good Explanations and the Challenges Explaining LLM Outputs

A philosophical and technical exploration of what constitutes a 'good explanation' and the challenges of providing such explanations for LLM outputs.

AI/ML arXiv cs.AI

Dr-DCI: Scaling Direct Corpus Interaction via Dynamic Workspace Expansion

Introduces DR-DCI, a framework that combines retriever-steered direct corpus interaction to scale agentic search over large document collections.

AI/ML arXiv cs.AI

Relational Structural Causal Models

Proposes Relational Structural Causal Models to allow AI to reason about interventions and counterfactuals in environments with varying objects and relations.

AI/ML arXiv cs.AI

Trust Between AI Agents: Measuring Formation, Breakage, and Recovery, with Implications for Governing Multi-Agent Systems

Develops a behavioral measure for trust between AI agents based on costly verification in cooperative survival games.

AI/ML arXiv cs.AI

PrologMCP: A Standardized Prolog Tool Interface for LLM Agents

Introduces PrologMCP, an open-source server that exposes Prolog as a tool for LLM agents via the Model Context Protocol (MCP) to improve deductive reasoning.

AI/ML arXiv cs.AI

Semantics-Enhanced Retrieval-Augmented Time Series Forecasting

Presents SERAF, a multimodal RAG framework that uses both time series similarity and self-generated textual descriptions for better forecasting.

AI/ML arXiv cs.AI

AI Engram: In Search of Memory Traces in Artificial Intelligence

Introduces a geometric framework to identify 'AI engrams' (memory traces) in deep neural networks, allowing for surgical manipulation of learned knowledge.

AI/ML arXiv cs.AI

Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability

Proposes Metric Match, a subset selection method to estimate the reliability of LLM judges using limited human annotations.

AI/ML arXiv cs.AI

OSGuard: A Benchmark for Safety in Computer-Use Agents

Introduces OSGuard, a benchmark suite for evaluating the safety of computer-use agents, focusing on action-level guardrails and risk-augmented execution.

Hardware/Chips Hacker News

I Hacked into the Worst E-Bike and Fixed It [video]

A video demonstration of hacking into a poorly secured e-bike and implementing custom fixes.

AI/ML arXiv cs.AI

Federated Causal Inference from Multi-Site Observational Data via Propensity Score Aggregation

A novel federated learning approach for causal inference that estimates treatment effects from decentralized observational data without needing individual-level data access.

AI/ML arXiv cs.AI

Sentinel: Decoding Context Utilization via Attention Probing for Efficient LLM Context Compression

Sentinel is a lightweight sentence-level compression framework for LLMs that uses attention probing to remove noisy retrieved contexts in RAG pipelines.

AI/ML arXiv cs.AI

DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation

DiffusionBlocks enables memory-efficient block-wise training for transformers by treating residual connections as a denoising process, reducing memory needs proportional to the number of blocks.

AI/ML arXiv cs.AI

UltraSketchLLM: Sub-1-Bit LLM Compression via Sketch and Hardware-Friendly Operators

UltraSketchLLM achieves extreme LLM weight compression down to 0.5 bit per weight using data sketching and hardware-friendly operators.