All Articles
19232 articles total
Towards Verifiable Agentic Data Science: Solving Irregular TSQA Via Tool-Grounded Reasoning
Introduction of IRTS-ToolBench, a benchmark for evaluating LLM-based AI agents on irregular time series question answering (TSQA).
CONCORD: Asynchronous Sparse Aggregation for Device-Cloud RAG under Document Isolation
CONCORD is an asynchronous sparse aggregation framework for device-cloud RAG that maintains document isolation for privacy while improving throughput.
CogGuard: Cognitive and Operational Profiling for Proactive Warning in Edge Intelligent Services
CogGuard is a proactive-warning framework for edge intelligent services that decouples profile construction from online score prediction for efficiency.
Attribute Inference from Interactive Targeted Ads
Research on attribute inference from interactive targeted ads, providing a model and benchmark to evaluate privacy risks and defense mechanisms.
Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning
Visual-Seeker is a visual-native multimodal deep search agent that uses active visual reasoning to improve factual grounding in open-world scenarios.
Laser Phase Plate Cryo-Electron Microscopy
Discussion regarding Laser Phase Plate Cryo-Electron Microscopy.
A Definition of Good Explanations and the Challenges Explaining LLM Outputs
A philosophical and technical exploration of what constitutes a 'good explanation' and the challenges of providing such explanations for LLM outputs.
Dr-DCI: Scaling Direct Corpus Interaction via Dynamic Workspace Expansion
Introduces DR-DCI, a framework that combines retriever-steered direct corpus interaction to scale agentic search over large document collections.
Relational Structural Causal Models
Proposes Relational Structural Causal Models to allow AI to reason about interventions and counterfactuals in environments with varying objects and relations.
Trust Between AI Agents: Measuring Formation, Breakage, and Recovery, with Implications for Governing Multi-Agent Systems
Develops a behavioral measure for trust between AI agents based on costly verification in cooperative survival games.
PrologMCP: A Standardized Prolog Tool Interface for LLM Agents
Introduces PrologMCP, an open-source server that exposes Prolog as a tool for LLM agents via the Model Context Protocol (MCP) to improve deductive reasoning.
Semantics-Enhanced Retrieval-Augmented Time Series Forecasting
Presents SERAF, a multimodal RAG framework that uses both time series similarity and self-generated textual descriptions for better forecasting.
AI Engram: In Search of Memory Traces in Artificial Intelligence
Introduces a geometric framework to identify 'AI engrams' (memory traces) in deep neural networks, allowing for surgical manipulation of learned knowledge.
Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability
Proposes Metric Match, a subset selection method to estimate the reliability of LLM judges using limited human annotations.
OSGuard: A Benchmark for Safety in Computer-Use Agents
Introduces OSGuard, a benchmark suite for evaluating the safety of computer-use agents, focusing on action-level guardrails and risk-augmented execution.
I Hacked into the Worst E-Bike and Fixed It [video]
A video demonstration of hacking into a poorly secured e-bike and implementing custom fixes.
Federated Causal Inference from Multi-Site Observational Data via Propensity Score Aggregation
A novel federated learning approach for causal inference that estimates treatment effects from decentralized observational data without needing individual-level data access.
Sentinel: Decoding Context Utilization via Attention Probing for Efficient LLM Context Compression
Sentinel is a lightweight sentence-level compression framework for LLMs that uses attention probing to remove noisy retrieved contexts in RAG pipelines.
DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation
DiffusionBlocks enables memory-efficient block-wise training for transformers by treating residual connections as a denoising process, reducing memory needs proportional to the number of blocks.
UltraSketchLLM: Sub-1-Bit LLM Compression via Sketch and Hardware-Friendly Operators
UltraSketchLLM achieves extreme LLM weight compression down to 0.5 bit per weight using data sketching and hardware-friendly operators.