All Articles
18737 articles total
The time the x86 emulator team found code so bad they fixed it during emulation
A discussion on a case where an x86 emulator team identified and fixed bugs in the original software being emulated due to the poor quality of the original code.
Fusion is not one-size-fits-all: Cross-Modal Representation Alignment for Time-to-Event Modeling
A new framework for cross-modal representation alignment between CT imaging and EHR data to improve time-to-event prediction in clinical settings.
Risk-Aware LLM Agents for Geospatial Data Retrieval: Design and Preliminary Adversarial Evaluation
An LLM-driven framework designed for retrieving remote sensing data from geospatial catalogues using natural language queries via a multi-agent architecture.
Cognitive Debt: AI as Intellectual Leverage and the Dynamics of Systemic Fragility
A formal theory on 'cognitive debt,' exploring how substituting first-principles reasoning with AI leads to systemic fragility and intellectual erosion.
VGPT-RSI for RH-Adjacent Formal Progress: Boundary Certificates, Verified Finite Lagarias Inequalities, and Explicit Failure Localization
The use of VGPT-RSI, an AI-assisted reasoning system, to produce formally verified partial progress on the Riemann Hypothesis using Coq and interval arithmetic.
Towards Verifiable Agentic Data Science: Solving Irregular TSQA Via Tool-Grounded Reasoning
Introduction of IRTS-ToolBench, a benchmark for evaluating LLM-based AI agents on irregular time series question answering (TSQA).
CONCORD: Asynchronous Sparse Aggregation for Device-Cloud RAG under Document Isolation
CONCORD is an asynchronous sparse aggregation framework for device-cloud RAG that maintains document isolation for privacy while improving throughput.
CogGuard: Cognitive and Operational Profiling for Proactive Warning in Edge Intelligent Services
CogGuard is a proactive-warning framework for edge intelligent services that decouples profile construction from online score prediction for efficiency.
Attribute Inference from Interactive Targeted Ads
Research on attribute inference from interactive targeted ads, providing a model and benchmark to evaluate privacy risks and defense mechanisms.
Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning
Visual-Seeker is a visual-native multimodal deep search agent that uses active visual reasoning to improve factual grounding in open-world scenarios.
Laser Phase Plate Cryo-Electron Microscopy
Discussion regarding Laser Phase Plate Cryo-Electron Microscopy.
A Definition of Good Explanations and the Challenges Explaining LLM Outputs
A philosophical and technical exploration of what constitutes a 'good explanation' and the challenges of providing such explanations for LLM outputs.
Dr-DCI: Scaling Direct Corpus Interaction via Dynamic Workspace Expansion
Introduces DR-DCI, a framework that combines retriever-steered direct corpus interaction to scale agentic search over large document collections.
Relational Structural Causal Models
Proposes Relational Structural Causal Models to allow AI to reason about interventions and counterfactuals in environments with varying objects and relations.
Trust Between AI Agents: Measuring Formation, Breakage, and Recovery, with Implications for Governing Multi-Agent Systems
Develops a behavioral measure for trust between AI agents based on costly verification in cooperative survival games.
PrologMCP: A Standardized Prolog Tool Interface for LLM Agents
Introduces PrologMCP, an open-source server that exposes Prolog as a tool for LLM agents via the Model Context Protocol (MCP) to improve deductive reasoning.
Semantics-Enhanced Retrieval-Augmented Time Series Forecasting
Presents SERAF, a multimodal RAG framework that uses both time series similarity and self-generated textual descriptions for better forecasting.
AI Engram: In Search of Memory Traces in Artificial Intelligence
Introduces a geometric framework to identify 'AI engrams' (memory traces) in deep neural networks, allowing for surgical manipulation of learned knowledge.
Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability
Proposes Metric Match, a subset selection method to estimate the reliability of LLM judges using limited human annotations.
OSGuard: A Benchmark for Safety in Computer-Use Agents
Introduces OSGuard, a benchmark suite for evaluating the safety of computer-use agents, focusing on action-level guardrails and risk-augmented execution.