Software Engineering Hacker News

The time the x86 emulator team found code so bad they fixed it during emulation

A discussion on a case where an x86 emulator team identified and fixed bugs in the original software being emulated due to the poor quality of the original code.

AI/ML arXiv cs.AI

Fusion is not one-size-fits-all: Cross-Modal Representation Alignment for Time-to-Event Modeling

A new framework for cross-modal representation alignment between CT imaging and EHR data to improve time-to-event prediction in clinical settings.

AI/ML arXiv cs.AI

Risk-Aware LLM Agents for Geospatial Data Retrieval: Design and Preliminary Adversarial Evaluation

An LLM-driven framework designed for retrieving remote sensing data from geospatial catalogues using natural language queries via a multi-agent architecture.

AI/ML arXiv cs.AI

Cognitive Debt: AI as Intellectual Leverage and the Dynamics of Systemic Fragility

A formal theory on 'cognitive debt,' exploring how substituting first-principles reasoning with AI leads to systemic fragility and intellectual erosion.

AI/ML arXiv cs.AI

VGPT-RSI for RH-Adjacent Formal Progress: Boundary Certificates, Verified Finite Lagarias Inequalities, and Explicit Failure Localization

The use of VGPT-RSI, an AI-assisted reasoning system, to produce formally verified partial progress on the Riemann Hypothesis using Coq and interval arithmetic.

AI/ML arXiv cs.AI

Towards Verifiable Agentic Data Science: Solving Irregular TSQA Via Tool-Grounded Reasoning

Introduction of IRTS-ToolBench, a benchmark for evaluating LLM-based AI agents on irregular time series question answering (TSQA).

AI/ML arXiv cs.AI

CONCORD: Asynchronous Sparse Aggregation for Device-Cloud RAG under Document Isolation

CONCORD is an asynchronous sparse aggregation framework for device-cloud RAG that maintains document isolation for privacy while improving throughput.

AI/ML arXiv cs.AI

CogGuard: Cognitive and Operational Profiling for Proactive Warning in Edge Intelligent Services

CogGuard is a proactive-warning framework for edge intelligent services that decouples profile construction from online score prediction for efficiency.

Cybersecurity arXiv cs.AI

Attribute Inference from Interactive Targeted Ads

Research on attribute inference from interactive targeted ads, providing a model and benchmark to evaluate privacy risks and defense mechanisms.

AI/ML arXiv cs.AI

Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning

Visual-Seeker is a visual-native multimodal deep search agent that uses active visual reasoning to improve factual grounding in open-world scenarios.

Other Hacker News

Laser Phase Plate Cryo-Electron Microscopy

Discussion regarding Laser Phase Plate Cryo-Electron Microscopy.

AI/ML arXiv cs.AI

A Definition of Good Explanations and the Challenges Explaining LLM Outputs

A philosophical and technical exploration of what constitutes a 'good explanation' and the challenges of providing such explanations for LLM outputs.

AI/ML arXiv cs.AI

Dr-DCI: Scaling Direct Corpus Interaction via Dynamic Workspace Expansion

Introduces DR-DCI, a framework that combines retriever-steered direct corpus interaction to scale agentic search over large document collections.

AI/ML arXiv cs.AI

Relational Structural Causal Models

Proposes Relational Structural Causal Models to allow AI to reason about interventions and counterfactuals in environments with varying objects and relations.

AI/ML arXiv cs.AI

Trust Between AI Agents: Measuring Formation, Breakage, and Recovery, with Implications for Governing Multi-Agent Systems

Develops a behavioral measure for trust between AI agents based on costly verification in cooperative survival games.

AI/ML arXiv cs.AI

PrologMCP: A Standardized Prolog Tool Interface for LLM Agents

Introduces PrologMCP, an open-source server that exposes Prolog as a tool for LLM agents via the Model Context Protocol (MCP) to improve deductive reasoning.

AI/ML arXiv cs.AI

Semantics-Enhanced Retrieval-Augmented Time Series Forecasting

Presents SERAF, a multimodal RAG framework that uses both time series similarity and self-generated textual descriptions for better forecasting.

AI/ML arXiv cs.AI

AI Engram: In Search of Memory Traces in Artificial Intelligence

Introduces a geometric framework to identify 'AI engrams' (memory traces) in deep neural networks, allowing for surgical manipulation of learned knowledge.

AI/ML arXiv cs.AI

Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability

Proposes Metric Match, a subset selection method to estimate the reliability of LLM judges using limited human annotations.

AI/ML arXiv cs.AI

OSGuard: A Benchmark for Safety in Computer-Use Agents

Introduces OSGuard, a benchmark suite for evaluating the safety of computer-use agents, focusing on action-level guardrails and risk-augmented execution.