AI/ML arXiv cs.AI

LeanFlow: A Case Study in Workflow-Driven Lean Autoformalization

LeanFlow is an LLM agent system that automates the translation of mathematical papers into formal Lean projects.

AI/ML arXiv cs.AI

Optimizing Hypergraph-Based RAG: Toward Better Fact Extraction and Chunk Retrieval

HyperGraphRAG enhances knowledge graph retrieval by using hypergraphs for richer semantics and self-consistency prompting for fact extraction.

AI/ML arXiv cs.AI

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference

MiniCache is a caching framework that uses small models as interfaces to enable efficient, reusable computation across structurally similar LLM requests.

AI/ML arXiv cs.AI

Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain

Telco-GAIA is a bilingual, multi-modal benchmark for evaluating tool-using agents in the telecommunications domain.

AI/ML arXiv cs.AI

SiGMA: Sign-Guided Merging and Adaptation for Multimodal Continual Instruction Tuning

SiGMA is a framework designed to mitigate negative interference during multimodal continual instruction tuning to prevent catastrophic forgetting.

AI/ML arXiv cs.AI

Reliability-Aware LLM Alignment from Inconsistent Human Feedback

RGPO is a robust framework designed to mitigate the impact of inconsistent human feedback during RLHF for LLM alignment.

Software Engineering Hacker News

Modula-3 History Collection on Computer History Museum

The Computer History Museum has released a collection of historical materials related to the Modula-3 programming language.

AI/ML arXiv cs.AI

CRAWO: Custom Resources for Adaptive Workload Orchestration

CRAWO is a new framework for coordinating AI pipelines across distributed edge environments using a hardware-aware allocator and Kubernetes CRDs.

AI/ML arXiv cs.AI

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making

DFAH-Bench introduces a replay benchmark to measure behavioral instability in financial AI agents, revealing a gap between decision agreement and process consistency.

AI/ML arXiv cs.AI

Attention-based Experience Replay Framework for Continual Learning of Agnostic Time Series Forecasting Models

A new attention-based experience replay framework is proposed for continual learning in time series forecasting to mitigate catastrophic forgetting.

Cybersecurity arXiv cs.AI

Isolating LLM Alignment from Regex: Zero Coverage and Metric-Dependent Divergence Under Adversarial Mutation

Research shows that LLM alignment often provides zero additional coverage over regex filters for natural-language harmful requests but is more effective against adversarial probes.

AI/ML arXiv cs.AI

Workload-Aware Caching for Multi-Agent Systems

A workload-aware caching policy for multi-agent systems is introduced, reducing latency by up to 64.7% by considering recomputation cost and DAG dependencies.

AI/ML arXiv cs.AI

From Errors to Rules: Iterative Prompt Optimization for Text Classification

ERGO is a new iterative prompt optimization method for text classification that focuses on diagnosing and rewriting rules based on classification failures.

AI/ML arXiv cs.AI

AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs

AISE-Bench is a full-cycle curated benchmark for information seeking on academic knowledge graphs, providing 1,133 QA pairs and execution trajectories.

AI/ML arXiv cs.AI

ExecuGraph: A Multi-Agent, Execution-Grounded Framework for Reliable Backend Code Synthesis with Large Language Models

ExecuGraph is a multi-agent framework for reliable backend code synthesis that uses execution-based validation and a typed directed workflow.

AI/ML arXiv cs.AI

FlowEdit: Information-Theoretic Control of LLM Reasoning Flows for Ill-posed Problems Involving Conflicts

FlowEdit uses information-theoretic principles to regulate LLM reasoning flows, helping models handle ill-posed problems with conflicting conditions.

Other Hacker News

Zitron: The Subprime Datacenter Crisis

A discussion regarding the potential for a subprime datacenter crisis, likely referring to infrastructure instability or financial risks in the data center sector.

AI/ML arXiv cs.AI

Routing Without Training: Controllable-Ratio LLM Offloading via Reliability Gating

Introduction of CARGO, a training-free framework for efficient LLM offloading between local and cloud environments based on model agreement.

AI/ML arXiv cs.AI

PersonaTrail: Benchmarking Personalized Web Agents through Browsing Trails

Introduction of PersonaTrail, a benchmark for evaluating personalized web agents using realistic browsing history, and the PACMem framework for structured memory.

AI/ML arXiv cs.AI

Tractable Hierarchical Control of Autoregressive Language Models

A method for steering autoregressive LLM generation towards syntactically valid output by distilling them into tractable probabilistic models.