AI/ML arXiv cs.AI

ITPEval: Benchmarking Formal Translation Across Interactive Theorem Provers

ITPEval is a benchmark for evaluating how well automated formal proofs can be translated between different interactive theorem provers like Lean 4 and Rocq.

AI/ML arXiv cs.AI

FORCE-Bench: A Benchmark, Dataset, and Evaluation Harness for Agentic AI in Enterprise Finance

FORCE-Bench provides an expert-annotated benchmark for evaluating agentic AI systems on complex operational finance workflows.

Cybersecurity arXiv cs.AI

The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception in Agentic AI

A study formalizes the 'Chronos Vulnerability', a new security threat where the internal belief system of stateful autonomous agents is compromised via memory-based attacks.

AI/ML arXiv cs.AI

Sophisticated Policies from Epistemic Priors

Researchers explore Sophisticated Inference in active inference, showing it combines epistemic drive with closed-loop control to solve stochastic environments.

AI/ML arXiv cs.AI

Knowledge-Centric Self-Improvement

The Knowledge-Centric Self-Improvement (KSI) protocol allows agents to improve via a shared, persistent knowledge base rather than through agent-centric optimization.

Other arXiv cs.AI

Edge Intelligence in Civil Aviation: Paradigms, Techniques, and Applications

This paper reviews the application of edge intelligence and distributed AI to reduce latency and bandwidth in the safety-critical civil aviation sector.

AI/ML arXiv cs.AI

Profile-Graph Memory for LLM Agents: Implicit Cross-Entity Traversal through Narrative Profiles

Introduces ProGraph, a two-layer memory architecture for LLM agents that improves multi-hop reasoning and precision recall over existing RAG and memory frameworks.

AI/ML arXiv cs.AI

Lifted Representation Hypothesis in Language Models

Proposes the 'lifted representation hypothesis', suggesting LLMs update memory via shared latent structures rather than isolated facts, and identifies vulnerabilities in nested rule learning.

AI/ML arXiv cs.AI

GraphContainer: A Unified Platform for Comparing and Debugging Graph RAG Methods

Presents GraphContainer, a platform that unifies and visualizes diverse Graph RAG workflows to facilitate comparison and debugging of retrieval behaviors.

AI/ML arXiv cs.AI

AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally

Introduces AdaRoPE, a variant of Rotary Position Embedding that uses learnable, head-specific rotation frequencies and scaling to improve long-context performance.

AI/ML arXiv cs.AI

Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models

Develops a transparent SAE-feature steering pipeline for behavioral control in LLMs using a statistically grounded Borda consensus approach.

AI/ML arXiv cs.AI

Logic-Guided Data Extraction with Answer Set Programming and Large Language Models

Combines LLM-based data extraction with Answer Set Programming (ASP) to ensure logical consistency and reduce the number of LLM calls required for complex extraction tasks.

AI/ML arXiv cs.AI

Geometry-Guided Constraint Learning for LLM Safety Classification

Explores the use of Sparse Autoencoders (SAE) to optimize safety classification in LLMs, suggesting safety boundaries admit low-dimensional linear descriptions in SAE feature space.

AI/ML arXiv cs.AI

Rethinking Uncertainty Evaluation in Large Language Models

Critiques current uncertainty evaluation in LLMs, arguing that calibration is insufficient and proposing the C1 metrics to measure coherent probabilistic beliefs.

AI/ML arXiv cs.AI

Spectral-LSH: Sub-Quadratic Prompt Compression via Krylov-Projected Locality-Sensitive Hashing

Introduces Spectral-LSH, a training-free prompt compression method using Krylov-projected Locality-Sensitive Hashing to reduce prefill attention costs.

AI/ML arXiv cs.AI

Beyond Tracking or Shortcut: Composition-Bounded Predictive States in Poker Autoregressive Models

Analyzes opponent-range representation in poker autoregressive models, finding that most predictive information stems from visible betting patterns rather than hidden states.

Software Engineering Hacker News

git's –end-of-options Flag

Discussion regarding the use of the --end-of-options flag in git to explicitly separate options from positional arguments.

AI/ML arXiv cs.AI

FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads

FineServe introduces a multi-model LLM serving workload dataset and generator to better benchmark routing and scheduling in production environments.

AI/ML arXiv cs.AI

Hybrid LSTM-Graph Neural Framework for Robust Financial Fraud Detection and Adversarial Resilience

FraudShield AI is a hybrid LSTM-Graph Neural framework designed to detect sophisticated financial fraud by combining temporal sequences and relational context.

AI/ML arXiv cs.AI

OpenEvoShield: Dual Non-Stationary Continual Defense for Open-World Multi-Agent System Attacks

OpenEvoShield is a co-evolutionary continual defense framework designed to protect multi-agent LLM systems from dynamic adversarial attacks.