All Articles
16513 articles total
ITPEval: Benchmarking Formal Translation Across Interactive Theorem Provers
ITPEval is a benchmark for evaluating how well automated formal proofs can be translated between different interactive theorem provers like Lean 4 and Rocq.
FORCE-Bench: A Benchmark, Dataset, and Evaluation Harness for Agentic AI in Enterprise Finance
FORCE-Bench provides an expert-annotated benchmark for evaluating agentic AI systems on complex operational finance workflows.
The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception in Agentic AI
A study formalizes the 'Chronos Vulnerability', a new security threat where the internal belief system of stateful autonomous agents is compromised via memory-based attacks.
Sophisticated Policies from Epistemic Priors
Researchers explore Sophisticated Inference in active inference, showing it combines epistemic drive with closed-loop control to solve stochastic environments.
Knowledge-Centric Self-Improvement
The Knowledge-Centric Self-Improvement (KSI) protocol allows agents to improve via a shared, persistent knowledge base rather than through agent-centric optimization.
Edge Intelligence in Civil Aviation: Paradigms, Techniques, and Applications
This paper reviews the application of edge intelligence and distributed AI to reduce latency and bandwidth in the safety-critical civil aviation sector.
Profile-Graph Memory for LLM Agents: Implicit Cross-Entity Traversal through Narrative Profiles
Introduces ProGraph, a two-layer memory architecture for LLM agents that improves multi-hop reasoning and precision recall over existing RAG and memory frameworks.
Lifted Representation Hypothesis in Language Models
Proposes the 'lifted representation hypothesis', suggesting LLMs update memory via shared latent structures rather than isolated facts, and identifies vulnerabilities in nested rule learning.
GraphContainer: A Unified Platform for Comparing and Debugging Graph RAG Methods
Presents GraphContainer, a platform that unifies and visualizes diverse Graph RAG workflows to facilitate comparison and debugging of retrieval behaviors.
AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally
Introduces AdaRoPE, a variant of Rotary Position Embedding that uses learnable, head-specific rotation frequencies and scaling to improve long-context performance.
Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models
Develops a transparent SAE-feature steering pipeline for behavioral control in LLMs using a statistically grounded Borda consensus approach.
Logic-Guided Data Extraction with Answer Set Programming and Large Language Models
Combines LLM-based data extraction with Answer Set Programming (ASP) to ensure logical consistency and reduce the number of LLM calls required for complex extraction tasks.
Geometry-Guided Constraint Learning for LLM Safety Classification
Explores the use of Sparse Autoencoders (SAE) to optimize safety classification in LLMs, suggesting safety boundaries admit low-dimensional linear descriptions in SAE feature space.
Rethinking Uncertainty Evaluation in Large Language Models
Critiques current uncertainty evaluation in LLMs, arguing that calibration is insufficient and proposing the C1 metrics to measure coherent probabilistic beliefs.
Spectral-LSH: Sub-Quadratic Prompt Compression via Krylov-Projected Locality-Sensitive Hashing
Introduces Spectral-LSH, a training-free prompt compression method using Krylov-projected Locality-Sensitive Hashing to reduce prefill attention costs.
Beyond Tracking or Shortcut: Composition-Bounded Predictive States in Poker Autoregressive Models
Analyzes opponent-range representation in poker autoregressive models, finding that most predictive information stems from visible betting patterns rather than hidden states.
git's –end-of-options Flag
Discussion regarding the use of the --end-of-options flag in git to explicitly separate options from positional arguments.
FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads
FineServe introduces a multi-model LLM serving workload dataset and generator to better benchmark routing and scheduling in production environments.
Hybrid LSTM-Graph Neural Framework for Robust Financial Fraud Detection and Adversarial Resilience
FraudShield AI is a hybrid LSTM-Graph Neural framework designed to detect sophisticated financial fraud by combining temporal sequences and relational context.
OpenEvoShield: Dual Non-Stationary Continual Defense for Open-World Multi-Agent System Attacks
OpenEvoShield is a co-evolutionary continual defense framework designed to protect multi-agent LLM systems from dynamic adversarial attacks.