AI/ML arXiv cs.AI

SUNTA: Hierarchical Video Prediction with Surprise-based Chunking

Introduces SUNTA, a hierarchical video prediction method that uses surprise-based chunking to maintain accurate long-horizon predictions.

AI/ML arXiv cs.AI

ContextNest: Verifiable Context Governance for Autonomous AI Agent

Introduces ContextNest, an open specification and reference implementation for governed AI knowledge vaults to ensure provenance and integrity in AI agents.

AI/ML arXiv cs.AI

Enhancing Fitness Intelligence through Domain-Specific LLM Post-Training

Introduces FitOne, a series of domain-specific LLMs post-trained for scientific fitness coaching to improve reliability over general-purpose models.

AI/ML arXiv cs.AI

Coding-agents can replicate scientific machine learning papers

Presents a workflow called Paper-replication that enables coding agents to systematically replicate computational claims in scientific machine learning papers.

AI/ML arXiv cs.AI

A$^{2}$utoLPBench: An Auto-Generated, Agent-Friendly LP Benchmark via Inverse-KKT Construction

Introduces A2utoLPBench, an auto-generated benchmark for testing LLM agents on linear programming problems, designed to prevent data leakage.

Other Hacker News

I Type Holes in Keyboard Covers – This One Survived

A personal anecdote about keyboard covers and their durability.

AI/ML arXiv cs.AI

ElephantAgent: Contextual State Continuity in Agentic Systems

ElephantAgent introduces a protocol for Contextual State Continuity to protect agentic systems from contextual state poisoning attacks.

AI/ML arXiv cs.AI

A-TMA: Decoupling State-Aware Memory Failures in Long-Term Agent Memory

The ATMA framework addresses 'ghost memory' in long-term agent memory by decoupling bank maintenance, retrieval, and answer-time resolution.

AI/ML arXiv cs.AI

Atomic Task Graph: A Unified Framework for Agentic Planning and Execution

Atomic Task Graph (ATG) provides a unified framework for agentic planning and execution using explicit graphs to improve efficiency and error recovery.

AI/ML arXiv cs.AI

OntoLearner: A Modular Python Library for Ontology Learning with Large Language Models

OntoLearner is an open-source Python library for ontology learning with LLMs, featuring 180 machine-readable ontologies and standardized benchmarking.

AI/ML arXiv cs.AI

Multimodal Knowledge Edit-Scoped Generalization for Online Recursive MLLM Editing

ScopeEdit is an online editor for multimodal LLMs that controls the propagation boundary of knowledge edits to prevent leakage and ensure cross-modal transfer.

AI/ML arXiv cs.AI

Episodic-to-Semantic Consolidation Without Identity Drift

A proposal for consolidating episodic memory into semantic knowledge for agents without altering their cryptographically certified identity.

AI/ML arXiv cs.AI

Traceable Fault Diagnosis for Battery Energy Storage Systems via Retrieval-Augmented Multi-Agent O&M Assistant

A multi-agent RAG-based assistant designed for traceable fault diagnosis in Battery Energy Storage Systems (BESS).

AI/ML arXiv cs.AI

InduceKV: Fixed-Footprint Continual Adaptation of Multimodal LLMs via Inducing KV Memories

InduceKV enables fixed-footprint continual adaptation of MLLMs by storing training prefixes as compact KV payloads in a retrieval-based memory.

AI/ML arXiv cs.AI

Hidden Forgetting in Continual Multimodal Learning: When Accuracy Survives but Grounding Fails

The RCL framework aims to prevent 'hidden evidence-use forgetting' in continual multimodal learning by preserving the evidence paths behind correct answers.

AI/ML Hacker News

Ask HN: Is anyone experimenting with different ways of using LLMs for coding?

A community discussion on Hacker News regarding various experimental methods for utilizing Large Language Models in coding workflows.

AI/ML arXiv cs.AI

Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge

Introduction of Pre-Flight, an open-source benchmark for evaluating LLM reasoning and safety in aviation operational knowledge.

AI/ML arXiv cs.AI

Actual causality in fault trees

A theoretical study applying Halpern & Pearl's theory of actual causality to fault trees to improve failure diagnostics in complex systems.

AI/ML arXiv cs.AI

CLAP: Closed-Loop Training, Evaluation, and Release Control for Domain Agent Post-training

Proposal of CLAP, a closed-loop method for post-training domain agents to improve data quality and release stability in manufacturing scenarios.

Cybersecurity arXiv cs.AI

Safety Targeted Embedding Exploit via Refinement

Research demonstrating a gradient-guided attack (STEER) that bypasses LLM safety filters by translating refusal-triggering words into low-resource languages.