All Articles
17646 articles total
SUNTA: Hierarchical Video Prediction with Surprise-based Chunking
Introduces SUNTA, a hierarchical video prediction method that uses surprise-based chunking to maintain accurate long-horizon predictions.
ContextNest: Verifiable Context Governance for Autonomous AI Agent
Introduces ContextNest, an open specification and reference implementation for governed AI knowledge vaults to ensure provenance and integrity in AI agents.
Enhancing Fitness Intelligence through Domain-Specific LLM Post-Training
Introduces FitOne, a series of domain-specific LLMs post-trained for scientific fitness coaching to improve reliability over general-purpose models.
Coding-agents can replicate scientific machine learning papers
Presents a workflow called Paper-replication that enables coding agents to systematically replicate computational claims in scientific machine learning papers.
A$^{2}$utoLPBench: An Auto-Generated, Agent-Friendly LP Benchmark via Inverse-KKT Construction
Introduces A2utoLPBench, an auto-generated benchmark for testing LLM agents on linear programming problems, designed to prevent data leakage.
I Type Holes in Keyboard Covers – This One Survived
A personal anecdote about keyboard covers and their durability.
ElephantAgent: Contextual State Continuity in Agentic Systems
ElephantAgent introduces a protocol for Contextual State Continuity to protect agentic systems from contextual state poisoning attacks.
A-TMA: Decoupling State-Aware Memory Failures in Long-Term Agent Memory
The ATMA framework addresses 'ghost memory' in long-term agent memory by decoupling bank maintenance, retrieval, and answer-time resolution.
Atomic Task Graph: A Unified Framework for Agentic Planning and Execution
Atomic Task Graph (ATG) provides a unified framework for agentic planning and execution using explicit graphs to improve efficiency and error recovery.
OntoLearner: A Modular Python Library for Ontology Learning with Large Language Models
OntoLearner is an open-source Python library for ontology learning with LLMs, featuring 180 machine-readable ontologies and standardized benchmarking.
Multimodal Knowledge Edit-Scoped Generalization for Online Recursive MLLM Editing
ScopeEdit is an online editor for multimodal LLMs that controls the propagation boundary of knowledge edits to prevent leakage and ensure cross-modal transfer.
Episodic-to-Semantic Consolidation Without Identity Drift
A proposal for consolidating episodic memory into semantic knowledge for agents without altering their cryptographically certified identity.
Traceable Fault Diagnosis for Battery Energy Storage Systems via Retrieval-Augmented Multi-Agent O&M Assistant
A multi-agent RAG-based assistant designed for traceable fault diagnosis in Battery Energy Storage Systems (BESS).
InduceKV: Fixed-Footprint Continual Adaptation of Multimodal LLMs via Inducing KV Memories
InduceKV enables fixed-footprint continual adaptation of MLLMs by storing training prefixes as compact KV payloads in a retrieval-based memory.
Hidden Forgetting in Continual Multimodal Learning: When Accuracy Survives but Grounding Fails
The RCL framework aims to prevent 'hidden evidence-use forgetting' in continual multimodal learning by preserving the evidence paths behind correct answers.
Ask HN: Is anyone experimenting with different ways of using LLMs for coding?
A community discussion on Hacker News regarding various experimental methods for utilizing Large Language Models in coding workflows.
Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge
Introduction of Pre-Flight, an open-source benchmark for evaluating LLM reasoning and safety in aviation operational knowledge.
Actual causality in fault trees
A theoretical study applying Halpern & Pearl's theory of actual causality to fault trees to improve failure diagnostics in complex systems.
CLAP: Closed-Loop Training, Evaluation, and Release Control for Domain Agent Post-training
Proposal of CLAP, a closed-loop method for post-training domain agents to improve data quality and release stability in manufacturing scenarios.
Safety Targeted Embedding Exploit via Refinement
Research demonstrating a gradient-guided attack (STEER) that bypasses LLM safety filters by translating refusal-triggering words into low-resource languages.