AI/ML arXiv cs.AI

MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference

MemDecay is a training-free, region-aware KV-cache eviction policy that optimizes LLM agent inference by managing token priorities based on semantic structure.

AI/ML arXiv cs.AI

WasteAssistant: Regulation-Guided Visual Question Answering Framework for Intelligent Waste Segregation and Sustainable Managemen

WasteAssistant is a VQA framework and dataset for intelligent waste segregation based on Indian waste management regulations.

Other Hacker News

The Second Life of Sanskrit

A discussion on the modern application and revival of the Sanskrit language.

Tech Business/VC Hacker News

StubHub's 'marketplace for fans' is run by a mass scalper, SEC filings reveal

SEC filings reveal that StubHub's marketplace is effectively operated by a mass scalper.

Software Engineering Hacker News

Show HN: A Free RSS reader with a configurable recommendation engine

A free RSS reader featuring a configurable recommendation engine for personalized content discovery.

AI/ML arXiv cs.AI

Learning the Brain's Dynamics as a Port-Hamiltonian System

Researchers model the human motor cortex as a port-Hamiltonian system to improve BCI decoders and neuromodulation signals.

AI/ML arXiv cs.AI

Context by Distinct Information: An Auditable Dirichlet-Process Working Memory for Long, Redundant Context Streams

Introduces a working-memory component for LLMs that scales with distinct information rather than tokens, improving efficiency in long, redundant contexts.

AI/ML arXiv cs.AI

Annotation-Free Furniture Codes: What They Encode, and How Far They Transfer

A study on using self-supervised tokens from object geometry to replace human-annotated labels and poses in 3D scene synthesis.

AI/ML arXiv cs.AI

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards

Presents RLVP, a reinforcement learning framework that uses continuous physics rewards to post-train LLMs for multi-PDE solver code generation.

AI/ML arXiv cs.AI

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples

Introduces ARMOR, a framework to stabilize on-policy LLM RL training by using off-policy anchor samples to prevent over-optimization.

Cybersecurity arXiv cs.AI

Temporary Authority, Permanent Effects: Commit-Time Authorization for LLM Agents

Analyzes the security of LLM agents regarding commit-time authorization and introduces CommitGuard to block stale durable-effect attempts.

AI/ML arXiv cs.AI

Confining Nondeterminism: AI-Driven Research Systems as DBMSs for Reliable, Non-Wasteful, Transparent, and Collaborative Research [Vision]

Proposes treating AI-driven research systems as DBMSs, where LLMs act as query compilers rather than executors to ensure reliability and transparency.

AI/ML Hacker News

Guardian Angels: LLM Personalization for Productivity and Security

A discussion on the use of Large Language Models for personalizing productivity and enhancing security.

AI/ML arXiv cs.AI

VINE: Taming Generative Control Policies for Reinforcement Learning

Introduces VINE, a sampling method for flow-matching policies in RL that enables stable end-to-end value-gradient optimization for robotic manipulation.

AI/ML arXiv cs.AI

ABot-N1: Toward a General Visual Language Navigation Foundation Model

Presents ABot-N1, a Visual Language Navigation foundation model that decouples cognition from control using a slow-fast architecture to improve urban-scale navigation.

AI/ML arXiv cs.AI

Structured Thoughts For Improved Reasoning And Context Pruning

Introduces Structured Thoughts, a framework that organizes LLM reasoning into scratch and conclusion blocks to improve performance and enable context pruning.

AI/ML arXiv cs.AI

The evolution of AI from image interpretation toward scientific inference in nanoparticle electron microscopy

A review of AI's evolution in nanoparticle electron microscopy, moving from simple image interpretation to complex scientific inference and materials discovery.

AI/ML arXiv cs.AI

A Stepwise Questioning Expert-Editor Multi-Agent Framework for Long-Document Summarization

Proposes a multi-agent framework using an expert-editor stepwise questioning approach to improve long-document summarization with LLMs.

AI/ML arXiv cs.AI

SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding

Introduces SynthDocBench, a synthetic benchmark for long-context visual document understanding to identify failure modes in frontier Vision Language Models.

Cybersecurity arXiv cs.AI

Large Language Models in Misinformation Ecosystems: Misuse, Defense, and Vulnerability

Analyzes the ecosystem-level security challenges of LLMs in misinformation, introducing a role-layer framework for understanding risks and defenses.