AI/ML arXiv cs.AI

Controlling Tool Use with Heading-Specific Activation Steering

Research on using activation steering vectors to control and suppress unnecessary tool invocation in tool-augmented LLMs.

AI/ML arXiv cs.AI

From Passive Retrieval to Active Memory Navigation: Learning to Use Memory as a Structured Action Space

Introduction of NapMem, a framework that treats long-term user memory as a structured action space for conversational agents.

Other Hacker News

Show HN: Neil the Seal Game

A new game called 'Neil the Seal' has been showcased on Hacker News.

AI/ML arXiv cs.AI

Prompt-to-Paper: Agentic AI System for Bioinformatics

Prompt-to-Paper is a multi-agent AI framework designed to automate the generation of bioinformatics manuscripts by grounding claims in verifiable literature and executing real computational experiments.

AI/ML arXiv cs.AI

From Graphs to Gradients: Physics-Inspired Structural Attribution for Cyber-Physical IoT Systems and Beyond

A new physics-inspired framework for cyber-physical IoT systems uses energy-based representations to provide structural attribution and explainable AI without needing directed causal graphs.

AI/ML arXiv cs.AI

CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programming

CSTutorBench is a new benchmark for evaluating Small Language Models (SLMs) acting as computer science tutors for block-based programming in VEX VR.

AI/ML arXiv cs.AI

Foundation Models for Automatic CAD Generation

The study introduces LLMForge, a multi-model text-to-CAD framework that uses JSON-schema validation and VLM-based critique to generate parametric 3D mechanical designs.

AI/ML arXiv cs.AI

Narrative World Model: Narratology-Grounded Writer Memory for Long-Form Fiction

The Narrative World Model (NWM) is a writer-memory system using a narratology-grounded temporal-state graph to improve long-form fiction writing by AI agents.

AI/ML arXiv cs.AI

FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents

FirstResearch introduces a structured 'Research Question Certificate' to make LLM-generated scientific discovery questions more auditable and grounded in first principles.

AI/ML arXiv cs.AI

Memory in the Loop: In-Process Retrieval as ExtendedWorking Memory for Language Agents

Research explores moving memory 'in-process' for language agents to reduce retrieval latency from milliseconds to microseconds, treating it as extended working memory.

AI/ML arXiv cs.AI

Akashic: A Low-Overhead LLM Inference Service with MemAttention

Akashic is a low-overhead LLM inference service using MemAttention to organize context into semantic chunks, reducing prefill costs and improving throughput.

AI/ML arXiv cs.AI

ArtisanCAD: An Industrial-Level CAD Agent with Expert-Grounded Knowledge Distillation

ArtisanCAD is an industrial-level CAD agent that uses CAD-IR (Intermediate Representation) to distill expert knowledge into reusable skills for production-ready B-Rep models.

Open Source Hacker News

LineageOS Statistics

A Hacker News discussion regarding statistics for LineageOS, a popular open-source Android distribution.

AI/ML arXiv cs.AI

Seduced by the Narrative: Assessing Rule Adherence in Semi-Open Textual Sandboxes

Researchers introduce CoC-Seduce, a benchmark to evaluate how LLMs can be manipulated into ignoring rules through 'Rhetorical Injection' in semi-open textual environments.

AI/ML arXiv cs.AI

Training Hybrid Block Diffusion Language Models with Partial Bidirectionality

The BDLM Mamba-attention hybrid architecture is proposed to increase throughput for long-context generation in diffusion language models by enabling exact caching.

AI/ML arXiv cs.AI

SovereignNegotiation-Bench: Evaluating User-Owned Personal Agents In Delegated Bargaining Under Privacy, Consent, Evidence, And Institutional Pressure

SovereignNegotiation-Bench is introduced to evaluate user-owned personal agents on their ability to negotiate while protecting user privacy and consent.

Cybersecurity arXiv cs.AI

Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models

Study reveals a critical vulnerability in cloud-edge LVLM inference where manipulating just 10% of vision tokens can drastically reduce model accuracy.

Cybersecurity arXiv cs.AI

JavaVulBench: A Java Vulnerability Benchmark with Realistic Splits, a Unified Multi-Backend Harness, and a Leakage-Aware Evaluation Mode

JavaVulBench is released, providing a comprehensive dataset and harness for evaluating Java vulnerability detection across multiple LLM backends.

AI/ML arXiv cs.AI

Differential Amplifier-Inspired AmpAttention for Multi-View Robotic Manipulation

AmpAttention is a new attention mechanism inspired by analog circuits to reduce attention drift and improve perception in multi-view robotic manipulation.

Cybersecurity arXiv cs.AI

Determinants and Limits of LLM Security-Tool Orchestration: A Study with HexStrike-AI

Research using the HexStrikeAI orchestrator evaluates the capabilities and limits of LLM agents driving security tool suites over the Model Context Protocol.