All Articles
17547 articles total
Controlling Tool Use with Heading-Specific Activation Steering
Research on using activation steering vectors to control and suppress unnecessary tool invocation in tool-augmented LLMs.
From Passive Retrieval to Active Memory Navigation: Learning to Use Memory as a Structured Action Space
Introduction of NapMem, a framework that treats long-term user memory as a structured action space for conversational agents.
Show HN: Neil the Seal Game
A new game called 'Neil the Seal' has been showcased on Hacker News.
Prompt-to-Paper: Agentic AI System for Bioinformatics
Prompt-to-Paper is a multi-agent AI framework designed to automate the generation of bioinformatics manuscripts by grounding claims in verifiable literature and executing real computational experiments.
From Graphs to Gradients: Physics-Inspired Structural Attribution for Cyber-Physical IoT Systems and Beyond
A new physics-inspired framework for cyber-physical IoT systems uses energy-based representations to provide structural attribution and explainable AI without needing directed causal graphs.
CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programming
CSTutorBench is a new benchmark for evaluating Small Language Models (SLMs) acting as computer science tutors for block-based programming in VEX VR.
Foundation Models for Automatic CAD Generation
The study introduces LLMForge, a multi-model text-to-CAD framework that uses JSON-schema validation and VLM-based critique to generate parametric 3D mechanical designs.
Narrative World Model: Narratology-Grounded Writer Memory for Long-Form Fiction
The Narrative World Model (NWM) is a writer-memory system using a narratology-grounded temporal-state graph to improve long-form fiction writing by AI agents.
FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents
FirstResearch introduces a structured 'Research Question Certificate' to make LLM-generated scientific discovery questions more auditable and grounded in first principles.
Memory in the Loop: In-Process Retrieval as ExtendedWorking Memory for Language Agents
Research explores moving memory 'in-process' for language agents to reduce retrieval latency from milliseconds to microseconds, treating it as extended working memory.
Akashic: A Low-Overhead LLM Inference Service with MemAttention
Akashic is a low-overhead LLM inference service using MemAttention to organize context into semantic chunks, reducing prefill costs and improving throughput.
ArtisanCAD: An Industrial-Level CAD Agent with Expert-Grounded Knowledge Distillation
ArtisanCAD is an industrial-level CAD agent that uses CAD-IR (Intermediate Representation) to distill expert knowledge into reusable skills for production-ready B-Rep models.
LineageOS Statistics
A Hacker News discussion regarding statistics for LineageOS, a popular open-source Android distribution.
Seduced by the Narrative: Assessing Rule Adherence in Semi-Open Textual Sandboxes
Researchers introduce CoC-Seduce, a benchmark to evaluate how LLMs can be manipulated into ignoring rules through 'Rhetorical Injection' in semi-open textual environments.
Training Hybrid Block Diffusion Language Models with Partial Bidirectionality
The BDLM Mamba-attention hybrid architecture is proposed to increase throughput for long-context generation in diffusion language models by enabling exact caching.
SovereignNegotiation-Bench: Evaluating User-Owned Personal Agents In Delegated Bargaining Under Privacy, Consent, Evidence, And Institutional Pressure
SovereignNegotiation-Bench is introduced to evaluate user-owned personal agents on their ability to negotiate while protecting user privacy and consent.
Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models
Study reveals a critical vulnerability in cloud-edge LVLM inference where manipulating just 10% of vision tokens can drastically reduce model accuracy.
JavaVulBench: A Java Vulnerability Benchmark with Realistic Splits, a Unified Multi-Backend Harness, and a Leakage-Aware Evaluation Mode
JavaVulBench is released, providing a comprehensive dataset and harness for evaluating Java vulnerability detection across multiple LLM backends.
Differential Amplifier-Inspired AmpAttention for Multi-View Robotic Manipulation
AmpAttention is a new attention mechanism inspired by analog circuits to reduce attention drift and improve perception in multi-view robotic manipulation.
Determinants and Limits of LLM Security-Tool Orchestration: A Study with HexStrike-AI
Research using the HexStrikeAI orchestrator evaluates the capabilities and limits of LLM agents driving security tool suites over the Model Context Protocol.