All Articles
16580 articles total
Specifying the Delegated-Autonomy Boundary: Requirements Engineering for Agentic AI
This paper introduces a requirements engineering framework for agentic AI, focusing on the 'delegated-autonomy boundary' via Agency Justification Records and Agentic Delegation Policies.
A RFID Based Campus Wide Payment System
A project implementing a campus-wide cashless payment system using RFID and Raspberry Pi.
A Large-Scale Measurement of AI Bill of Materials Completeness in Hugging Face Models
A study on the completeness of AI Bills of Materials (AIBOMs) for models on Hugging Face, finding significant gaps in AI-specific documentation.
Distilled Reinforcement Learning for LLM Post-training
Distilled RL integrates teacher supervision into the RL objective to improve the transfer of knowledge during LLM post-training.
Asynchronous Multimodal Diffusion Policy Composition via Latency-Aware Guidance Fusion
LAG-Fusion is a latency-aware guidance fusion framework for asynchronous multimodal diffusion policies in robotic imitation learning.
It's a shame what's happened to radio
A Hacker News discussion reflecting on the decline of traditional radio broadcasting.
Where Does Agent Reliability Come From? A Cross-Benchmark Decomposition of Verification Loops, Specialist Models, and Scaffolding in a Production Enterprise Agent
Research on increasing enterprise agent reliability using verification loops and specialized post-trained models to reduce common failure modes.
Solver-Hard Is Not Model-Hard: A Hardness-Controlled Diagnostic for LLM Constraint Reasoning
A study exploring the gap between solver hardness and model hardness for LLM constraint reasoning, revealing model sensitivity to problem surface.
EvoGUI: An Evolution-Aware Benchmark for GUI State-Transition Understanding
Introduction of EvoGUI, a diagnostic framework for evaluating how GUI agents understand state transitions in user interfaces.
ThAME: 3D Memory-Enabled Heterogeneous Accelerator for LLM Mixture of Experts
Proposal of ThAME, a 3D heterogeneous multi-chiplet architecture designed to optimize MoE LLM inference by reducing memory bandwidth and routing bottlenecks.
ALLUDE: A Unified Evaluation System for Configurable Attacks in Differentiable Environments
Presentation of ALLUDE, an open-source unified evaluation system for configuring and testing adversarial attacks against vision models in differentiable environments.
DepthART: Scaling Foundation Monocular Depth to Tiny Models
Development of DepthART, a compact monocular depth estimation model designed for efficient on-device deployment on tiny hardware like Jetson Nano.
Auto Research for Materials: Auditable AI-Scientist Workflows with Held-Out Transfer
Research on using closed-loop AI agents to discover reusable and transferable modeling changes for materials science predictions.
Teach it to stop, not just to click
An analysis of the variance and reliability of computer-use RL agents, providing a library for better k-seed reporting in research.
Noise-Robust Box-Supervised Infrared Small Target Detection via Physics-Inspired Soft Label Optimization
Introduction of HALO, a physics-inspired soft label optimization method for noise-robust small target detection in infrared imagery.
Juggling for Blind People
An article about juggling for blind people, likely focusing on accessibility or human interest.
A flaky test exposed a Redis client use-after-free
A report on how a flaky test identified a use-after-free vulnerability in a Redis client.
Show HN: An MCP server that turns async-work practices into tools
A new Model Context Protocol (MCP) server that converts asynchronous work practices into usable tools.
How an AI Anime Is Created
An exploration of the technical process behind creating AI-generated anime.
Ten Steps Towards Happiness
A piece offering ten steps toward achieving happiness, which is non-technical content.