All Articles
16860 articles total
L-MARS: Legal Multi-Agent System with Agentic Search and Citation-Faithfulness Audit
L-MARS is an open multi-agent legal QA system featuring agentic search and a citation-faithfulness audit mechanism.
Show HN: A zoomable timeline of 4M Wikipedia events
A showcase of a zoomable timeline visualization featuring 4 million Wikipedia events.
MoonBASIC: A modern BASIC for building 2D and 3D games
MoonBASIC is introduced as a modern BASIC language specifically designed for 2D and 3D game development.
Pebble founder Eric Migicovsky says his 30-day warranty is all about trust
Pebble's founder discusses the trust-based approach and 30-day warranty for their revived e-paper smartwatches.
Agents think in milliseconds, legacy infrastructure doesn't. LinkedIn, Walmart and Zendesk shared how they closed the gap at VB Transform 2026
Leaders from LinkedIn, Walmart, and Zendesk discuss overcoming infrastructure bottlenecks (Kubernetes latency, data pipelines) when scaling AI agents into production.
Mask-Aware Policy Gradients for Diffusion Language Models
Researchers propose Mask-Aware Policy Gradients for Diffusion Language Models to improve mathematical reasoning and coding benchmarks.
MM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue Localization
The MM-IssueLoc benchmark is introduced to evaluate how well multimodal AI can localize issues in software repositories using visual evidence like screenshots.
Symbal: Detecting Systematic Misalignments in Model-Generated Captions
Symbal is a new tool and benchmark for detecting systematic misalignments in captions generated by multimodal LLMs.
In-Place Tokenizer Expansion for Pre-trained LLMs
A new 'in-place tokenizer expansion' recipe allows pre-trained LLMs to efficiently add new languages, significantly increasing decode speed for those languages.
Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
A study on cost-aware evaluation of security agents, arguing that economic efficiency and tool use are as important as success rates in offensive and defensive security.
SceneBind: Binding What and Where Across Vision, Audio and Language
SceneBind is introduced as an omni-modal representation for joint semantic and 3D spatial understanding across vision, audio, and language.
Learning a few things about running SQLite
A community discussion on the nuances and practicalities of running SQLite in production environments.
Homomorphically encrypted CIFAR-10 inference in 200ms
A technical demonstration of achieving high-speed inference on the CIFAR-10 dataset using homomorphic encryption.
Parameter-efficient Prompt Tuning of Vision Foundation Model With Adaptive Focal Loss for Interpretable MCI Screening
Proposes a parameter-efficient framework for MCI screening using frozen DINOv2-Small and adaptive focal loss to improve interpretability and accuracy.
ANet Patu-1: The Value of Connection in the Agent Network
Introduces ANet Patu-1, a self-organizing consensus protocol for AI agents that optimizes collaboration and scales effectively regardless of model strength.
Towards Hierarchical Structure Understanding of Newspaper Images
Explores hierarchical structure understanding of newspaper images using both a modular pipeline and a new end-to-end transformer architecture called Tiramisu.
Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents
Presents a multi-agent framework using SFT, DPO, and RAG to simulate and audit political coalition formation with LLM agents.
NIFA: Nonlinear IMC enhanced FPGA for efficient ML inference
Introduces NIFA, an FPGA architecture using ADC-free IMC blocks to significantly improve energy and area efficiency for ML inference, particularly for Transformers.
Scaling Behavior Foundation Model for Humanoid Robots
Develops a Behavior Foundation Model for humanoid robots using a 'Humanoid Transformer' architecture to improve whole-body coordination and task generalization.
T^2MLR: Transformer with Temporal Middle-Layer Recurrence
Introduces T2MLR, a transformer architecture that uses temporal middle-layer recurrence to allow intermediate reasoning states to persist across decoding steps.