AI/ML arXiv cs.AI

L-MARS: Legal Multi-Agent System with Agentic Search and Citation-Faithfulness Audit

L-MARS is an open multi-agent legal QA system featuring agentic search and a citation-faithfulness audit mechanism.

Other Hacker News

Show HN: A zoomable timeline of 4M Wikipedia events

A showcase of a zoomable timeline visualization featuring 4 million Wikipedia events.

Software Engineering Hacker News

MoonBASIC: A modern BASIC for building 2D and 3D games

MoonBASIC is introduced as a modern BASIC language specifically designed for 2D and 3D game development.

Hardware/Chips The Verge

Pebble founder Eric Migicovsky says his 30-day warranty is all about trust

Pebble's founder discusses the trust-based approach and 30-day warranty for their revived e-paper smartwatches.

AI/ML VentureBeat

Agents think in milliseconds, legacy infrastructure doesn't. LinkedIn, Walmart and Zendesk shared how they closed the gap at VB Transform 2026

Leaders from LinkedIn, Walmart, and Zendesk discuss overcoming infrastructure bottlenecks (Kubernetes latency, data pipelines) when scaling AI agents into production.

AI/ML arXiv cs.AI

Mask-Aware Policy Gradients for Diffusion Language Models

Researchers propose Mask-Aware Policy Gradients for Diffusion Language Models to improve mathematical reasoning and coding benchmarks.

AI/ML arXiv cs.AI

MM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue Localization

The MM-IssueLoc benchmark is introduced to evaluate how well multimodal AI can localize issues in software repositories using visual evidence like screenshots.

AI/ML arXiv cs.AI

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Symbal is a new tool and benchmark for detecting systematic misalignments in captions generated by multimodal LLMs.

AI/ML arXiv cs.AI

In-Place Tokenizer Expansion for Pre-trained LLMs

A new 'in-place tokenizer expansion' recipe allows pre-trained LLMs to efficiently add new languages, significantly increasing decode speed for those languages.

Cybersecurity arXiv cs.AI

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

A study on cost-aware evaluation of security agents, arguing that economic efficiency and tool use are as important as success rates in offensive and defensive security.

AI/ML arXiv cs.AI

SceneBind: Binding What and Where Across Vision, Audio and Language

SceneBind is introduced as an omni-modal representation for joint semantic and 3D spatial understanding across vision, audio, and language.

Software Engineering Hacker News

Learning a few things about running SQLite

A community discussion on the nuances and practicalities of running SQLite in production environments.

AI/ML Hacker News

Homomorphically encrypted CIFAR-10 inference in 200ms

A technical demonstration of achieving high-speed inference on the CIFAR-10 dataset using homomorphic encryption.

AI/ML arXiv cs.AI

Parameter-efficient Prompt Tuning of Vision Foundation Model With Adaptive Focal Loss for Interpretable MCI Screening

Proposes a parameter-efficient framework for MCI screening using frozen DINOv2-Small and adaptive focal loss to improve interpretability and accuracy.

AI/ML arXiv cs.AI

ANet Patu-1: The Value of Connection in the Agent Network

Introduces ANet Patu-1, a self-organizing consensus protocol for AI agents that optimizes collaboration and scales effectively regardless of model strength.

AI/ML arXiv cs.AI

Towards Hierarchical Structure Understanding of Newspaper Images

Explores hierarchical structure understanding of newspaper images using both a modular pipeline and a new end-to-end transformer architecture called Tiramisu.

AI/ML arXiv cs.AI

Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents

Presents a multi-agent framework using SFT, DPO, and RAG to simulate and audit political coalition formation with LLM agents.

Hardware/Chips arXiv cs.AI

NIFA: Nonlinear IMC enhanced FPGA for efficient ML inference

Introduces NIFA, an FPGA architecture using ADC-free IMC blocks to significantly improve energy and area efficiency for ML inference, particularly for Transformers.

AI/ML arXiv cs.AI

Scaling Behavior Foundation Model for Humanoid Robots

Develops a Behavior Foundation Model for humanoid robots using a 'Humanoid Transformer' architecture to improve whole-body coordination and task generalization.

AI/ML arXiv cs.AI

T^2MLR: Transformer with Temporal Middle-Layer Recurrence

Introduces T2MLR, a transformer architecture that uses temporal middle-layer recurrence to allow intermediate reasoning states to persist across decoding steps.