AI/ML arXiv cs.AI

MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering

Introduction of MultAttnAttrib, a training-free method for multimodal attribution in long document QA, and its companion benchmark MultAttrEval.

AI/ML arXiv cs.AI

Risk Architecture for AI-Native Engineering Teams: An Organizational Framework for Agentic System Governance

A proposed organizational framework for governing agentic AI systems, addressing the gap between high-level policy and low-level threat taxonomies.

AI/ML arXiv cs.AI

IsoSci: A Benchmark of Isomorphic Cross-Domain Science Problems for Evaluating Reasoning versus Knowledge Retrieval in LLMs

Presentation of ISOSCI, a benchmark designed to separate reasoning ability from knowledge retrieval in LLMs using isomorphic cross-domain science problems.

AI/ML arXiv cs.AI

On the Utility and Factual Reliability of Pruned Mixture-of-Experts Models in the Biomedical Domain

Investigation into how structured expert pruning in Mixture-of-Experts (MoE) models affects factual reliability, specifically in the biomedical domain.

AI/ML arXiv cs.AI

Token Geometry

Introduction of Ember, a lightweight optimizer for embedding and LM-head matrices that significantly reduces VRAM usage compared to Adam.

AI/ML arXiv cs.AI

Grounded Optimization: A Layered Engineering Framework for Reducing LLM Hallucination in Automated Personal Document Rewriting

A five-layer engineering framework called Grounded Optimization designed to reduce hallucinations in LLM-based automated document rewriting.

Other arXiv cs.AI

Fully Unsupervised Detection of Physical Contacts on Subsea Cables via State-of-Polarization Monitoring

Development of an unsupervised detector using State-of-Polarization monitoring to identify physical contacts on subsea cables.

AI/ML arXiv cs.AI

Don't Let Gains FADE: Breaking Down Policy Gradient Weights in RL

Introduction of FADE, a self-adapting advantage function for RL post-training that improves reasoning while maintaining diversity and training stability.

Other Hacker News

Instead of banning AI, I made a classroom contract with my students

A discussion on establishing a classroom contract to integrate AI usage rather than banning it.

Other Hacker News

Closer to Rude Than Snide: An Interview with Leo Robson

An interview with Leo Robson discussing various topics.

Open Source Hacker News

Show HN: Bramble – Local-first password manager

Introduction of Bramble, a local-first password manager designed for privacy and local control.

AI/ML arXiv cs.AI

AI-enabled gravitational-waves searches for binary neutron stars at optimal sensitivity

Researchers developed Aframe, an AI-enabled search tool that detects binary neutron star mergers with sensitivity comparable to matched-filter pipelines but lower cost.

AI/ML arXiv cs.AI

How Should Transformers Encode Numeric Values in Electronic Health Records?

A study comparing numeric value encoding strategies for transformers in Electronic Health Records, suggesting hybrid token-based approaches as a robust default.

AI/ML arXiv cs.AI

Rethinking Generic Object Tracking Toward Human-Level Perceptual Intelligence

A dissertation proposing methods to improve generic object tracking (GOT) by enhancing target discrimination and geometric reasoning to mimic human perception.

AI/ML arXiv cs.AI

NeuroBridge: Bridging Multi-Task MRI Knowledge for Neurodegenerative Disease Diagnosis

NeuroBridge is a multi-task MRI framework that integrates self-supervised pretraining for improved diagnosis of Alzheimer's and other neurodegenerative diseases.

AI/ML arXiv cs.AI

Spin-Weighted Spherical Harmonics Enable Complete and Scalable $\mathrm{E}(3)$-Equivariant Networks

Introduction of SpinGTP, a scalable approach to E(3)-equivariant networks using Spin-Weighted Spherical Harmonics for better 3D atomistic system modeling.

Software Engineering arXiv cs.AI

GPUAlert: A Zero-Instrumentation Process-Boundary Monitor for Diagnosing GPU Training-Job Failures

GPUAlert is a zero-instrumentation wrapper that monitors GPU training jobs and emails structured failure notifications and logs without modifying the training script.

Tech Business/VC arXiv cs.AI

Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI

A study of Microsoft's rollout of Claude Code and GitHub Copilot CLI, finding that adopters merged 24% more pull requests.

Tech Business/VC Hacker News

Costco Is the Anti-Amazon

A discussion on how Costco's business model differs from Amazon's, focusing on the 'anti-Amazon' approach to retail.

Open Source Hacker News

Show HN: Mcpsnoop – Wireshark for MCP (transparent proxy and live TUI)

Mcpsnoop is a transparent proxy and live TUI designed as a 'Wireshark for MCP', providing visibility into Model Context Protocol traffic.