Open Source Hacker News

Show HN: Ex-Deloitte auditor open-sourced the whole SOC 2 method for your AI

An ex-Deloitte auditor has open-sourced the SOC 2 compliance method for AI systems, providing a framework for auditing AI-specific risks.

AI/ML Hacker News

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

A systematic study exploring the phenomenon of benchmark saturation in AI, where benchmarks no longer effectively differentiate model performance.

Hardware/Chips Hacker News

Looking inside a 1970s PROM chip that stores data in microscopic fuses (2019)

A detailed look at the inner workings of a 1970s PROM chip, demonstrating how data is stored using microscopic fuses.

Cybersecurity TechCrunch

Hackers steal over $130 million by exploiting bug in offline hardware wallets

A vulnerability in Coldcard hardware wallets has led to over $130 million in cryptocurrency thefts.

Other The Verge

T-Mobile’s $0-down financing plan bundles taxes and fees

T-Mobile introduces a new financing plan that allows customers to pay taxes and fees over 36 months.

AI/ML arXiv cs.AI

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents

Introduction of Argos, a reward agent that uses an adaptive verifier to improve multimodal reinforcement learning for AI agents.

AI/ML arXiv cs.AI

M3MAD-Bench: Multi-Dimensional Evaluation of Multi-Agent Debate Across Domains and Modalities

M3MAD-Bench is a new unified benchmark for evaluating multi-agent debate across diverse domains and modalities.

AI/ML arXiv cs.AI

RAPiD: Reward-Guided Consistency Distillation of Diffusion Planners for Real-Time Autonomous Driving

RAPiD is a reward-guided consistency distillation framework that significantly reduces inference latency for diffusion-based autonomous driving planners.

AI/ML arXiv cs.AI

Shaping Scientific Explanations to Expert Perspectives with Persona-Conditioned Reinforcement Learning

A framework for persona-conditioned reinforcement learning that adapts scientific explanations to different expert perspectives.

AI/ML arXiv cs.AI

What Makes a Sale? Simulating End-to-End Seller--Buyer Retail Dynamics with LLM Agents

RetailSim is an end-to-end retail simulation framework using LLM agents to model buyer-seller interactions and retail strategies.

AI/ML TechCrunch

Spotify expands AI remix and covers project with Merlin partnership

Spotify is partnering with Merlin and Universal Music Group to launch a paid AI-powered tool for creating music remixes and covers while ensuring artist compensation.

Other TechCrunch

Texas halts new data centers as governor calls for audits

Texas Governor Greg Abbott has paused new data center development pending the completion of an audit.

Tech Business/VC TechCrunch

Walmart completes its acquisition of TV advertising company Vibe.co

Walmart has completed the acquisition of Vibe.co, integrating the TV advertising company into its Walmart Connect platform.

Hardware/Chips The Verge

Samsung’s HDR10 Plus Advanced is launching this month on Prime Video

Samsung is launching HDR10 Plus Advanced on Prime Video, offering improved HDR metadata and tone mapping for its 2026 TV lineup.

Other The Verge

Texas says data centers must pass an audit before connecting to the grid

Texas is requiring new data centers to pass an audit focusing on grid stability, water consumption, and community impact before connecting to the energy grid.

AI/ML arXiv cs.AI

The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations

A research paper proposes the Socratic Test, an automated conversational assessment framework using Dynamic Assessment and Bloom's Taxonomy to map student cognitive boundaries.

AI/ML arXiv cs.AI

SATViz: Real-Time Visualization of Clausal Proofs

The paper introduces SATViz, a tool for the real-time visualization of clausal proofs in SAT instances using variable interaction graphs.

AI/ML arXiv cs.AI

Combining Large Language Models and Symbolic Reasoning for Multi-Robot Temporal Planning through Explainable Knowledge Bases

The PLANTOR framework combines LLMs and symbolic reasoning to generate multi-robot task plans through explainable knowledge bases.

AI/ML arXiv cs.AI

Shall We Play a Game? Language Models for Open-ended Wargames

A scoping review examines the role of LLMs in open-ended wargames, highlighting the need for models to act as both player agents and world adjudicators.

AI/ML arXiv cs.AI

Embedded Universal Predictive Intelligence: a coherent framework for multi-agent learning

The paper introduces a mathematical framework for embedded agency in multi-agent learning, extending AIXI theory to enable infinite-order theory of mind.