Hardware/Chips Hacker News

Show HN: Continuous Nvidia CUDA PC Sampling Profiler

A new continuous sampling profiler specifically designed for Nvidia CUDA GPUs to assist developers in identifying performance bottlenecks.

AI/ML arXiv cs.AI

UltraQuant: 4-bit KV Caching for Context-Heavy Agents

UltraQuant introduces a 4-bit KV caching mechanism for context-heavy AI agents, significantly improving time-to-first-token and output throughput on AMD GPUs.

AI/ML arXiv cs.AI

Optimal Order of Multi-Agent and General Many-Body Systems

A research paper proposing a general framework for analyzing multi-agent systems, focusing on the balance between collective productivity and systemic fragility.

AI/ML arXiv cs.AI

Contagion Networks: Evaluator Bias Propagation in Multi-Agent LLM Systems

Introduces Contagion Networks, a framework for measuring how evaluation biases propagate through multi-agent LLM systems, and suggests using committees to mitigate this.

Cybersecurity arXiv cs.AI

Calibration Without Comprehension: Diagnosing the Limits of Fine-Tuning LLMs for Vulnerability Detection in Systems Software

Analysis of LLMs' ability to detect vulnerabilities in systems software, concluding that fine-tuning often results in 'calibration without comprehension' rather than actual security reasoning.

AI/ML arXiv cs.AI

FreeStyle: Free Control of Style-Content Dual-Reference Generation from Community LoRA Mining

FreeStyle is a dual-reference generation framework that uses community LoRA mining to synthesize images while preserving content structure and separating style.

AI/ML arXiv cs.AI

Efficient and Sound Probabilistic Verification for AI Agents

A new framework for sound and efficient probabilistic verification of AI agents using distributionally robust optimization to ensure security policy bounds.

Cybersecurity arXiv cs.AI

Sovereign Execution Brokers: Enforcing Certificate-Bound Authority in Agentic Control Planes

Introduces the Sovereign Execution Broker (SEB), a runtime enforcement boundary that separates proposal, admission, and execution to secure agentic control planes.

Tech Business/VC Hacker News

John Jumper leaves Google to join Anthropic

John Jumper, a key figure in AlphaFold, leaves Google to join Anthropic.

Other Hacker News

Court Records Should Be Free

A discussion on the necessity of making court records freely available to the public.

AI/ML arXiv cs.AI

Robust $Q$-learning for mean-field control under Wasserstein uncertainty in common noise

Researchers present a robust Q-learning algorithm for mean-field control problems under Wasserstein uncertainty in common noise.

AI/ML arXiv cs.AI

AutoPass: Evidence-Guided LLM Agents for Compiler Performance Tuning

AutoPass is a multi-agent LLM framework that optimizes compiler performance tuning by analyzing internal compiler states and runtime feedback.

AI/ML arXiv cs.AI

CRAX: Fast Safe Reinforcement Learning Benchmarking

CRAX is a fast, safe reinforcement learning benchmark utilizing the MuJoCo XLA (MJX) physics engine for hardware-accelerated 3D dynamics.

AI/ML arXiv cs.AI

DataMagic: Transforming Tabular Data into Data Insight Video

DataMagic is an interactive system that uses a multi-agent architecture to transform raw tabular data into narrative data-insight videos.

Cybersecurity arXiv cs.AI

LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems

NRT-Bench is a multi-turn red-teaming benchmark for LLM agents operating safety-critical systems, simulated via a nuclear power plant control room.

Cybersecurity arXiv cs.AI

Multi-View Decompilation for LLM-Based Malware Classification

Research shows that using multiple decompiler views (e.g., Ghidra and RetDec) improves LLM-based malware classification and triage.

AI/ML arXiv cs.AI

Repurposing a Speech Classifier for Guided Diffusion-Based Speech Generation

A study on repurposing a frozen speech classifier as the backbone for diffusion-based speech generation to reduce memory and computational costs.

Cybersecurity arXiv cs.AI

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems

The paper introduces CMPE, a lightweight conversational misdirection method to defend agentic AI systems against automated jailbreak attacks.

Other Hacker News

Music generation using Algebra, presented as web MIDI shop

A project demonstrating music generation using algebraic principles, presented as an interactive web-based MIDI shop.

AI/ML arXiv cs.AI

FlowMaps: Modeling Long-Term Multimodal Object Dynamics with Flow Matching

Introduction of FlowMaps, a latent flow matching model that predicts future locations of dynamic objects in 3D spaces to improve robotic navigation.