AI/ML arXiv cs.AI

GPU-Parallel Linearization Error Bounds for Real-Time Robust Optimal Control of Nonlinear and Neural Network Dynamics

Develops GPUSLS-LEO, a GPU-parallel solver providing tight linearization error bounds for robust optimal control of nonlinear and neural network dynamics in real-time.

AI/ML arXiv cs.AI

Distill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation

Introduces Distill to Detect (D2D), a method to expose stealthy biases in LLMs by distilling distributional shifts into cartridge prefix adapters.

Other Hacker News

CarPlay Is Additive

A discussion thread on Hacker News regarding the additive nature of CarPlay integration.

Software Engineering Hacker News

Show HN: Gitstock–Transform you GitHub commit history into K-line and animations

Gitstock is a tool that transforms GitHub commit history into K-line charts and animations for visualization.

Hardware/Chips The Verge

Sony’s PlayStation disc factory is already being repurposed

Sony is repurposing its PlayStation disc manufacturing facility in Austria to produce optical microlenses as physical disc demand declines.

AI/ML arXiv cs.AI

MemSyco-Bench: Benchmarking Sycophancy in Agent Memory

Researchers introduce MemSyco-Bench, a benchmark designed to evaluate sycophancy in LLM-based agents' memory, where agents over-align with users at the cost of accuracy.

AI/ML arXiv cs.AI

Staleness-Learning Rate Scaling Laws for Asynchronous RLHF

This paper analyzes the effect of stale rollouts in asynchronous RLHF (GRPO) and derives scaling laws to maintain stability during policy optimization.

AI/ML arXiv cs.AI

LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models

LongVQUBench is presented as a new benchmark for evaluating the long-term video quality understanding of vision-language models across various temporal scopes.

Software Engineering arXiv cs.AI

Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering

A case study exploring how software engineering shifts from implementation effort to the governance of abundant AI-generated code.

AI/ML arXiv cs.AI

CausalMix: Data Mixture as Causal Inference for Language Model Training

CausalMix treats LLM data mixture optimization as a causal inference problem, allowing for better generalization across different data pool sizes and model scales.

AI/ML arXiv cs.AI

FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy Improvement

The Failure-Aware Retry (FAR) framework allows robots to learn from previous test-time failures to autonomously recover and improve policies.

AI/ML arXiv cs.AI

Towards Developing a Multimodal Chat Assistant for University Stakeholders: RAG-based Approach

A RAG-based multimodal chatbot designed for university stakeholders to provide reliable institutional information using quantized inference for constrained hardware.

Tech Business/VC Hacker News

"An AI Job Apocalypse?" – Goldman Sachs Report [pdf]

A report from Goldman Sachs discusses the potential for AI to cause widespread job displacement.

Other Hacker News

GitHub is proud to announce that you can now obtain your public repo on CD-ROM

GitHub introduces a novelty feature allowing users to obtain their public repositories on CD-ROM.

AI/ML Hacker News

Right to Local Intelligence

An exploration of the 'Right to Local Intelligence', advocating for the ability to run AI models locally.

Tech Business/VC TechCrunch

Last chance to apply — Startup Battlefield Australia applications close July 6

Call for applications for Startup Battlefield Australia, closing July 6.

AI/ML VentureBeat

Enterprises lost Claude Fable 5 for a few weeks. New data shows two-thirds had already built their hedge

Enterprises are increasingly hedging their AI strategies by combining closed frontier models with open-weight models to avoid vendor lock-in and outages.

AI/ML arXiv cs.AI

Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads

Researchers introduce LOCOS, a write-aware detector that identifies attention heads responsible for non-literal retrieval in LLMs.

AI/ML arXiv cs.AI

Reading Order Inference for Complex Document Layouts

A training-free, graph-based framework for inferring reading order in complex historical document layouts.

AI/ML arXiv cs.AI

Behavior-Adaptive Conversational Agents: Toward a Fluid Personality Framework

Proposed Fluid Personality Framework to adapt an AI agent's persona and personality expression based on context and user goals.