AI/ML arXiv cs.AI

Rethinking Global Average Pooling: Your Classifier Is Secretly a Multi-Instance Learner

A re-evaluation of Global Average Pooling in image classifiers, proposing a Multiple-Instance Learning interpretation to recover spatial class evidence.

AI/ML arXiv cs.AI

Regional Climate Model Emulation with Diffusion Approaches: What is the Added Value of Generative Machine Learning?

An evaluation of diffusion-based generative models for regional climate model emulation, introducing the ParamDiffusion framework.

AI/ML arXiv cs.AI

SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model

SIMMER, a new benchmark for detecting latent failures in LLM-based executable planning for autonomous agents in kitchen environments.

AI/ML arXiv cs.AI

CARE: Controlling LLM-Generated Policies through Auditable Review of Evidence in Scientific Experimentation

CARE, an auditable controller for high-throughput experimentation that uses LLMs to revise policies while maintaining a safe default optimizer.

AI/ML arXiv cs.AI

Sensitivity Shaping for Latent Modeling

A method for improving OOD detection in robotic generative dynamics models using support-conditioned control-sensitivity regularization.

Other Hacker News

Swedish parliament abolishes permanent residence visas for migrants

Sweden's parliament has abolished permanent residence visas for migrants, marking a significant shift in immigration policy.

Other Hacker News

Why I Email Complete Strangers

An essay exploring the philosophy and practical benefits of emailing strangers to build professional and personal networks.

Open Source Hacker News

Reviving an abandoned open-source project: 6 years of Atomic Calendar Revive

A retrospective on the six-year journey of reviving and maintaining an abandoned open-source calendar project.

Other Hacker News

Commander Keen Games (free book)

A free book providing technical or historical insights into the Commander Keen games.

AI/ML arXiv cs.AI

CADET: Physics-Grounded Causal Auditing and Training-Free Deconfounding of End-to-End Driving Planners

Introduction of CADET, a training-free framework to audit and repair causal confusion in end-to-end autonomous driving planners.

AI/ML arXiv cs.AI

tap: A File-Based Protocol for Heterogeneous LLM Agent Collaboration

Presentation of tap, a file-based protocol enabling heterogeneous LLM agents from different vendors to collaborate on shared codebases.

AI/ML arXiv cs.AI

MoDiCoL: A Modular Diagnostic Continual Learning Dataset for Robust Speech Recognition

MoDiCoL, a modular diagnostic dataset for continual learning in speech recognition, designed to improve robustness against real-world shifts.

AI/ML arXiv cs.AI

The Perceived Fragility of Explanations in Audio Models: Manipulation of Attribution with Unchanged Predictions

Research demonstrating that explanations in audio deepfake detection models can be manipulated via psychoacoustic perturbations without changing predictions.

AI/ML arXiv cs.AI

A Fixed-Point Neural Operator for Size- and Functional-Transferable Hamiltonian Prediction

HamEvo, a neural operator for predicting the Kohn-Sham Hamiltonian, offering significantly faster inference than conventional DFT.

AI/ML arXiv cs.AI

Fodor and Pylyshyn's Systematicity Challenge Still Stands

A critical analysis arguing that the challenge of systematicity in human cognition remains unsolved by current neural network architectures.

Other Hacker News

Techno-libertarians are flocking to the Caribbean

A discussion regarding techno-libertarians moving to the Caribbean to establish sovereign zones or escape government regulation.

Other Hacker News

The Dead Economy Theory

An exploration of the 'Dead Economy Theory', likely discussing economic stagnation or systemic failure through a theoretical lens.

Software Engineering Hacker News

What job interviews taught me about Kubernetes

A personal reflection on the nature of Kubernetes job interviews and what they reveal about the industry's perception of the tool.

Tech Business/VC TechCrunch

The US government’s Anthropic models ban was never about an AI jailbreak

An analysis of the US government's ban on specific Anthropic AI models, highlighting government interference in the AI industry.

AI/ML arXiv cs.AI

PLAIground: SLO-Driven Runtime Model Selection for Compound AI Systems in the Edge-Cloud-Space Continuum

Introduces PLAIground, a framework for SLO-driven runtime model selection in Compound AI systems across edge, cloud, and space environments.