AI/ML arXiv cs.AI

JL1-CC&QA: Extending the JL1-CD Benchmark with Change Captioning and Question Answering

The JL1-CC&QA benchmark extends remote sensing change detection with change captioning and question answering to provide semantic understanding of land-cover changes.

Other arXiv cs.AI

A Technical Typology of AI Systems in Public Administration

The paper proposes a technical typology to categorize AI systems in public administration to improve accountability and procedural justice.

AI/ML arXiv cs.AI

CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield

CHERRY introduces compute-efficient language model training via selective supervision, depth compression with recurrent recovery, and a mixture of efficient experts.

Software Engineering Hacker News

Avoiding Fallback in Distributed Systems

A discussion on strategies to avoid fallback mechanisms in distributed systems to maintain system stability and predictability.

Software Engineering Hacker News

Show HN: Unobin compiles Infrastructure as Code to one binary

Unobin is a tool that compiles Infrastructure as Code (IaC) into a single binary, simplifying deployment and distribution.

AI/ML arXiv cs.AI

ECHO: Prune to act, trace to learn with selective turn memory in agentic RL

Introducing ECHO, a selective turn-memory framework for agentic RL that prevents history collapse and enables traceable learning in long-horizon language agents.

AI/ML arXiv cs.AI

Improving Certified Robustness via Adversarial Distillation

AD-CERT is a certified training objective that uses adversarial distillation with IBP to improve the robustness of neural networks against adversarial perturbations.

AI/ML arXiv cs.AI

Sparsity-Inducing Divergence Losses for Biometric Verification

Q-Margin introduces a novel alpha-divergence loss function for biometric verification, improving performance at low False Acceptance Rates and enabling memory-efficient training.

AI/ML arXiv cs.AI

WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models

WorldRoamBench is an open-world benchmark for evaluating the long-horizon stability, physics, and memory of interactive world models.

AI/ML arXiv cs.AI

Histogram-constrained Image Generation

Histogram-constrained Image Generation (HIG) is a novel control mechanism for diffusion models that uses optimal transport to enforce user-specified distributional constraints.

AI/ML arXiv cs.AI

When to Truncate a Feature Ranking: A Residual-Overlap Stopping Rule for Subset Selection

A paper presenting a risk-calibrated stopping rule for subset selection in supervised feature rankings, reducing high-dimensional datasets to a few dozen variables without losing predictive performance.

AI/ML arXiv cs.AI

ShopX: A Foundation Model for Intent-to-Item Fulfillment in Agentic Shopping

ShopX is a foundation model for agentic shopping that unifies intent understanding and item-space fulfillment using semantic IDs (SIDs).

AI/ML arXiv cs.AI

RCT: A Robot-Collected Touch-Vision-Language Dataset for Tactile Generalization

The RCT dataset is a robot-collected touch-vision-language dataset designed to improve tactile generalization in robotic manipulation of open-world objects.

Other Hacker News

Fable 5 update: Still willing to cybercrime

A Hacker News thread discussing a 'Fable 5' update regarding cybercrime.

AI/ML arXiv cs.AI

Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment

Research on how optimizer choice in LLM fine-tuning impacts emergent misalignment, finding that Muon preserves alignment better than Adam or Lion.

Cybersecurity arXiv cs.AI

Comparative Analysis of Machine Learning based Intrusion Detection in Realistic IoT Networks

A comparative study of ML algorithms for intrusion detection in IoT networks, identifying Random Forest as the top performer using the Gotham2025 dataset.

AI/ML arXiv cs.AI

Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning

Introduction of ViToS, a dual-stream RL framework that improves medical multimodal reasoning by pruning visual tokens to increase inference speed and accuracy.

AI/ML arXiv cs.AI

Preserve the Hard, Regenerate the Rest: Uncertainty-Guided Synthetic Training Data Augmentation with Diffusion Models

A new synthetic data augmentation strategy using diffusion models and predictive entropy to improve semantic segmentation in complex datasets.

AI/ML arXiv cs.AI

Learning Structurally Consistent Representations for Multi-View Radar Semantic Segmentation

A framework for multi-view radar semantic segmentation using learnable hypergraphs and Unbalanced Optimal Transport for structurally consistent representations.

AI/ML arXiv cs.AI

Automating Cause-Effect Specification with Knowledge Graphs and Large Language Models

A semantic-AI framework combining knowledge graphs and LLMs to automate the generation of cause-and-effect specifications for process control safety.