Open Source Hacker News

9front "This Was Supposed to Be Fun" Released

The release of 9front 'This Was Supposed to Be Fun', a new version of the community-driven Plan 9 fork.

AI/ML arXiv cs.AI

HERO: History-Enriched Rollout Training for Long-Horizon Autoregressive Neural Operators

Introduces HERO, a training method for neural operators to improve long-horizon accuracy and stability in predicting partial differential equations.

AI/ML arXiv cs.AI

Have I Seen You? Embedding Behavior Signals Synthetic Face Dataset Membership

Research demonstrating that synthetic face datasets can still leak membership information about the real faces used to train the generators.

AI/ML arXiv cs.AI

Implicit Machine Learning Force Fields Accelerate Molecular Dynamics Simulations

Introduces implicit machine learning force fields (I-MLFFs) to accelerate molecular dynamics simulations by reducing compute and memory footprints.

AI/ML arXiv cs.AI

Memory Provenance Laundering in LLM Agents: A Non-Amplification Firewall for Persistent Memory

Proposes the Provenance-Preserving Memory Firewall (PPMF) to prevent 'memory provenance laundering' in LLM agents with persistent memory.

AI/ML arXiv cs.AI

ActFovea: Runtime Safeguarding for VLA Policies via Spatiotemporal Visual-Action Consistency

Presents ActFovea, a safeguarding framework for Vision-Language-Action policies in robotics to mitigate runtime disturbances via spatiotemporal consistency.

AI/ML arXiv cs.AI

CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning

Introduces CLIFT, a method for closed-loop iterative fine-tuning of closed-weight robot foundation models using managed SFT APIs.

Cybersecurity Hacker News

Show HN: Nightcrawler – A local AI pentesting agent running on a smartphone

Nightcrawler is a local AI-powered pentesting agent designed to run on smartphones for mobile security auditing.

AI/ML arXiv cs.AI

Learning Lookahead Lemmas for Neural Network Verification

Researchers propose a lookahead-driven inprocessing framework to improve the efficiency and performance of neural network verifiers like Marabou and alpha-beta-CROWN.

AI/ML arXiv cs.AI

Autonomous Repair for Multi-Agent Systems via Monte-Carlo Tree Search

The MARS framework utilizes Monte-Carlo Tree Search to automate the repair of errors in multi-agent systems, introducing the StateMAS benchmark for evaluation.

AI/ML arXiv cs.AI

Benchmarking Frontier Large Language Models Against Official Crash Database Coding Using Police Crash Narratives

A study benchmarks frontier LLMs against official police crash databases, finding that basic keyword-rule baselines can be as effective as advanced models for specific attributes.

AI/ML arXiv cs.AI

Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art

A comparative analysis of NLP-based deception detection in legal contexts highlights strong domain sensitivity and the need for specialized adaptation over general LLMs.

AI/ML arXiv cs.AI

Federated Foundation Models Fine-Tuning with Heterogeneous Compressed Clients

FedSLM is a parameter-centric federated learning framework that allows heterogeneous compressed clients to fine-tune foundation models with reduced GPU memory requirements.

Open Source arXiv cs.AI

metasignal: A Python Package for Comprehensive Metacognitive Analysis and Decision-Making

metasignal is an open-source Python package for signal detection theory and metacognitive measurement, facilitating decision-making and psychological research.

AI/ML arXiv cs.AI

DoubleHelix: Structured Cross-Modal Fusion for Audio-Visual Speech Recognition with LLMs

DoubleHelix introduces a structured cross-modal fusion framework for audio-visual speech recognition, improving robustness in noisy environments.

AI/ML arXiv cs.AI

Multi-Granularity Position Embedding of Graphs via Granular-Ball for Link Prediction

The MGLP method introduces multi-granularity position embedding for graphs using Granular-Ball refinement to improve link prediction accuracy.

AI/ML arXiv cs.AI

InferQ: A Database-Oriented Benchmark for Quantum Circuits Simulation

InferQ is a database-oriented benchmark that evaluates using RDBMSs like PostgreSQL and DuckDB for quantum circuit simulation via SQL workloads.

Other Hacker News

Train Simulator Controller

A discussion regarding a Train Simulator Controller.

AI/ML The Verge

China’s Alibaba takes another swipe at America’s AI supremacy

Alibaba released Qwen3.8-Max, claiming it rivals US frontier models like those from OpenAI and Anthropic.

AI/ML arXiv cs.AI

Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates

Research presenting a method to speed up Latent Adversarial Training (LAT) for LLMs using low-rank defense and circuit-guided surrogates.