All Articles
15946 articles total
9front "This Was Supposed to Be Fun" Released
The release of 9front 'This Was Supposed to Be Fun', a new version of the community-driven Plan 9 fork.
HERO: History-Enriched Rollout Training for Long-Horizon Autoregressive Neural Operators
Introduces HERO, a training method for neural operators to improve long-horizon accuracy and stability in predicting partial differential equations.
Have I Seen You? Embedding Behavior Signals Synthetic Face Dataset Membership
Research demonstrating that synthetic face datasets can still leak membership information about the real faces used to train the generators.
Implicit Machine Learning Force Fields Accelerate Molecular Dynamics Simulations
Introduces implicit machine learning force fields (I-MLFFs) to accelerate molecular dynamics simulations by reducing compute and memory footprints.
Memory Provenance Laundering in LLM Agents: A Non-Amplification Firewall for Persistent Memory
Proposes the Provenance-Preserving Memory Firewall (PPMF) to prevent 'memory provenance laundering' in LLM agents with persistent memory.
ActFovea: Runtime Safeguarding for VLA Policies via Spatiotemporal Visual-Action Consistency
Presents ActFovea, a safeguarding framework for Vision-Language-Action policies in robotics to mitigate runtime disturbances via spatiotemporal consistency.
CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning
Introduces CLIFT, a method for closed-loop iterative fine-tuning of closed-weight robot foundation models using managed SFT APIs.
Show HN: Nightcrawler – A local AI pentesting agent running on a smartphone
Nightcrawler is a local AI-powered pentesting agent designed to run on smartphones for mobile security auditing.
Learning Lookahead Lemmas for Neural Network Verification
Researchers propose a lookahead-driven inprocessing framework to improve the efficiency and performance of neural network verifiers like Marabou and alpha-beta-CROWN.
Autonomous Repair for Multi-Agent Systems via Monte-Carlo Tree Search
The MARS framework utilizes Monte-Carlo Tree Search to automate the repair of errors in multi-agent systems, introducing the StateMAS benchmark for evaluation.
Benchmarking Frontier Large Language Models Against Official Crash Database Coding Using Police Crash Narratives
A study benchmarks frontier LLMs against official police crash databases, finding that basic keyword-rule baselines can be as effective as advanced models for specific attributes.
Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art
A comparative analysis of NLP-based deception detection in legal contexts highlights strong domain sensitivity and the need for specialized adaptation over general LLMs.
Federated Foundation Models Fine-Tuning with Heterogeneous Compressed Clients
FedSLM is a parameter-centric federated learning framework that allows heterogeneous compressed clients to fine-tune foundation models with reduced GPU memory requirements.
metasignal: A Python Package for Comprehensive Metacognitive Analysis and Decision-Making
metasignal is an open-source Python package for signal detection theory and metacognitive measurement, facilitating decision-making and psychological research.
DoubleHelix: Structured Cross-Modal Fusion for Audio-Visual Speech Recognition with LLMs
DoubleHelix introduces a structured cross-modal fusion framework for audio-visual speech recognition, improving robustness in noisy environments.
Multi-Granularity Position Embedding of Graphs via Granular-Ball for Link Prediction
The MGLP method introduces multi-granularity position embedding for graphs using Granular-Ball refinement to improve link prediction accuracy.
InferQ: A Database-Oriented Benchmark for Quantum Circuits Simulation
InferQ is a database-oriented benchmark that evaluates using RDBMSs like PostgreSQL and DuckDB for quantum circuit simulation via SQL workloads.
Train Simulator Controller
A discussion regarding a Train Simulator Controller.
China’s Alibaba takes another swipe at America’s AI supremacy
Alibaba released Qwen3.8-Max, claiming it rivals US frontier models like those from OpenAI and Anthropic.
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates
Research presenting a method to speed up Latent Adversarial Training (LAT) for LLMs using low-rank defense and circuit-guided surrogates.