AI/ML arXiv cs.AI

Soft-TransFormers for Continual Learning

Soft-TransFormers (Soft-TF) is a continual learning framework that uses soft subnetworks to prevent forgetting while adapting pre-trained Transformers.

AI/ML arXiv cs.AI

A Self-Supervised Framework for Space Object Behaviour Characterisation

A self-supervised framework using a Perceiver-VAE architecture for characterizing space object behavior and detecting anomalies in orbital populations.

AI/ML arXiv cs.AI

Parameter-Efficient Continual Fine-Tuning: A Survey

A survey on Parameter-Efficient Continual Fine-Tuning (PECFT), exploring the synergy between continual learning and PEFT to combat catastrophic forgetting.

Other Hacker News

Museum of the Human Web

A showcase of the 'Museum of the Human Web', reflecting on the early internet and its human-centric design.

AI/ML Hacker News

Anthropomorphism in Children's Interactions with LLM Chatbots

An exploration of how children anthropomorphize their interactions with LLM chatbots.

Other Hacker News

Designing a 4D Digital Archive for Ikebana

A project detailing the design of a 4D digital archive specifically for the art of Ikebana.

AI/ML arXiv cs.AI

Fluid Reasoning Representations

Introduces Fluid Reasoning Representations (FRRs) to explain how LLMs organize concepts during extended test-time thinking.

AI/ML arXiv cs.AI

LLM-Grounded Explainable AI for Supply Chain Risk Early Warning via Temporal Graph Attention Networks

Proposes a framework using Temporal Graph Attention Networks and LLMs for explainable supply chain risk early warning systems.

AI/ML arXiv cs.AI

FinRAG-12B: A Production-Validated Recipe for Grounded Question Answering in Banking

Presents FinRAG-12B, a domain-specific LLM for banking that optimizes for grounding, citation, and calibrated refusal.

AI/ML arXiv cs.AI

Frontier LLM-based agents can overcome the ontology curation bottleneck for natural phenotypes

Demonstrates that frontier LLM agents can perform phenotype annotation at a level comparable to human biocurators.

AI/ML arXiv cs.AI

AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning

Introduces AgentJet, a distributed swarm training framework for Agentic Reinforcement Learning with a decoupled multi-node architecture.

AI/ML arXiv cs.AI

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery

A comprehensive survey on AI for mathematical reasoning, covering language models, neuro-symbolic systems, and verified discovery.

AI/ML arXiv cs.AI

The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism

Formalizes the Theory of Mind Utility (ToM-U) to specify how agents infer others' beliefs through Local Epistemic World Models.

AI/ML Hacker News

Show HN: Cactus Hybrid: We taught Gemma 4 to know when it's wrong

A project called Cactus Hybrid demonstrates teaching Gemma 4 to recognize when its own outputs are incorrect.

AI/ML Hacker News

How we made our LeRobot video reader up to 15× faster

The developers of LeRobot discuss technical optimizations that improved their video reader's performance by up to 15x.

Other Hacker News

Fretboard Memorisation with Modular Arithmetic

An exploration of using modular arithmetic to assist in fretboard memorization for musical instruments.

Other Hacker News

Show HN: ValuePair – a friendship app that cares about values first

Introduction of ValuePair, a social application designed to prioritize shared values in friendship matching.

AI/ML arXiv cs.AI

FormGym: Doing Paperwork with Agents

The FormGym benchmark and FieldFinder tool address the difficulty of automated form-filling in image-only domains for AI agents.

AI/ML arXiv cs.AI

Learning, Reasoning, Refinement: A Framework for Kahneman's Dual-System Intelligence in GUI Agents

CogniGUI introduces a cognitive framework for GUI agents using dual-process theory and the ScreenSeek benchmark for better adaptability.

AI/ML arXiv cs.AI

Assistax: A Multi-Agent Hardware-Accelerated Reinforcement Learning Benchmark for Assistive Robotics

Assistax is a GPU-accelerated reinforcement learning benchmark for assistive robotics built on JAX and MuJoCo-MJX for high-throughput simulation.