AI/ML arXiv cs.AI

BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation

Introduction of BrainG3N, a dual-purpose tokenizer for 3D brain MRI generation using a volumetric masked-autoencoder and diffusion transformer.

AI/ML arXiv cs.AI

Denoising Implicit Feedback for Cold-start Recommendation

DIF is a model-agnostic denoising method for improving cold-start recommendations in short video applications by filtering noisy implicit feedback.

AI/ML arXiv cs.AI

Deontic Policies for Runtime Governance of Agentic AI Systems

Researchers propose AgenticRei, a governance framework for LLM agents using a deontic policy language based on OWL to handle complex obligations and conflicts outside the LLM.

Other arXiv cs.AI

Measuring Curriculum Alignment across Topical Coverage, Competency, and Cognitive Depth: A Longitudinal Framework Applied to CS2013 and CS2023

A study introduces a longitudinal framework to measure how undergraduate computer science curricula align with international guidelines like CS2013 and CS2023.

AI/ML arXiv cs.AI

Diffusion Language Models: An Experimental Analysis

This paper provides a systematic experimental analysis of eight state-of-the-art Diffusion Language Models (DLMs), comparing their efficiency and quality across various benchmarks.

AI/ML arXiv cs.AI

Hidden Anchors in Multi-Agent LLM Deliberation

The authors model multi-agent LLM deliberation as a dynamical system with 'hidden anchors' (internal beliefs) to explain how agents can reach conclusions beyond their initial collective knowledge.

AI/ML arXiv cs.AI

DeXposure-Claw: An Agentic System for DeFi Risk Supervision

DeXposure-Claw is an agentic system for DeFi risk supervision that combines a graph time-series foundation model with deterministic monitors to provide auditable supervisory tickets.

AI/ML arXiv cs.AI

LLM Doesn't Know What It Doesn't Know: Detecting Epistemic Blind Spots via Cross-Model Attribution Divergence on Clinical Tabular Data

Research on clinical tabular data reveals that LLM confidence is often misleading and proposes a cross-model calibrator using attribution divergence to improve reliability.

AI/ML arXiv cs.AI

REVEAL++: Differentiable Phenotypic Grouping for Vision-Language Retinal Modeling of Alzheimer's Disease Risk

REVEAL++ introduces a continuous formulation of phenotypic grouping in contrastive learning to improve the prediction of Alzheimer's disease risk from retinal images and clinical narratives.

AI/ML arXiv cs.AI

Emergent Alignment

The paper introduces 'Emergent Alignment,' a technique using a 'conscience step' and DPO to enable LLMs to self-correct unethical outputs without an external judge.

AI/ML arXiv cs.AI

ITNet: A Learnable Integral Transform That Subsumes Convolution, Attention, and Recurrence

ITNet is introduced as a unified architecture based on a learnable integral transform that mathematically subsumes convolutions, attention, and recurrence.

AI/ML arXiv cs.AI

Uncertainty Decomposition for Clarification Seeking in LLM Agents

A new prompt-based uncertainty decomposition method improves the ability of LLM agents to proactively seek clarification when task specifications are ambiguous.

Other Hacker News

Horizons JPL Solar System Data Demo and NASA DSN Updates: Datastar, Common Lisp

A discussion on NASA's JPL Solar System Data Demo and DSN updates, mentioning Datastar and Common Lisp.

AI/ML arXiv cs.AI

Signals of Provenance: Practices & Challenges of Navigating Indicators in AI-Generated Media for Sighted and Blind Individuals

Research on the challenges sighted and blind individuals face when identifying AI-generated media markers, suggesting improved accessibility for provenance indicators.

AI/ML arXiv cs.AI

Revisiting Active Speaker Detection: An In-the-Wild Benchmark for Generalization and Robustness

Introduction of UniTalk, a new benchmark dataset for active speaker detection designed to improve model generalization in challenging, real-world environments.

AI/ML arXiv cs.AI

ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark

ASyMOB is presented as a high-resolution benchmark for symbolic mathematics to distinguish genuine reasoning from pattern memorization in LLMs.

AI/ML arXiv cs.AI

Self-Evolving Multi-Agent Systems via Textual Backpropagation

Proposed 'Agentic Neural Network' (ANN) framework that allows multi-agent systems to self-evolve roles and coordination through a process mirroring backpropagation.

AI/ML arXiv cs.AI

Grids Often Outperform Implicit Neural Representations at Compressing Dense Signals

A study finding that regularized grids with interpolation often outperform Implicit Neural Representations (INRs) for compressing dense signals.

AI/ML arXiv cs.AI

From Memorization to Parameter Interference: How Overtraining Experts Harms Model Merging

Research demonstrating that overtraining expert models can lead to parameter interference, which harms the performance of subsequent model merging.

AI/ML arXiv cs.AI

Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs

Introduction of Partial Model Collapse (PMC), a method for machine unlearning in LLMs that removes private data by deliberately triggering model collapse.