AI/ML arXiv cs.AI

Rashomon Alignment

Introduces Rashomon Alignment (RA), a geometrical measure to assess functional similarity between ML models regardless of specific data distributions.

AI/ML arXiv cs.AI

From Deterministic to Generative Deep Learning for Urban Air Quality Reconstruction from Sparse Observations

A diffusion-based generative framework for reconstructing urban air quality from sparse observations, outperforming deterministic deep learning models.

AI/ML arXiv cs.AI

Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction

Presents HYSET, a hyperedge-based set-level tool retrieval system that evaluates the joint utility of candidate tool sets for LLM agents.

AI/ML arXiv cs.AI

Shared Voxel-Map-Based Cooperative Indoor UAV Guidance with a Multi-Agent Soft Actor-Critic Controller

A cooperative indoor UAV guidance framework using a shared voxel-map world model and a multi-agent Soft Actor-Critic controller.

AI/ML arXiv cs.AI

Image Quality Dependent Degradation for AI Systems

A design strategy for automated driving systems to maintain safety in poor image quality by adjusting confidence thresholds using normalizing flows.

AI/ML arXiv cs.AI

SpectONet: A Physics-Guided Spectral Deep Operator Network for Euler-Bernoulli Beam Dynamics

Introduces SpectONet, a physics-guided spectral deep operator network for efficient and accurate Euler-Bernoulli beam vibration analysis.

AI/ML arXiv cs.AI

Lowering the implementation barrier of neutral-atom quantum computing with agentic workflows

An agentic workflow designed to automate the translation of theoretical quantum protocols into experiments on neutral-atom quantum processors.

AI/ML Hacker News

Sell Your AI Skills?

A Hacker News community discussion regarding the monetization and sale of AI-related skills in the current job market.

AI/ML arXiv cs.AI

Contrastive Representation Learning of Longitudinal Disease Trajectories on Temporal Graphs

Proposes a contrastive representation learning framework using temporal graphs and graph neural networks to model disease trajectories from longitudinal clinical data.

AI/ML arXiv cs.AI

A Human-in-the-Loop Corpus for LLM-Based Simplification of Scientific Summaries

Introduces a human-in-the-loop workflow and a new corpus for simplifying complex scientific summaries using LLMs to make research more accessible to non-specialists.

AI/ML arXiv cs.AI

Construction-Driven Injection: Linguistically-Grounded Edit-Based Code-Mixing Fingerprints for Large Language Models

Presents a unified framework for injecting linguistically-grounded, code-mixing fingerprints into LLMs to enable black-box ownership verification and prevent unauthorized redistribution.

AI/ML arXiv cs.AI

F(AI)2R: Who Did What, and Who Checked? Verifiable AI Provenance as an Executable Skill

Introduces F(AI)2R and 'aiprov', a PROV-O extension for creating machine-readable, verifiable provenance graphs for AI-assisted research artefacts.

AI/ML arXiv cs.AI

OmniPhys: Knowledge-Graph-Driven Benchmarking and Collective Optimization for Physical Commonsense in Text-to-Image Generation

Introduces OmniPhys, a benchmark for testing physical commonsense in text-to-image models, and OmniPrompt, a framework to optimize physical consistency in generative AI.

Cybersecurity arXiv cs.AI

KQFuzz: Knowledge-Guided Fuzzing for Quantum Libraries via Large Language Models

Presents KQFuzz, an LLM-based fuzzer that uses codebase knowledge to find bugs in quantum computing libraries like Qiskit, PennyLane, and Cirq.

AI/ML arXiv cs.AI

Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing

Discusses the failure risks of current AI governance frameworks when applied to general-purpose AI, specifically within the context of public policing.

AI/ML arXiv cs.AI

MyMentorLLM: A psychotherapy GenAI environment with multimodal voice/text patients, trainees and experts for deliberate practice

Describes MyMentorLLM, a multimodal simulation environment using GenAI to provide deliberate practice for psychotherapists training in CBT.

AI/ML arXiv cs.AI

DynaBridge: Dynamic Summary-Guided Cross-Task Multimodal Fusion for DASS-Structured Mental Health Assessment

Introduces DynaBridge, a multimodal fusion framework that uses LLM-generated summaries to assess mental health risks based on DASS-21 psychometric structures.

AI/ML Hacker News

LLM Honeypot

A discussion on Hacker News regarding the creation of a honeypot designed specifically for Large Language Models.

AI/ML arXiv cs.AI

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation

Introduction of Argus-Unified, a compact multimodal model that unifies image understanding and generation with significantly lower compute and data requirements.

AI/ML arXiv cs.AI

Multi-Scale Structural Features for Continual, Comprehensible Visual Recognition in a Developmental Learning Framework

A new visual feature representation for a developmental learning framework that enables continual, human-interpretable visual recognition without destructive forgetting.