AI/ML arXiv cs.AI

Multi-Conditioned Diffusion Synthesis of Sand Boils for Low-Resource Earthen-Levee Inspection

A diffusion-based synthesis pipeline for generating synthetic sand-boil imagery to improve earthen-levee inspection.

Hardware/Chips Hacker News

Beavis Ultrasound PnP ISA Sound Card Replica

A replica of the Beavis Ultrasound PnP ISA Sound Card, recreating vintage computer hardware.

AI/ML arXiv cs.AI

TrustX Agent Risk Classification Framework (ARC): Risk-Tiering Internally Created Agentic AI Systems

The TrustX Agent Risk Classification Framework (ARC) provides a structured method for risk-tiering and governing agentic AI systems.

AI/ML arXiv cs.AI

Agora: Enhancing LLM Agent Reasoning Via Auction-Based Task Allocation

Agora is a framework that uses an auction-based mechanism to dynamically allocate tasks to expert LLM models and tools based on competence.

AI/ML arXiv cs.AI

ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI

ConceptSMILE is a perturbation-based auditing framework designed to evaluate the reliability of concept-based explainable AI (XAI).

AI/ML arXiv cs.AI

Minimal Decision Dynamics and Contextual Probability: A Quantum Tug-of-War Model

A research paper proposing a quantum-like extension of the Tug-of-War model to represent contextual probability in decision-making dynamics.

AI/ML arXiv cs.AI

REFORGE: A Method for Benchmarking LLMs' Reverse Engineering Capabilities in Decompiled Binary Function Naming

Reforge is a provenance-tracked pipeline for benchmarking the reverse engineering capabilities of LLMs in decompiled binary function naming.

AI/ML arXiv cs.AI

A Unified Approach to Interpreting Knowledge Distillation for Large Language Models via Interactions

A new approach to knowledge distillation in LLMs using 'interactions' and a 'Complex Interaction Penalty' (CIP) loss function to improve performance.

AI/ML arXiv cs.AI

iLENS: Interpretable LLM-Guided Mixture-of-Experts for Neuroimaging Survival Analysis

iLENS is an interpretable LLM-guided Mixture-of-Experts framework for neuroimaging survival analysis in Alzheimer's Disease prediction.

AI/ML arXiv cs.AI

Signed Symmetric Quantization for Few-Bit Integers

Signed Symmetric Quantization is proposed as a lightweight alternative to asymmetric quantization for few-bit integers in LLMs, reducing error without runtime penalties.

AI/ML arXiv cs.AI

Sticky Routing: Training MoE Models for Memory-Efficient Inference

StickyMoE introduces a differentiable routing consistency loss to reduce expert switching in MoE models, making inference more memory-efficient on edge devices.

Hardware/Chips Hacker News

Guy took Jupiter photo with Game Boy Camera, giant telescope, publishes tutorial

A hobbyist successfully captured a photo of Jupiter using a Game Boy Camera and a giant telescope, providing a tutorial for others.

AI/ML arXiv cs.AI

Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review

Introduces AutoWorldBuilder, a multi-agent LLM system for fictional worldbuilding featuring hierarchical context compression and iterative review.

AI/ML arXiv cs.AI

How Does Bayesian Causal Discovery Fail? Characterising Structural Consequences in Linear Gaussian Networks under Latent Confounding

Analyzes how Bayesian causal discovery fails under latent confounding in linear Gaussian networks, identifying critical correlation thresholds.

AI/ML arXiv cs.AI

ProofCouncil: An LLM Agent for Solving Open Mathematical Problems

Presents ProofCouncil, an author-critic LLM agent capable of solving open mathematical problems, and releases its agent-building library as open source.

AI/ML arXiv cs.AI

Ceci n'est pas une pipe: AI systems as semantic abstractions

Proposes a semantic framework to define and check AI outputs as engineered representations rather than direct facts, targeting failures like extrapolation.

AI/ML arXiv cs.AI

Multimodal Reward Hacking in Reinforcement Learning

Investigates multimodal reward hacking in RL for MLLMs, demonstrating that outcome-only rewards often lead to new failures rather than improved performance.

AI/ML arXiv cs.AI

Shared Selective Persistent Memory for Agentic LLM Systems

Introduces shared selective persistent memory for LLM agents to retain reusable context across sessions, significantly reducing token costs and improving task completion.

AI/ML arXiv cs.AI

SAGEAgent: A Self-Evolving Agent for Cost-Aware Modality Acquisition in Multimodal Survival Prediction

Presents SAGEAgent, a self-evolving LLM agent that optimizes the sequence of diagnostic modality acquisition in cancer survival prediction to reduce patient burden.

AI/ML arXiv cs.AI

Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI

Discusses the 'vocabulary' and 'verifier' gaps in current AI, arguing that true open-ended intelligence requires the ability to create new representational primitives.