AI/ML arXiv cs.AI

Traceable Scholarship: Page Anchors and Ariadne's Thread for Humanistic Inquiry in the Age of Generative AI

The paper proposes 'Traceable Scholarship' to ensure AI-generated academic text is backed by verifiable page anchors and evidence, implemented via AIH-Infra.

AI/ML arXiv cs.AI

OPOD: On-Policy Omni Distillation

On-Policy Omni Distillation (OPOD) improves omni-modal models by routing student responses to modality-specific teachers for targeted guidance.

AI/ML arXiv cs.AI

Representing Entity Importance in AI Knowledge Systems: A Dual-Signal Framework of Audience Evaluation and Structural Authority

A new dual-signal framework suggests that AI knowledge systems should maintain separate representations for audience evaluation and structural authority rather than a single importance score.

AI/ML arXiv cs.AI

SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration

SciExplore is a new benchmark for evaluating how LLM agents handle complex, realistic scientific information-seeking and reasoning workflows.

AI/ML arXiv cs.AI

Clustered Edge Intelligence: Beyond Just Convergence of Edge Computing and AI

The paper introduces 'Clustered Edge Intelligence' (CEI), a framework to treat derived intelligence as a first-class, shareable entity across the edge-cloud continuum.

AI/ML arXiv cs.AI

From Scalars to Time Series: Rethinking Implicit Neural Representations for Time-Varying Volumetric Data

A new approach to Implicit Neural Representations (INRs) uses sequence-level supervision over spatial time series to reduce training costs and improve quality for volumetric data.

Other Hacker News

Projects every RC live races and results

A Hacker News thread discussing RC live races and results.

AI/ML arXiv cs.AI

ArbiGraph: Arbitrarily Scalable Verifiable Task Graphs for Evaluating Context Management

Introduction of ArbiGraph, a benchmark generator for evaluating how tool-assisted language agents manage context across complex reasoning workflows.

AI/ML arXiv cs.AI

The Human-AI Substitution Principle: When will you be replaced by AI in your organization?

An analytical model studying Human-AI Task Allocation (HAT) to determine the structural conditions under which AI replaces human labor in organizations.

AI/ML arXiv cs.AI

Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling

A study on Refusal-Gated Decoding, a method to maintain LLM safety guardrails and refusal behavior during high-temperature sampling.

AI/ML arXiv cs.AI

Can an AI System Be Creative? A Critical Perspective from Art and Engineering

A philosophical and technical exploration of AI's capacity for creativity, arguing that AI lacks the ability for transformational creativity and serendipity.

AI/ML arXiv cs.AI

Profiling Lightweight Large Language Models

A paper introducing a PTME-based framework for precision-aware profiling of lightweight LLMs to optimize deployment in resource-constrained environments.

AI/ML arXiv cs.AI

Enhancing Explainable Cardiac Diagnosis with Guide-Grounded Multimodal LLMs

A multimodal LLM framework that improves cardiac diagnosis by grounding ECG report generation in structured clinical knowledge guides.

AI/ML arXiv cs.AI

Efficient and Interpretable Body-Based Emotion Recognition with Lightweight Temporal Convolutional Networks

Research on using lightweight Temporal Convolutional Networks (TCNs) as an efficient, interpretable alternative to graph-based models for body-based emotion recognition.

AI/ML arXiv cs.AI

Auditing Provenance Sensitivity in LLM Agent Action Selection

An audit of LLM agent action selection to determine how provenance and source authority influence tool and argument selection.

AI/ML arXiv cs.AI

Auditing Evidence Use in Medical LLM Diagnosis

A behavioral audit of evidence use in medical LLMs to identify where diagnostic accuracy may hide failures in using clinical evidence.

Software Engineering Hacker News

The PImpl idiom and the C++26 std:indirect type

A discussion on the PImpl (Pointer to Implementation) idiom and the introduction of std::indirect type in C++26.

Other Hacker News

Russia's businesses under strain from Ukraine's attacks on Wildberries

Russian businesses are facing operational strain due to Ukrainian attacks targeting the Wildberries platform.

AI/ML arXiv cs.AI

DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers

Introduction of DynamicMCPBench, a reusable framework for evaluating LLM agents over live Model Context Protocol (MCP) servers using effect-scored benchmarks.

AI/ML arXiv cs.AI

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use

AppWorld-UL is a new 'user-in-the-loop' benchmark designed to evaluate how tool-use agents interact with users to resolve ambiguities and constraints.