All Articles
16483 articles total
Traceable Scholarship: Page Anchors and Ariadne's Thread for Humanistic Inquiry in the Age of Generative AI
The paper proposes 'Traceable Scholarship' to ensure AI-generated academic text is backed by verifiable page anchors and evidence, implemented via AIH-Infra.
OPOD: On-Policy Omni Distillation
On-Policy Omni Distillation (OPOD) improves omni-modal models by routing student responses to modality-specific teachers for targeted guidance.
Representing Entity Importance in AI Knowledge Systems: A Dual-Signal Framework of Audience Evaluation and Structural Authority
A new dual-signal framework suggests that AI knowledge systems should maintain separate representations for audience evaluation and structural authority rather than a single importance score.
SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration
SciExplore is a new benchmark for evaluating how LLM agents handle complex, realistic scientific information-seeking and reasoning workflows.
Clustered Edge Intelligence: Beyond Just Convergence of Edge Computing and AI
The paper introduces 'Clustered Edge Intelligence' (CEI), a framework to treat derived intelligence as a first-class, shareable entity across the edge-cloud continuum.
From Scalars to Time Series: Rethinking Implicit Neural Representations for Time-Varying Volumetric Data
A new approach to Implicit Neural Representations (INRs) uses sequence-level supervision over spatial time series to reduce training costs and improve quality for volumetric data.
Projects every RC live races and results
A Hacker News thread discussing RC live races and results.
ArbiGraph: Arbitrarily Scalable Verifiable Task Graphs for Evaluating Context Management
Introduction of ArbiGraph, a benchmark generator for evaluating how tool-assisted language agents manage context across complex reasoning workflows.
The Human-AI Substitution Principle: When will you be replaced by AI in your organization?
An analytical model studying Human-AI Task Allocation (HAT) to determine the structural conditions under which AI replaces human labor in organizations.
Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling
A study on Refusal-Gated Decoding, a method to maintain LLM safety guardrails and refusal behavior during high-temperature sampling.
Can an AI System Be Creative? A Critical Perspective from Art and Engineering
A philosophical and technical exploration of AI's capacity for creativity, arguing that AI lacks the ability for transformational creativity and serendipity.
Profiling Lightweight Large Language Models
A paper introducing a PTME-based framework for precision-aware profiling of lightweight LLMs to optimize deployment in resource-constrained environments.
Enhancing Explainable Cardiac Diagnosis with Guide-Grounded Multimodal LLMs
A multimodal LLM framework that improves cardiac diagnosis by grounding ECG report generation in structured clinical knowledge guides.
Efficient and Interpretable Body-Based Emotion Recognition with Lightweight Temporal Convolutional Networks
Research on using lightweight Temporal Convolutional Networks (TCNs) as an efficient, interpretable alternative to graph-based models for body-based emotion recognition.
Auditing Provenance Sensitivity in LLM Agent Action Selection
An audit of LLM agent action selection to determine how provenance and source authority influence tool and argument selection.
Auditing Evidence Use in Medical LLM Diagnosis
A behavioral audit of evidence use in medical LLMs to identify where diagnostic accuracy may hide failures in using clinical evidence.
The PImpl idiom and the C++26 std:indirect type
A discussion on the PImpl (Pointer to Implementation) idiom and the introduction of std::indirect type in C++26.
Russia's businesses under strain from Ukraine's attacks on Wildberries
Russian businesses are facing operational strain due to Ukrainian attacks targeting the Wildberries platform.
DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers
Introduction of DynamicMCPBench, a reusable framework for evaluating LLM agents over live Model Context Protocol (MCP) servers using effect-scored benchmarks.
AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use
AppWorld-UL is a new 'user-in-the-loop' benchmark designed to evaluate how tool-use agents interact with users to resolve ambiguities and constraints.