AI/ML arXiv cs.AI

Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation

A study evaluating federated pre-training reliability, finding that direct next-token prediction is a more reliable evaluation signal than downstream fine-tuning.

AI/ML arXiv cs.AI

Sensitivity Analysis of GRU, LSTM and Transformer Encoder in Classification of Automated Driving Systems

Sensitivity analysis of GRU, LSTM, and Transformer encoders for classifying automated driving systems (ADS) using telematics data.

AI/ML arXiv cs.AI

Guarantees on Dynamical System Distinguishability for LLM Token Generation

A theoretical framework for the distinguishability of LLM token generation based on modeling token embeddings as trajectories of a dynamical system.

AI/ML arXiv cs.AI

LAWFUL: Law-Aligned Witness for Faithful Use of Latents

Presentation of LAWFUL, a framework to verify if neural networks have learned governing physical laws and use them in internal computations.

AI/ML arXiv cs.AI

MPP-GNN: Subject-Adaptive Community Detection for fMRI-Based Alzheimer's Disease Classification

Introduction of MPP-GNN, a subject-adaptive community detection GNN for better Alzheimer's disease classification using fMRI data.

AI/ML arXiv cs.AI

Metaphor-Induced Algorithmic Steering: Cross-Domain Procedural Transfer in LLM Code Generation

Exploration of 'metaphorical algorithmic steering' where metaphors in prompts lead LLMs to generate less efficient code.

AI/ML arXiv cs.AI

Technological Advances in Detecting and Managing Cognitive Impairment in Older Adults: Trends, Challenges, and Future Directions

A review of technological advances in AI and neuroimaging for the detection and management of cognitive impairment in older adults.

AI/ML arXiv cs.AI

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction

Introduces ExtractBench, a benchmark for schema-guided enterprise document extraction that evaluates value accuracy, record completeness, and grounding.

AI/ML arXiv cs.AI

Scaffolding Critical Engagement with GenAI: Transforming Ethnic Minority Preparatory Students' Collaborative Discourse in Prompt Engineering Tasks

Research on how pedagogical scaffolding can shift ethnic minority students from passive consumption to critical co-creation when using GenAI.

Hardware/Chips arXiv cs.AI

Topology-Aware Data Movement for Disaggregated GPU Inference

Proposes a topology-aware transfer orchestrator to reduce latency in disaggregated GPU inference by optimizing KV cache movement.

AI/ML arXiv cs.AI

The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models

Analyzes asymmetric effects of knowledge distillation on bias in small LLMs and proposes the Per-Condition Calibration Diagnosis (PCCD) protocol.

AI/ML arXiv cs.AI

The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?

Introduces the Agentic Formalism Trap and the Evaluative Dissonance Index to quantify how LLM-as-a-Judge systems can be misled by structural proceduralism.

AI/ML arXiv cs.AI

Seeing Differently: Modeling Interpretive Perspectives in Computational Creativity using a Four-World Framework

Proposes a four-world framework for modeling interpretive perspectives in computational creativity, using a persona-based evaluation approach.

Software Engineering arXiv cs.AI

Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation

Introduces MobileForge, a benchmark for project-level multi-screen mobile app generation that evaluates build, navigation, and maintainability.

AI/ML arXiv cs.AI

ConnectED: A Curriculum-Aligned AI System for Vietnamese Instructional Lesson Planning and Student Learning

Presents ConnectED, an AI system for Vietnamese instructional lesson planning based on the VietEduQwen model.

AI/ML arXiv cs.AI

Why It Hurts: Identifying the Drivers of Negative Thoughts in Emotional Support Conversations

Introduces the AppraiSal benchmark and the PRISM probabilistic framework to help LLMs identify salient cognitive appraisal dimensions in emotional support.

AI/ML arXiv cs.AI

COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention

Presents COSI-Lab, a multimodal dataset of social interactions at a scientific workshop to model multi-perspective social intention.

Other Hacker News

Don't be a meat proxy

A discussion on the dangers of acting as a 'meat proxy'—performing tasks for AI that it cannot yet do, thereby hindering its own evolution and personal efficiency.

Other Hacker News

More German than many Germans

An exploration of German language and culture, likely discussed in the context of linguistic nuances or identity.

Software Engineering Hacker News

Rust project goals: Immobile types and guaranteed destructors

Updates on the Rust project goals, specifically focusing on the implementation of immobile types and guaranteed destructors to improve memory safety and predictability.