AI/ML arXiv cs.AI

The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models

Analyzes asymmetric effects of knowledge distillation on bias in small LLMs and proposes the Per-Condition Calibration Diagnosis (PCCD) protocol.

AI/ML arXiv cs.AI

The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?

Introduces the Agentic Formalism Trap and the Evaluative Dissonance Index to quantify how LLM-as-a-Judge systems can be misled by structural proceduralism.

AI/ML arXiv cs.AI

Seeing Differently: Modeling Interpretive Perspectives in Computational Creativity using a Four-World Framework

Proposes a four-world framework for modeling interpretive perspectives in computational creativity, using a persona-based evaluation approach.

Software Engineering arXiv cs.AI

Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation

Introduces MobileForge, a benchmark for project-level multi-screen mobile app generation that evaluates build, navigation, and maintainability.

AI/ML arXiv cs.AI

ConnectED: A Curriculum-Aligned AI System for Vietnamese Instructional Lesson Planning and Student Learning

Presents ConnectED, an AI system for Vietnamese instructional lesson planning based on the VietEduQwen model.

AI/ML arXiv cs.AI

Why It Hurts: Identifying the Drivers of Negative Thoughts in Emotional Support Conversations

Introduces the AppraiSal benchmark and the PRISM probabilistic framework to help LLMs identify salient cognitive appraisal dimensions in emotional support.

AI/ML arXiv cs.AI

COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention

Presents COSI-Lab, a multimodal dataset of social interactions at a scientific workshop to model multi-perspective social intention.

Other Hacker News

Don't be a meat proxy

A discussion on the dangers of acting as a 'meat proxy'—performing tasks for AI that it cannot yet do, thereby hindering its own evolution and personal efficiency.

Other Hacker News

More German than many Germans

An exploration of German language and culture, likely discussed in the context of linguistic nuances or identity.

Software Engineering Hacker News

Rust project goals: Immobile types and guaranteed destructors

Updates on the Rust project goals, specifically focusing on the implementation of immobile types and guaranteed destructors to improve memory safety and predictability.

Other Hacker News

Òrbites – Connection and constellation puzzle game

A showcase of 'Òrbites', a puzzle game centered around connection and constellations.

AI/ML arXiv cs.AI

AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction

Introduces AMTFV, a framework that decouples verification modeling from execution to improve LLM self-correction in mathematical problem solving.

AI/ML arXiv cs.AI

COntExt: Towards Context-Aware Ontology Extension from Operational Metrics

Presents COntExt, a framework for context-aware ontology extension using operational metrics to reduce the manual effort of maintaining formal ontologies.

AI/ML arXiv cs.AI

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

Introduces LEMUR, a framework for multi-objective reinforcement learning that aligns agents with multiple human preferences without pre-defined reward functions.

AI/ML arXiv cs.AI

DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat

Introduces DungeonBench, a tactical reasoning benchmark for LLMs based on Dungeons & Dragons combat, testing resource budgeting and rule-aware discipline.

AI/ML arXiv cs.AI

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers

Introduces AgentHPOBench, a benchmark to evaluate the ability of LLM agents to act as sequential hyperparameter optimizers in ML tasks.

Other arXiv cs.AI

Development of FDD-ON: an Ontology for VAV HVAC System Fault Detection and Diagnostics

Proposes FDD-ON, a modular ontology for fault detection and diagnostics in VAV HVAC systems to improve interoperability and AI-driven maintenance.

Other Hacker News

Convergence Is Not Enough

A discussion on Hacker News regarding the concept of convergence and whether it is sufficient for progress.

AI/ML arXiv cs.AI

CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents

Presents CAGE, a framework for certifying authorization in tool-using LLM agents to prevent unsafe actions caused by small binding errors or numerical drift.

AI/ML arXiv cs.AI

MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft

Introduces MirrorCraft, a benchmark for evaluating LLM agents in Minecraft under hidden rule changes to test their adaptability beyond fixed mechanics.