All Articles
15966 articles total
The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models
Analyzes asymmetric effects of knowledge distillation on bias in small LLMs and proposes the Per-Condition Calibration Diagnosis (PCCD) protocol.
The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?
Introduces the Agentic Formalism Trap and the Evaluative Dissonance Index to quantify how LLM-as-a-Judge systems can be misled by structural proceduralism.
Seeing Differently: Modeling Interpretive Perspectives in Computational Creativity using a Four-World Framework
Proposes a four-world framework for modeling interpretive perspectives in computational creativity, using a persona-based evaluation approach.
Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation
Introduces MobileForge, a benchmark for project-level multi-screen mobile app generation that evaluates build, navigation, and maintainability.
ConnectED: A Curriculum-Aligned AI System for Vietnamese Instructional Lesson Planning and Student Learning
Presents ConnectED, an AI system for Vietnamese instructional lesson planning based on the VietEduQwen model.
Why It Hurts: Identifying the Drivers of Negative Thoughts in Emotional Support Conversations
Introduces the AppraiSal benchmark and the PRISM probabilistic framework to help LLMs identify salient cognitive appraisal dimensions in emotional support.
COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention
Presents COSI-Lab, a multimodal dataset of social interactions at a scientific workshop to model multi-perspective social intention.
Don't be a meat proxy
A discussion on the dangers of acting as a 'meat proxy'—performing tasks for AI that it cannot yet do, thereby hindering its own evolution and personal efficiency.
More German than many Germans
An exploration of German language and culture, likely discussed in the context of linguistic nuances or identity.
Rust project goals: Immobile types and guaranteed destructors
Updates on the Rust project goals, specifically focusing on the implementation of immobile types and guaranteed destructors to improve memory safety and predictability.
Òrbites – Connection and constellation puzzle game
A showcase of 'Òrbites', a puzzle game centered around connection and constellations.
AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction
Introduces AMTFV, a framework that decouples verification modeling from execution to improve LLM self-correction in mathematical problem solving.
COntExt: Towards Context-Aware Ontology Extension from Operational Metrics
Presents COntExt, a framework for context-aware ontology extension using operational metrics to reduce the manual effort of maintaining formal ontologies.
LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback
Introduces LEMUR, a framework for multi-objective reinforcement learning that aligns agents with multiple human preferences without pre-defined reward functions.
DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat
Introduces DungeonBench, a tactical reasoning benchmark for LLMs based on Dungeons & Dragons combat, testing resource budgeting and rule-aware discipline.
AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers
Introduces AgentHPOBench, a benchmark to evaluate the ability of LLM agents to act as sequential hyperparameter optimizers in ML tasks.
Development of FDD-ON: an Ontology for VAV HVAC System Fault Detection and Diagnostics
Proposes FDD-ON, a modular ontology for fault detection and diagnostics in VAV HVAC systems to improve interoperability and AI-driven maintenance.
Convergence Is Not Enough
A discussion on Hacker News regarding the concept of convergence and whether it is sufficient for progress.
CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents
Presents CAGE, a framework for certifying authorization in tool-using LLM agents to prevent unsafe actions caused by small binding errors or numerical drift.
MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft
Introduces MirrorCraft, a benchmark for evaluating LLM agents in Minecraft under hidden rule changes to test their adaptability beyond fixed mechanics.