AI/ML arXiv cs.AI

Thermodynamic Measure of Intelligence

A theoretical paper proposing a thermodynamic measure of intelligence based on the amplification of rare but valid futures through recursive self-simulation.

AI/ML arXiv cs.AI

A Multi-Agent system for Multi-Objective constrained optimization

Presents MAMO, a multi-agent reinforcement learning approach for multi-objective constrained optimization in dynamic environments.

AI/ML arXiv cs.AI

Navigating Unreliable Parametric and Contextual Knowledge: Explicit Knowledge Conflict Resolution for LLM Inference

Introduces MACR, a multi-agent reasoning framework to resolve knowledge conflicts between an LLM's parametric knowledge and external context.

AI/ML arXiv cs.AI

Confidence-Aware Automated Assessment of Student-Drawn Scientific Models

A study on using Vision Transformers with confidence-aware scoring to automate the assessment of student-drawn scientific models.

AI/ML arXiv cs.AI

Lagrange: An Open-Vocabulary, Energy-Based Sparse Framework for Generalized End-to-End Driving

Introduces Lagrange, an open-vocabulary sparse driving framework using Masked Latent Fields to improve end-to-end autonomous driving.

AI/ML arXiv cs.AI

Leveraging systems' non-linearity to tackle the scarcity of data in the design of Intelligent Fault Diagnosis Systems

A method for intelligent fault diagnosis systems using deep transfer learning and non-linearity to handle data scarcity in vibration-based analysis.

AI/ML arXiv cs.AI

SoftSkill: Behavioral Compression for Contextual Adaptation

Proposes SoftSkill, a method that compresses natural-language agent skills into compact latent behavioral priors to improve LLM performance.

AI/ML arXiv cs.AI

Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining

A diagnostic study on mining interaction trajectories to automate the generation of skill libraries for computer-using agents.

Hardware/Chips Hacker News

The ISA Doesn't Matter Where It Counts

A community discussion on Hacker News regarding the relative importance of Instruction Set Architecture (ISA) in modern computing.

AI/ML arXiv cs.AI

Learning to Prompt: Improving Student Engagement with Adaptive LLM-based High-School Tutoring

Researchers developed a subject-aware prompting system for LLM-based high-school tutoring that adapts teaching strategies based on pedagogical features.

AI/ML arXiv cs.AI

RACL: Reasoning-Agent Control Layers for Continuous Metaheuristic Learning

Introduction of RACL, a reasoning-agent control layer that optimizes the search behavior of existing metaheuristic optimizers using a reasoning agent.

AI/ML arXiv cs.AI

BIM-Edit: Benchmarking Large Language Models for IFC-Based Building Information Modeling

BIM-Edit is a new benchmark for evaluating the ability of LLMs to perform natural-language editing of Building Information Models in IFC format.

AI/ML arXiv cs.AI

Modularity-Free Conflict-Averse Training for Generalized PINNs

The ModSync framework addresses capacity-induced failure modes in Physics-informed neural networks (PINNs) to improve convergence and accuracy in solving PDEs.

AI/ML arXiv cs.AI

Implicit Semantic-Aware Communication Based on Hypergraph Reasoning

HISR is a proposed hypergraph-based reasoning framework designed to improve the accuracy of implicit semantic interpretation in communication systems.

AI/ML arXiv cs.AI

Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact

A study suggests that the apparent psychological profiles of LLMs are largely artifacts of measurement bias rather than inherent personality traits.

AI/ML arXiv cs.AI

Beyond Accuracy: Measuring Logical Compliance of Predictive Models

The Rule Violation Score (RVS) is introduced as a new metric to measure how well predictive models comply with logical and domain-specific constraints.

AI/ML arXiv cs.AI

Augmenting Game AI with Deep Reinforcement Learning

A proposal for a framework to integrate deep reinforcement learning into game AI to create more believable and authentic non-player characters.

AI/ML arXiv cs.AI

QMFOL: Benchmarking Large Language Model Reasoning via Quantifiable Monadic First-Order Logic Test Case Generation

QMFOL is an automated framework for generating monadic first-order logic tasks to precisely evaluate the deductive reasoning capabilities of LLMs.

Software Engineering Hacker News

So You Want to Define a Well-Known URI

A discussion on the standards and practices for defining well-known URIs.

AI/ML Hacker News

Zen and the Art of Machine Learning Research

Insights and philosophies regarding the practice of machine learning research.