AI/ML arXiv cs.AI

Mechanical Conscience: A Mathematical Framework for Dependability of Machine Intelligence

Introduces 'Mechanical Conscience,' a mathematical framework designed to regulate the behavioral trajectories of distributed intelligent systems to ensure dependability.

AI/ML arXiv cs.AI

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework

Proposes a multi-dimensional behavioral framework to measure LLM reasoning quality beyond simple correctness, focusing on consistency, robustness, and efficiency.

AI/ML arXiv cs.AI

Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches

A comprehensive survey on the adoption and maturity of Reasoning Language Models (RLMs) across 28 different scientific disciplines.

AI/ML arXiv cs.AI

Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation

Introduces the TBS (Think-Before-Speak) framework for multi-agent social simulations, separating private internal reasoning from public utterances.

AI/ML arXiv cs.AI

TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data?

Presents TerraBench, a benchmark and the TerraAgent framework for evaluating how AI agents reason over heterogeneous Earth-system data like satellite imagery and GIS.

AI/ML arXiv cs.AI

Diagnosing and Mitigating Compounding Failures in Agentic Persuasion via Taxonomic Strategy Retrieval

Introduces Taxonomic Strategy RAG (TS-RAG) to prevent compounding errors and semantic leakage in agentic persuasion tasks.

AI/ML arXiv cs.AI

Heuresis: Search Strategies for Autonomous AI Research Agents Across Quality, Diversity and Novelty

Introduces Heuresis, a framework for autonomous AI research agents to explore machine learning ideas across quality, diversity, and novelty.

Cybersecurity arXiv cs.AI

Hey, That's My Model! Introducing Chain & Hash, An LLM Fingerprinting Technique

Presents Chain & Hash, a cryptographic fingerprinting technique to prove ownership of LLMs and detect misuse or theft.

AI/ML arXiv cs.AI

Enhancing Hardware Fault Tolerance in Machines with Reinforcement Learning Policy Gradient Algorithms

Compares PPO and SAC reinforcement learning algorithms for enhancing hardware fault tolerance in autonomous machines.

Other Hacker News

Crossword Heatmap

A discussion or project related to creating a heatmap for crossword puzzles.

AI/ML arXiv cs.AI

Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles

Introduction of PRIME, a framework using logic grid puzzles to evaluate and quantify implicit social biases in LLM reasoning.

AI/ML arXiv cs.AI

GameDevBench: Evaluating Agentic Capabilities Through Game Development

Presentation of GameDevBench, a multimodal benchmark designed to evaluate AI agents' capabilities in complex game development tasks.

AI/ML arXiv cs.AI

SleepLM: Natural-Language Intelligence for Human Sleep

SleepLM is a family of foundation models that align human sleep polysomnography with natural language for better interpretation and interaction.

AI/ML arXiv cs.AI

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks

Proposal of an Explicit Logic Channel (ELC) to validate and enhance MLLMs on zero-shot tasks by mimicking human logical reasoning.

AI/ML arXiv cs.AI

XSkill: Continual Learning from Experience and Skills in Multimodal Agents

XSkill is a dual-stream framework that allows multimodal agents to continually learn from experiences and skills without parameter updates.

AI/ML arXiv cs.AI

SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models

SocialOmni is a benchmark designed to evaluate the social interactivity and conversational competence of Omni-modal Large Language Models.

AI/ML arXiv cs.AI

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification

Introduction of Rule-VLN, a benchmark for rule-compliant urban navigation, and SNRM, a module to improve safety awareness in AI agents.

AI/ML arXiv cs.AI

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale

EvoMaster is a self-evolving agent framework for scientific discovery that enables agents to iteratively refine hypotheses and accumulate knowledge.

AI/ML arXiv cs.AI

LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent

LiteResearcher is a scalable RL training framework that uses a virtual world to train deep research agents more efficiently and effectively.

AI/ML arXiv cs.AI

Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?

An audit of performance-optimization benchmarks for coding agents (GSO, SWE-Perf, SWE-fficiency) reveals significant instabilities and scoring flaws that may misrepresent agent progress.