All Articles
17650 articles total
Mechanical Conscience: A Mathematical Framework for Dependability of Machine Intelligence
Introduces 'Mechanical Conscience,' a mathematical framework designed to regulate the behavioral trajectories of distributed intelligent systems to ensure dependability.
Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework
Proposes a multi-dimensional behavioral framework to measure LLM reasoning quality beyond simple correctness, focusing on consistency, robustness, and efficiency.
Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches
A comprehensive survey on the adoption and maturity of Reasoning Language Models (RLMs) across 28 different scientific disciplines.
Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation
Introduces the TBS (Think-Before-Speak) framework for multi-agent social simulations, separating private internal reasoning from public utterances.
TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data?
Presents TerraBench, a benchmark and the TerraAgent framework for evaluating how AI agents reason over heterogeneous Earth-system data like satellite imagery and GIS.
Diagnosing and Mitigating Compounding Failures in Agentic Persuasion via Taxonomic Strategy Retrieval
Introduces Taxonomic Strategy RAG (TS-RAG) to prevent compounding errors and semantic leakage in agentic persuasion tasks.
Heuresis: Search Strategies for Autonomous AI Research Agents Across Quality, Diversity and Novelty
Introduces Heuresis, a framework for autonomous AI research agents to explore machine learning ideas across quality, diversity, and novelty.
Hey, That's My Model! Introducing Chain & Hash, An LLM Fingerprinting Technique
Presents Chain & Hash, a cryptographic fingerprinting technique to prove ownership of LLMs and detect misuse or theft.
Enhancing Hardware Fault Tolerance in Machines with Reinforcement Learning Policy Gradient Algorithms
Compares PPO and SAC reinforcement learning algorithms for enhancing hardware fault tolerance in autonomous machines.
Crossword Heatmap
A discussion or project related to creating a heatmap for crossword puzzles.
Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles
Introduction of PRIME, a framework using logic grid puzzles to evaluate and quantify implicit social biases in LLM reasoning.
GameDevBench: Evaluating Agentic Capabilities Through Game Development
Presentation of GameDevBench, a multimodal benchmark designed to evaluate AI agents' capabilities in complex game development tasks.
SleepLM: Natural-Language Intelligence for Human Sleep
SleepLM is a family of foundation models that align human sleep polysomnography with natural language for better interpretation and interaction.
Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks
Proposal of an Explicit Logic Channel (ELC) to validate and enhance MLLMs on zero-shot tasks by mimicking human logical reasoning.
XSkill: Continual Learning from Experience and Skills in Multimodal Agents
XSkill is a dual-stream framework that allows multimodal agents to continually learn from experiences and skills without parameter updates.
SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models
SocialOmni is a benchmark designed to evaluate the social interactivity and conversational competence of Omni-modal Large Language Models.
Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification
Introduction of Rule-VLN, a benchmark for rule-compliant urban navigation, and SNRM, a module to improve safety awareness in AI agents.
EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale
EvoMaster is a self-evolving agent framework for scientific discovery that enables agents to iteratively refine hypotheses and accumulate knowledge.
LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent
LiteResearcher is a scalable RL training framework that uses a virtual world to train deep research agents more efficiently and effectively.
Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?
An audit of performance-optimization benchmarks for coding agents (GSO, SWE-Perf, SWE-fficiency) reveals significant instabilities and scoring flaws that may misrepresent agent progress.