AI/ML arXiv cs.AI

PEARL: Auditable Repair for Scientific Reasoning Graph Extraction

PEARL is a training-free framework that converts noisy LLM-generated scientific reasoning graphs into auditable, semantically valid graphs.

AI/ML arXiv cs.AI

The Autonomous Agency Scale: A Behavioral Framework for Measuring Self-Directed Behavior in AI Systems

The Autonomous Agency Scale (AAS) is introduced as a behavioral framework to measure self-directed behavior in AI systems across seven dimensions of agency.

AI/ML arXiv cs.AI

Towards Agentic Agent-based Models: Feasibility, Performance, and Statistical Model Checking

This study explores the feasibility and impact of integrating LLM-driven decisions into Agent-based Models (ABMs) using the Mesa Python library.

AI/ML arXiv cs.AI

OntoExtend: A Framework for Requirement-driven and Scalable Ontology Extension with LLMs

OntoExtend is a RAG-based framework designed to scale ontology extension by using LLMs to propose grounded extensions based on requirements.

AI/ML arXiv cs.AI

SAGE: Subgoal-Conditioned Action Generation for Latent World Model Planning

SAGE introduces a prior-conditioned planner that uses latent subgoals to improve long-horizon planning in world models.

AI/ML arXiv cs.AI

Do Maps Still Matter for Machines: Revisiting the Role of Choropleth Maps in Foundation Model Spatial Understanding

The ChoroplethMap-Bench benchmark demonstrates that combining maps with symbolic data significantly improves foundation models' spatial reasoning.

AI/ML arXiv cs.AI

PAMD: Structured Adaptive Distances for Bisimulation Representations in Visual Reinforcement Learning

PAMD introduces a Pairwise Adaptive Mahalanobis Distance to improve latent state similarity measurements in visual reinforcement learning.

AI/ML arXiv cs.AI

Rethinking Heterogeneous LLM Merging: A Weighted Model Averaging Perspective

This research explores a simple, training-free approach to merging heterogeneous LLMs using dimensional adaptation and weighted averaging.

Homelab/Self-Hosting arXiv cs.AI

AdaHome: An Adaptive Smart Home Assistant using Local Small Language Models

AdaHome is an adaptive smart home assistant that uses local small language models and an intent-aware planning framework to reduce latency and enhance privacy.

Tech Business/VC Hacker News

Postmortem of a British Startup: Tract

A postmortem analysis of the British startup Tract, discussing the reasons behind its failure.

Tech Business/VC TechCrunch

Gritt exits stealth with $34 million for robots to build solar plants—then, everything else

Robotics startup Gritt emerges from stealth with $34 million in funding to automate construction site tasks, specifically for solar plants.

Hardware/Chips The Verge

Who’s afraid of the big, bad GPU?

An exploration of the environmental and social costs of the AI GPU boom, including energy consumption, water usage, and e-waste.

AI/ML arXiv cs.AI

WuYu-EnvLE-Bench: A Benchmark for Evaluating Large Language Models in Environmental Law Enforcement

Introduction of WuYu-EnvLE-Bench, a benchmark for evaluating LLMs in environmental law enforcement cases.

Cybersecurity arXiv cs.AI

Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models

MIND, a cognitive jailbreak framework for text-to-image models that models latent defense mechanisms to generate adversarial prompts.

AI/ML arXiv cs.AI

Financial Audit Assistance using Misinformation Detection and Explanation

A system using unsupervised techniques for misinformation detection and explanation in financial audit assistance.

AI/ML arXiv cs.AI

PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model

PGN (Pangu Navigator), an offline vision-language navigation system based on the OpenPangu-7B multimodal foundation model.

Hardware/Chips arXiv cs.AI

A Hardware-oriented Approach for Efficient Bayesian Inference Computation and Deployment

A hardware-oriented approach to accelerate discrete Bayesian inference on embedded GPUs via memory layout restructuring and tensor clustering.

AI/ML arXiv cs.AI

Exploratory and Assimilating Reflection: Reflective Recall Cycle for Long-term Memory

The EAR framework introduces Exploratory and Assimilating Reflection to improve long-term memory retrieval for LLM-based autonomous agents.

AI/ML arXiv cs.AI

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding

ST-Veto, a training-free method to improve diffusion multimodal LLM reasoning by vetoing temporally unstable and weakly grounded tokens.

AI/ML arXiv cs.AI

Mechanistic Attention Guidance for Agent Memory Refinement

Researchers propose AGMR, a framework that uses retrieval-head attention signals to guide targeted memory updates for AI agents, improving efficiency and performance over text-only methods.