AI/ML arXiv cs.AI

When a Verified World Model Still Loses: Play-Adequacy vs Prediction-Accuracy in LLM-Synthesized Code World Models

Analysis of 'Code World Models' showing that high prediction accuracy on transitions does not guarantee success in planning or actual gameplay.

AI/ML arXiv cs.AI

ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System

ReasFlow is an autonomous multi-agent system designed for reasoning-centric scientific discovery in applied mathematics, capable of generating research papers.

AI/ML arXiv cs.AI

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

Presentation of RxBrain, a foundation model for embodied cognition that integrates language-visual reasoning and imagination for robotic planning.

AI/ML arXiv cs.AI

How Artificial Intelligence LLM Engines Shape the Global Conflict Information Environment

A study on how LLM answer engines are susceptible to Generative Engine Optimization (GEO) and hallucinations, especially in conflict-related information.

AI/ML arXiv cs.AI

Align AI to Dynamic Human-AI Workflows

A theoretical paper arguing for a shift from static human preference emulation to interactive and complementary alignment in human-AI workflows.

AI/ML arXiv cs.AI

The Steering Budget: Examples beat Knobs

Research suggesting that providing concrete examples is significantly more effective for steering generative models than adjusting prompt 'knobs' or tags.

Other Hacker News

Old Icons

A Hacker News thread discussing old icons.

AI/ML arXiv cs.AI

Intelligent Three Level Learning Architecture for Autonomous UAV Swarms in Search and Rescue

A novel three-level hierarchical learning architecture for autonomous UAV swarms in search and rescue, combining neuroplasticity, MARL, and meta-learning.

AI/ML arXiv cs.AI

HG-RAG: Hierarchy-Guided Retrieval-Augmented Generation for Structured Knowledge Graphs

HG-RAG is a framework that improves RAG by performing graph-traversal over hierarchical knowledge graphs to enhance relational reasoning.

AI/ML arXiv cs.AI

IMEX Interaction-Based Model Explanation

The IMEX approach provides an interaction-based model explanation method to identify variable contributions and significant interactions in black-box models.

AI/ML arXiv cs.AI

RegNetAgents: A Multi-Agent Framework for Cross-Network Regulatory Driver Identification in Cancer Genomics

RegNetAgents is a multi-agent framework using LangGraph and MCP to identify regulatory drivers in cancer genomics across heterogeneous networks.

AI/ML arXiv cs.AI

DialogueVPR: Towards Conversational Visual Place Recognition

DialogueVPR introduces a conversational approach to visual place recognition and the DlgQuest-Cities benchmark for interactive geo-localization.

AI/ML arXiv cs.AI

Interpretable Language Model for Closed-Loop Type 1 Diabetes Control

LLM-T1D combines RL precision with LLM reasoning to create a transparent and reliable insulin pump controller for Type 1 Diabetes.

AI/ML arXiv cs.AI

Human AI Construction of Bayesian Networks for Operational Decision Support -- A Virtual Survey Approach

A methodology using AI agents and LLMs to bridge expert opinion and data-driven learning for constructing Bayesian Belief Networks.

AI/ML arXiv cs.AI

Capability from Access Structure, Not Scale: Lower Bounds and Pre-Registered Tests for Hybrid Sequence Models

The Capability Convergence Hypothesis argues that model capability depends on access structure (hybrid sequence models) rather than just scale.

AI/ML arXiv cs.AI

ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability

ToolAnchor is a framework that uses counterfactual anchor contexts to help LLM agents adapt to expanded toolsets and overcome behavioral inertia.

Other Hacker News

GrapheneOS recommended for domestic abuse victims

Hacker News discussion regarding the recommendation of GrapheneOS for individuals experiencing domestic abuse due to its enhanced privacy and security features.

AI/ML arXiv cs.AI

Partially Observed Structural Causal Models

Introduces Partially Observed Structural Causal Models (POSCMs), a framework for causal modeling in settings where upstream contexts co-determine interaction structures and downstream mechanisms.

AI/ML arXiv cs.AI

Stable Attention Response for Reliable Precipitation Nowcasting

Proposes HARECast, a framework that regulates head-wise attention response energy to improve the stability and reliability of precipitation nowcasting.

AI/ML arXiv cs.AI

TuxBot: Semantic-Aware Online OS Tuning with Large Language Models

Presents TuxBot, an LLM-guided framework for steady-state OS tuning that uses semantic-aware guidance to optimize Linux kernel and sysctl parameters.