AI/ML arXiv cs.AI

Understanding Rollout Error in Graph World Models

Researchers propose Error-Aware Graph World Models (GWMs) to reduce rollout error and planning regret in environments represented as graphs.

AI/ML arXiv cs.AI

Grounded Iterative Language Planning: How Parameterized World Models Reduce Hallucination Propagation in LLM Agents

The Grounded Iterative Language Planning (GILP) framework combines parameterized world models with LLM reasoning to significantly reduce hallucinations in agents.

AI/ML arXiv cs.AI

ATOD: Annealed Turn-aware On-policy Distillation for Multi-turn Autonomous Agents

ATOD is a new hybrid distillation algorithm that blends on-policy distillation and RL to improve the training of small language-model agents for long-horizon tasks.

AI/ML arXiv cs.AI

NormAct: A Benchmark for Hidden Social Norm Compliance in Embodied Planning

The NormAct benchmark evaluates whether embodied AI agents can comply with hidden social norms, introducing NormPerceptor to help agents detect such norms.

AI/ML arXiv cs.AI

Verifiable Geometry Problem Solving: Solver-Driven Autoformalization and Theorem Proposing

SD-GPS is a solver-driven framework for geometry problem solving that uses a symbolic solver as an execution oracle to improve autoformalization and theorem proposing.

AI/ML arXiv cs.AI

RelBall: Relation Ball with Quaternion Rotation for Knowledge Graph Completion

RelBall is introduced as a Knowledge Graph Completion model using quaternion rotations and modulus transformations to better handle complex relational patterns and hierarchies.

AI/ML arXiv cs.AI

Lifted Causal Inference

The paper introduces Lifted Causal Inference (LCI) using parametric causal factor graphs to efficiently compute causal effects in relational domains.

AI/ML arXiv cs.AI

JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications

JD.com details Oxygen AIIC, an industrial-scale LLM/VLM-centric platform for managing item knowledge across billions of SKUs.

AI/ML arXiv cs.AI

Ontology-Guided Evidence Path Inference for Multi-hop Knowledge Graph Question Answering

The OPI framework uses ontology-guided evidence path inference to reduce the search space and improve accuracy in multi-hop knowledge graph question answering.

Other Hacker News

Age verification is just a precursor to automated attribution of speech

A discussion on how age verification requirements may lead to the automated attribution of speech and loss of online anonymity.

Other Hacker News

Idler Magazine

A mention of Idler Magazine, likely a general interest or lifestyle publication.

AI/ML arXiv cs.AI

AI-Model Network: Concept, Current State and Future

Proposes AI-ModelNet, a novel paradigm for interconnecting heterogeneous AI models to enable capability sharing and collaborative reasoning.

AI/ML arXiv cs.AI

When Does Personality Composition Matter for Multi-Agent LLM Teams?

Studies how personality traits in multi-agent LLM teams affect task performance, finding that impact varies significantly by task structure.

AI/ML arXiv cs.AI

Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning

Introduces a three-stage training paradigm to give LLM agents internal world models for better prospective planning in long-horizon tasks.

AI/ML arXiv cs.AI

Odyssey: Constructing Verifiable Local Truth-Preserving Foundation Models

Presents ODYSSEY, a categorical framework for constructing verifiable local truth-preserving foundation models using foundries and Kan extensions.

AI/ML arXiv cs.AI

DysLexLens: A Low-Resource LLM Framework for Analysing Dyslexic Learners Insights from Online Forums

Proposes DysLexLens, a low-resource LLM framework for analyzing dyslexic learners' experiences with AI tools using forum data.

AI/ML arXiv cs.AI

MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy

Introduces MER-R1, an RL framework for multimodal emotion recognition that optimizes the synergy between 'fast' intuition and 'slow' reasoning.

AI/ML arXiv cs.AI

ToE: A Hierarchical and Explainable Claim Verification Framework with Dynamic Multi-source Evidence Retrieval and Aggregation

Presents the Tree of Evidence (ToE) framework for automated fact-checking to combat AI-generated misinformation and GEO poisoning.

AI/ML arXiv cs.AI

Towards Reliable and Robust LLM Planning: Symbolic Feedback-Driven Iterative Self-Refinement Framework

Proposes a symbolic feedback-driven self-refinement framework to improve the reliability and feasibility of LLM planning in long-horizon tasks.

Other Hacker News

Tell Congress: Don't Force Age Checks Online

A discussion urging Congress not to implement mandatory online age verification checks.