AI/ML arXiv cs.AI

Robust Feasible Route Construction through Collaborative Partition Optimization

Collaborative Routing Constructors (CoRC) is a new framework for solving large-scale Capacitated Vehicle Routing Problems by allowing subproblems to exchange customers and vehicles.

Software Engineering Hacker News

Tiny-C Reference Manual Excerpt

An excerpt from the Tiny-C reference manual, likely focusing on language specifications or implementation details.

AI/ML arXiv cs.AI

Evaluating Generative Agents with Actions Grounded in Socially Distributed Task Environments using Incognita

Introduces Incognita, a framework for evaluating generative agents in socially distributed task environments where knowledge is partitioned among participants.

AI/ML arXiv cs.AI

Reinforcement Learning for Evidence-Seeking Diagnostic Reasoning with Large Language Models

Proposes a framework for medical diagnosis using RLVR and introduces RAGES, a retrieval-augmented examination simulator for biologically plausible feedback.

AI/ML arXiv cs.AI

Beyond Forecasting: The Belief-to-Trade Layer in Prediction-Market Agents

Presents Raven-Agent, an autonomous trading agent for prediction markets that outperforms other policies in risk-adjusted returns.

AI/ML arXiv cs.AI

Human-Centric Reflective Architecture for Human-AI Collaborative Decision-Making

Describes the Human-Centric Reflective Architecture (HCRA) to improve human-AI collaborative decision-making through iterative, reflective RL processes.

AI/ML arXiv cs.AI

Silicon Sampling via Cross-Survey Transfer

Evaluates 'silicon sampling' (LLMs simulating survey respondents) using a cross-survey transfer framework to test individual-level predictability.

AI/ML arXiv cs.AI

APeB: Benchmarking Personalization Ability of Large Language Model Agents

Introduces APeB, a benchmark for measuring the personalization abilities of LLM agents in product search tasks using underspecified queries.

AI/ML arXiv cs.AI

Organizational Memory for Agentic Business Process Execution

Proposes an 'organizational memory' architecture to store and govern shared procedural knowledge for LLM-based business process execution.

AI/ML arXiv cs.AI

Embodied Operators and Benchmarking: Toward Reusable and Deployable Embodied Intelligence Systems

Defines 'embodied operators' as reusable functional modules for robotics and proposes a taxonomy and benchmark for evaluating them.

AI/ML arXiv cs.AI

Reflective Dialogue or Prompt Refinement? Effects of Tutor Scaffolding on Students' Independent LLM Use for Programming

Studies the impact of Socratic-Guidance vs. Prompt-Refinement tutors on students' ability to use LLMs for programming independently.

AI/ML arXiv cs.AI

iFLYTEK-Embodied-Omni Technical Report

iFLYTEK-Embodied-Omni is a unified multimodal foundation model that jointly models vision, language, and action for general-purpose embodied agents.

AI/ML arXiv cs.AI

Internal Pluralism and the Limits of Pairwise Comparisons

This research explores the limitations of pairwise comparisons in preference learning, proposing a model that accounts for internal pluralism and indecision.

AI/ML arXiv cs.AI

ASK in the Dark: Uncertainty-Gated LLM Assistance under Partial Observability

The ASK+ framework improves SLM-guided reinforcement learning agents under partial observability by providing trajectory-aware context and structured chain-of-thought reasoning.

Open Source arXiv cs.AI

Automated Data Readiness for Scientific AI

REDI is an open-source framework for automating the transformation and readiness assessment of large-scale scientific datasets for AI training.

AI/ML arXiv cs.AI

SwarmResearch: Orchestrating Coding Agents for Open-Ended Discovery

SwarmResearch introduces an orchestrator-subagent harness that uses a Shepherd Agent to steer a population of Search Agents for open-ended coding discovery.

AI/ML arXiv cs.AI

Object-Centric Environment Modeling for Agentic Tasks

Object-Centric Environment Modeling (OCM) enables LLM agents to build executable object-centric models of their environment to improve interaction and reduce invalid actions.

AI/ML arXiv cs.AI

MedCalc-Pro: Solving Complex Medical Calculations with LLM Agents

MedCalc-Pro is a new benchmark and agent framework designed to solve complex medical calculations requiring multi-tool selection and nested-tool calling.

AI/ML arXiv cs.AI

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models

Oyster-II is a reinforcement learning-based constructive safety alignment framework for LLMs that avoids blanket refusals while maintaining high safety and helpfulness.

Open Source arXiv cs.AI

VERITAS: Towards a General-Purpose Replication Tool for Scientific Research

VERITAS is a domain-agnostic replication framework using CLI coding agents to automate the independent verification of scientific research claims.