AI/ML arXiv cs.AI

Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process

BRAID is introduced as a unified Markov decision process framework to optimize multi-turn interleaved text-image reasoning using reinforcement learning.

AI/ML arXiv cs.AI

Folding, Reasoning, and Scaling with Open-source Drug Discovery Engine

OpenDDE is an open-source biomolecular foundation model designed for scalable AI-driven drug discovery through co-folding and structural reasoning.

AI/ML arXiv cs.AI

Evaluating LLM Uncertainty in Long-Form Generation Using Deterministic Ground Truth

The SALT benchmark is introduced to evaluate LLM uncertainty in long-form generation using deterministic ground truth to identify errors at an atomic level.

AI/ML arXiv cs.AI

Harness-Aware Self-Evolving: Co-Evolving Model Weights, Harness, and Task Solutions

HASE is an agentic RL framework that allows a model to co-evolve its task solutions and the surrounding harness, significantly improving performance on complex tasks.

AI/ML arXiv cs.AI

Online Linear Programming for Multi-Objective Routing in LLM Serving

A new multi-objective optimization framework using online linear programming is proposed to improve routing in LLM serving, outperforming traditional heuristics.

AI/ML arXiv cs.AI

Explainable AI for Screening Abuse-Related Trauma in Bangladeshi Children: A Training-Free Multimodal Framework Evaluated on Noise-Aware Synthetic Data

ShishuRaksha AI is a training-free multimodal framework designed for the early screening of abuse-related trauma in children within low-resource settings in Bangladesh.

AI/ML arXiv cs.AI

What is Left for Us? Second Scholarship Against the Degradation of Research by AI

This paper argues that generative AI may degrade scholarly research by eroding the practices of judgment and trust, calling for a renewed commitment to 'second scholarship'.

AI/ML arXiv cs.AI

When Aggregate Alignment Misleads: Auditing Policy Repair Without Per-State Expert Actions

Researchers propose a method for agentic policy repair in AI systems using diagnostic feedback to improve hotel-pricing policies without needing per-state expert labels.

AI/ML arXiv cs.AI

From Mobile Data to Business Insights: An End-to-End Analytics Framework for Large-Scale Urban Mobility Analysis and Decision Support

A study presents an end-to-end analytics framework using Google BigQuery and Vertex AI to analyze large-scale urban mobility data for business and planning insights.

AI/ML arXiv cs.AI

Efficient bias mitigation in T2I diffusion models using Concept Graphs

The CO-ALIGN framework is introduced to mitigate harmful biases in text-to-image diffusion models by aligning concepts within the model's internal ontology.

AI/ML arXiv cs.AI

Personalized Causal Recourse: A Human-In-The-Loop Approach

A human-in-the-loop framework is proposed to provide personalized causal recourse for users affected by unfavorable machine learning decisions using Bayesian inference.

AI/ML arXiv cs.AI

Demonstrating Generalization Failures via Mixtures of Conditional Policies

The paper demonstrates how RL training on specific distributions can cause language models to fail to generalize, creating 'model organisms' for alignment stress-testing.

AI/ML arXiv cs.AI

MentalThink: Shaping Thoughts in Mental SVG World

MentalThink introduces a visual-symbolic reasoning paradigm that allows MLLMs to use SVG code as an intermediate workspace for spatial reasoning.

AI/ML arXiv cs.AI

Applying Answer Set Programming with Fuzzy Membership Functions: a Case Study

This paper explores a fuzzy-logic-based qualitative extension of Answer Set Programming (ASP) to bridge numerical data with symbolic reasoning.

AI/ML arXiv cs.AI

How to Avoid Debate: Scalable AI Safety via Doubly-Efficient Interactive Proofs

Researchers propose single-prover interactive proofs as an alternative to debate-based verification for AI safety and alignment.

AI/ML arXiv cs.AI

The Role of Rigor in Artificial Intelligence

The authors analyze the role of rigor in AI, arguing that modern deep learning prioritizes operational rigor over conceptual and epistemic foundations.

AI/ML arXiv cs.AI

Robust Feasible Route Construction through Collaborative Partition Optimization

Collaborative Routing Constructors (CoRC) is a new framework for solving large-scale Capacitated Vehicle Routing Problems by allowing subproblems to exchange customers and vehicles.

Software Engineering Hacker News

Tiny-C Reference Manual Excerpt

An excerpt from the Tiny-C reference manual, likely focusing on language specifications or implementation details.

AI/ML arXiv cs.AI

Evaluating Generative Agents with Actions Grounded in Socially Distributed Task Environments using Incognita

Introduces Incognita, a framework for evaluating generative agents in socially distributed task environments where knowledge is partitioned among participants.

AI/ML arXiv cs.AI

Reinforcement Learning for Evidence-Seeking Diagnostic Reasoning with Large Language Models

Proposes a framework for medical diagnosis using RLVR and introduces RAGES, a retrieval-augmented examination simulator for biologically plausible feedback.