AI/ML arXiv cs.AI

Context Recycling for Long-Horizon LLM Inference

ContextForge is introduced as a system for context recycling in LLMs to maintain task-relevant information across long conversations while reducing token overhead.

AI/ML arXiv cs.AI

Reducing Conversational Escalation in Large Language Model Dialogue with Nonviolent Communication Constraints

Researchers demonstrate that applying Nonviolent Communication (NVC) constraints to LLM prompts can reduce conversational escalation in conflict-prone interactions.

AI/ML Hacker News

Why current LLM costs are not sustainable

A Hacker News discussion exploring the economic sustainability of current Large Language Model (LLM) cost structures.

AI/ML arXiv cs.AI

Joint Learning of Experiential Rules and Policies for Large Language Model Agents

Introduces JERP, a method for LLM agents to jointly learn experiential rules and policies from interaction trajectories to improve decision performance in complex environments.

AI/ML arXiv cs.AI

OpenRCA 2.0: From Outcome Labels to Causal Process Supervision

Presents OpenRCA 2.0 and the PAVE protocol, a new benchmark for evaluating LLM agents' ability to perform step-wise causal root cause analysis.

AI/ML arXiv cs.AI

TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference

Proposes TOPS, a training-free and model-agnostic module for efficient MLLM inference by pruning redundant visual tokens based on task relevance and semantic diversity.

AI/ML arXiv cs.AI

A Process Harness for Uplifting Legacy Workflows to Agentic BPM: Design and Realization in CUGA FLO

Introduces CUGA FLO, a framework for uplifting legacy workflows into Agentic Business Process Management (BPM) using a policy-governed agentic layer.

Cybersecurity arXiv cs.AI

Vulnerability of Natural Language Classifiers to Evolutionary Generated Adversarial Text

Presents GAversary, a hybrid Genetic Algorithm for generating effective adversarial text attacks against natural language models.

AI/ML arXiv cs.AI

Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement

Introduces BINEVAL, an interpretable LLM evaluation framework that uses atomic binary questions to provide transparent, multi-dimensional scores.

AI/ML arXiv cs.AI

EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecasting

Presents EO-WM, a video diffusion transformer designed for physically informed Earth observation forecasting using satellite imagery.

AI/ML arXiv cs.AI

Simulation-based inference for rapid Bayesian parameter estimation in epidemiological models: a comparison with MCMC

A study comparing Simulation-based inference (SBI) with MCMC for rapid Bayesian parameter estimation in epidemiological models.

Cybersecurity arXiv cs.AI

Prompt Injection in Automated R\'esum\'e Screening with Large Language Models: Single and Multi-Injection Settings

Explores the vulnerability of automated LLM-based résumé screening to prompt injection attacks by job applicants.

AI/ML arXiv cs.AI

A Pipeline for Generating Longitudinal Synthetic Clinical Notes Using Large Language Models

A new pipeline for generating longitudinal synthetic clinical notes using LLMs to provide safe, privacy-preserving data for clinical AI development.

AI/ML arXiv cs.AI

Generative Retrieval via Diffusion Transformer with Metric-Ordered Sequence Training and Hybrid-Policy Preference Optimization

Introduction of MO-DiT+HPPO, a framework for pattern-preserving attribute retrieval using Diffusion Transformers and preference optimization.

AI/ML arXiv cs.AI

Learning to Recover Task Experts from a Multi-Task Merged Model

The ReTeX framework allows for the recovery of task-specific expert performance from a single merged model checkpoint by predicting parameter offsets.

AI/ML arXiv cs.AI

Diagnosing Task Insensitivity in Language Agents

Research identifying 'task insensitivity' in LLM agents and proposing Task-Perturbed NLL Optimization to improve OOD generalization.

AI/ML arXiv cs.AI

Where Do CoT Training Gains Land in LLM based Agents?

An investigation into whether CoT training improves the reasoning process or simply the prompt-action prediction quality in LLM agents.

AI/ML arXiv cs.AI

Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds

The Look-Before-Move framework enables active visual attention and camera planning in dynamic 3D story worlds for embodied AI.

AI/ML arXiv cs.AI

Einstein World Models

Einstein World Models (EWMs) propose using visual-temporal rollouts as visual thought experiments to enhance LLM reasoning.

AI/ML arXiv cs.AI

Adaptive Utility driven Resource Orchestration for Resilient AI (AURORA-AI)

AURORA-AI is an adaptive resource orchestration framework designed to maintain resilient and fair AI deployment under non-stationary conditions.