AI/ML arXiv cs.AI

Simulation-based inference for rapid Bayesian parameter estimation in epidemiological models: a comparison with MCMC

A study comparing Simulation-based inference (SBI) with MCMC for rapid Bayesian parameter estimation in epidemiological models.

Cybersecurity arXiv cs.AI

Prompt Injection in Automated R\'esum\'e Screening with Large Language Models: Single and Multi-Injection Settings

Explores the vulnerability of automated LLM-based résumé screening to prompt injection attacks by job applicants.

AI/ML arXiv cs.AI

A Pipeline for Generating Longitudinal Synthetic Clinical Notes Using Large Language Models

A new pipeline for generating longitudinal synthetic clinical notes using LLMs to provide safe, privacy-preserving data for clinical AI development.

AI/ML arXiv cs.AI

Generative Retrieval via Diffusion Transformer with Metric-Ordered Sequence Training and Hybrid-Policy Preference Optimization

Introduction of MO-DiT+HPPO, a framework for pattern-preserving attribute retrieval using Diffusion Transformers and preference optimization.

AI/ML arXiv cs.AI

Learning to Recover Task Experts from a Multi-Task Merged Model

The ReTeX framework allows for the recovery of task-specific expert performance from a single merged model checkpoint by predicting parameter offsets.

AI/ML arXiv cs.AI

Diagnosing Task Insensitivity in Language Agents

Research identifying 'task insensitivity' in LLM agents and proposing Task-Perturbed NLL Optimization to improve OOD generalization.

AI/ML arXiv cs.AI

Where Do CoT Training Gains Land in LLM based Agents?

An investigation into whether CoT training improves the reasoning process or simply the prompt-action prediction quality in LLM agents.

AI/ML arXiv cs.AI

Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds

The Look-Before-Move framework enables active visual attention and camera planning in dynamic 3D story worlds for embodied AI.

AI/ML arXiv cs.AI

Einstein World Models

Einstein World Models (EWMs) propose using visual-temporal rollouts as visual thought experiments to enhance LLM reasoning.

AI/ML arXiv cs.AI

Adaptive Utility driven Resource Orchestration for Resilient AI (AURORA-AI)

AURORA-AI is an adaptive resource orchestration framework designed to maintain resilient and fair AI deployment under non-stationary conditions.

AI/ML arXiv cs.AI

Semantic Early-Stopping for Iterative LLM Agent Loops

A study on semantic early-stopping for LLM agent loops to reduce token consumption without sacrificing answer quality.

AI/ML arXiv cs.AI

How to evaluate clustering with ground truth?

A review of external validity indexes for evaluating clustering with ground truth, recommending the centroid index (CI).

AI/ML Hacker News

US Govt to individually approve who gets GPT 5.6

Rumors or discussions regarding US government intervention in the approval process for GPT-5.6 access.

Other Hacker News

22-year-old Mozart's handwritten notebook unearthed in 'major discovery'

The discovery of a handwritten music notebook by Mozart in a major archaeological/historical finding.

Tech Business/VC Hacker News

Micron locks in historically high memory prices for five years

Micron's strategy to secure high memory prices through long-term contracts.

AI/ML arXiv cs.AI

KARLA: Knowledge-base Augmented Retrieval for Language Models

Introduction of KARLA, a method for LLMs to pull factual knowledge from external databases during token generation to improve accuracy and transparency.

Other arXiv cs.AI

Computational Analysis of Heart Rate Variability in Healthy Adults

A study evaluating various Heart Rate Variability (HRV) indices to improve clinical utility and standardization for healthy adults.

AI/ML arXiv cs.AI

The Capability Frontier: Benchmarks Miss 82% of Model Performance

The 'Capability Frontier' paper argues that current benchmarks significantly underestimate LLM performance by not accounting for model specialization and optimal generation sampling.

AI/ML arXiv cs.AI

Context-Aware Synthesis of Optimization Pipelines for Warehouse Optimization

An open-source framework, CASOP, for synthesizing and evaluating optimization pipelines for warehouse order fulfillment.

AI/ML arXiv cs.AI

LCAi: Life Cycle Assessment with big data fusion and retrieval-augmented generation-assisted interpretation

LCAi framework uses RAG and big data fusion to assist in interpreting Life Cycle Assessments for environmental impact analysis.