AI/ML arXiv cs.AI

A Machine-Learned Comorbidity Index

Proposal of a Machine-Learned Comorbidity Index (MLCI) that uses nHSIC to better capture nonlinear risk-outcome relationships in clinical data.

AI/ML arXiv cs.AI

MapSatisfyBench: Benchmarking Satisfaction-Aware Map Agents through Behavior-Grounded Implicit Decision Factors

Introduction of MapSatisfyBench, a benchmark to evaluate how well map agents can identify and satisfy implicit user decision factors.

AI/ML arXiv cs.AI

Dissecting model behavior through agent trajectories

Research on the 'intent-execution gap' in AI agents, introducing 'Simple Strands Agent' (SSA) to analyze and improve model-harness alignment.

AI/ML arXiv cs.AI

Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation

CEO-Bench, a new benchmark evaluating LLMs' ability to handle strategic resource reallocation and synthesize conflicting C-suite advice.

Other Hacker News

Stop Killing Games fails to secure EU law despite 1.3M signatures

An initiative to prevent the permanent deletion of games failed to secure new EU laws despite significant public support.

Other Hacker News

The Amphibious Villagers of Indonesia

A feature on the amphibious villagers of Indonesia.

Tech Business/VC Hacker News

Leaked OpenAI financials show $38.5B loss and compute burn

Leaked financial documents reveal that OpenAI has incurred massive losses and significant compute costs.

AI/ML arXiv cs.AI

Beyond Parallel Sampling: Diverse Query Initialization for Agentic Search

Researchers introduce DivInit, a training-free intervention that improves agentic search by ensuring diverse initial queries to reduce redundancy.

AI/ML arXiv cs.AI

When Rules Learn: A Self-Evolving Agent for Legal Case Retrieval

A new self-evolving framework for rule-driven query rewriting is proposed to enhance BM25 retrieval for legal case search without parameter training.

AI/ML arXiv cs.AI

SkillChain-Gym: A Benchmark for Reskilling-Aware Production-Inventory Control under Disruptions

SkillChain-Gym is introduced as a benchmark for production-inventory control that accounts for workforce skill decay and reskilling needs.

AI/ML arXiv cs.AI

Skill-Constrained Model Predictive Control for Resilient Manufacturing Supply Chains

The study evaluates a closed-loop model predictive controller for resilient manufacturing supply chains constrained by worker skill levels using the SkillChain-Gym benchmark.

AI/ML arXiv cs.AI

Nothing from Something: Can a Language Model Discover 0?

Research examines whether LLMs can discover the concept of 'zero' through mathematical discovery and out-of-distribution generalization.

AI/ML arXiv cs.AI

Quantifying Consistency in LLM Logical Reasoning via Structural Uncertainty

A new framework called structural uncertainty is proposed to quantify the consistency of LLM logical reasoning using self-preference rankings.

AI/ML arXiv cs.AI

MemTrace: Probing What Final Accuracy Misses in Long-Term Memory

MemTrace is introduced as a benchmark to probe long-term memory in LLM agents, revealing that evidence utilization is a larger bottleneck than retrieval.

Software Engineering Hacker News

All about the IBM 1130 Computing System

A discussion on Hacker News regarding the IBM 1130 Computing System, detailing its technical specifications and history.

AI/ML arXiv cs.AI

An Ensemble Deep Learning Approach for Reliable and Scalable Lemon Leaf Disease Classification

Research on using an ensemble of InceptionV3 and MobileNetV2 models to classify lemon leaf diseases with high accuracy.

AI/ML arXiv cs.AI

Improved Knowledge Distillation for Land-Use Image Classification

A proposed Knowledge Distillation framework to compress VGG16 knowledge into a lightweight MobileNetV2 for land-use image classification.

AI/ML arXiv cs.AI

Mask Proposal Voting Based on Geodesic Framework for Robust Image Segmentation

A novel mask proposal voting framework based on a geodesic framework to improve robustness in image segmentation tasks.

AI/ML arXiv cs.AI

An Empirical Study on Learning Latent Representations for Emotional Speech Synthesis

An empirical study on enhancing emotional speech synthesis (ESS) by integrating speaker embedding and prosody bottlenecks into FastSpeech 2.

AI/ML arXiv cs.AI

Policy Regret for Embedding Model Routing: Contextual Bandits with Low-Rank Experts

A formalization of embedding model routing as an adversarial contextual linear bandit and the introduction of the Hypentropy Policy Gradient (HPG) algorithm.