All Articles
18687 articles total
A Machine-Learned Comorbidity Index
Proposal of a Machine-Learned Comorbidity Index (MLCI) that uses nHSIC to better capture nonlinear risk-outcome relationships in clinical data.
MapSatisfyBench: Benchmarking Satisfaction-Aware Map Agents through Behavior-Grounded Implicit Decision Factors
Introduction of MapSatisfyBench, a benchmark to evaluate how well map agents can identify and satisfy implicit user decision factors.
Dissecting model behavior through agent trajectories
Research on the 'intent-execution gap' in AI agents, introducing 'Simple Strands Agent' (SSA) to analyze and improve model-harness alignment.
Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation
CEO-Bench, a new benchmark evaluating LLMs' ability to handle strategic resource reallocation and synthesize conflicting C-suite advice.
Stop Killing Games fails to secure EU law despite 1.3M signatures
An initiative to prevent the permanent deletion of games failed to secure new EU laws despite significant public support.
The Amphibious Villagers of Indonesia
A feature on the amphibious villagers of Indonesia.
Leaked OpenAI financials show $38.5B loss and compute burn
Leaked financial documents reveal that OpenAI has incurred massive losses and significant compute costs.
Beyond Parallel Sampling: Diverse Query Initialization for Agentic Search
Researchers introduce DivInit, a training-free intervention that improves agentic search by ensuring diverse initial queries to reduce redundancy.
When Rules Learn: A Self-Evolving Agent for Legal Case Retrieval
A new self-evolving framework for rule-driven query rewriting is proposed to enhance BM25 retrieval for legal case search without parameter training.
SkillChain-Gym: A Benchmark for Reskilling-Aware Production-Inventory Control under Disruptions
SkillChain-Gym is introduced as a benchmark for production-inventory control that accounts for workforce skill decay and reskilling needs.
Skill-Constrained Model Predictive Control for Resilient Manufacturing Supply Chains
The study evaluates a closed-loop model predictive controller for resilient manufacturing supply chains constrained by worker skill levels using the SkillChain-Gym benchmark.
Nothing from Something: Can a Language Model Discover 0?
Research examines whether LLMs can discover the concept of 'zero' through mathematical discovery and out-of-distribution generalization.
Quantifying Consistency in LLM Logical Reasoning via Structural Uncertainty
A new framework called structural uncertainty is proposed to quantify the consistency of LLM logical reasoning using self-preference rankings.
MemTrace: Probing What Final Accuracy Misses in Long-Term Memory
MemTrace is introduced as a benchmark to probe long-term memory in LLM agents, revealing that evidence utilization is a larger bottleneck than retrieval.
All about the IBM 1130 Computing System
A discussion on Hacker News regarding the IBM 1130 Computing System, detailing its technical specifications and history.
An Ensemble Deep Learning Approach for Reliable and Scalable Lemon Leaf Disease Classification
Research on using an ensemble of InceptionV3 and MobileNetV2 models to classify lemon leaf diseases with high accuracy.
Improved Knowledge Distillation for Land-Use Image Classification
A proposed Knowledge Distillation framework to compress VGG16 knowledge into a lightweight MobileNetV2 for land-use image classification.
Mask Proposal Voting Based on Geodesic Framework for Robust Image Segmentation
A novel mask proposal voting framework based on a geodesic framework to improve robustness in image segmentation tasks.
An Empirical Study on Learning Latent Representations for Emotional Speech Synthesis
An empirical study on enhancing emotional speech synthesis (ESS) by integrating speaker embedding and prosody bottlenecks into FastSpeech 2.
Policy Regret for Embedding Model Routing: Contextual Bandits with Low-Rank Experts
A formalization of embedding model routing as an adversarial contextual linear bandit and the introduction of the Hypentropy Policy Gradient (HPG) algorithm.