AI/ML arXiv cs.AI

Comparing Spectrogram Front-Ends for Abnormal Heart-Sound Detection with a Convolutional Neural Network

A study comparing spectrogram front-ends for abnormal heart-sound detection using CNNs, finding PCEN and multi-resolution spectrograms outperform standard logmel.

Hardware/Chips arXiv cs.AI

Fully-sensorized smart-eyewear platform for on-device Machine Learning

ARGO is a smart eyewear platform utilizing the STM32N6 NPU for on-device ML and real-time urban obstacle recognition with an optimized YOLOv11 model.

AI/ML arXiv cs.AI

International Agreements to Limit Frontier AI: Objectives and Exit

A research paper proposing frameworks and conditions for international agreements to limit the development of frontier AI to mitigate global risks.

AI/ML arXiv cs.AI

RouteCost: A Production-Inspired Multi-Stage Framework for Pre-Order Shipping Cost Estimation in E-Commerce

RouteCost is introduced as a multi-stage framework for improving pre-order shipping cost estimation in e-commerce through demand forecasting and residual correction.

AI/ML arXiv cs.AI

From Weights to Words: Expressing and Editing Preference Model Inferences in Natural Language

The 'weights to words' method allows users to inspect and edit preference model inferences by translating high-dimensional choice data into natural language dimensions.

AI/ML arXiv cs.AI

Token-Level Cross-Modal Transformer with Contrastive Multi-Task Learning for Breast Cancer Subtype Classification and Survival Prediction

A paper proposing a token-level cross-modal transformer for joint breast cancer subtype classification and survival prediction using genomic and clinical data.

AI/ML arXiv cs.AI

HantaWatch: Federated Learning for Hantavirus Genomic Surveillance

HantaWatch is a federated learning framework designed for collaborative Hantavirus genomic surveillance without sharing raw sensitive data.

Open Source arXiv cs.AI

OpenMHC: Accelerating the Science of Wearable Foundation Models

The OpenMHC project releases the largest open-access wearable health dataset and open-source foundation model implementations to democratize wearable AI research.

AI/ML arXiv cs.AI

The Failures of Marginal Influence-Based Attribution Methods for Global Time Series Explanations

Research demonstrating that common time-series attribution methods like SHAP fail to be DAG-faithful, conflating direct and mediated temporal dependencies.

AI/ML Hacker News

The Growing Compute Shortage

A Hacker News discussion regarding the increasing scarcity of compute resources for AI development.

Tech Business/VC The Verge

America needs to stop getting shocked by Chinese AI

Discussion on the competitive landscape of AI, specifically the emergence of high-performing models from Chinese startups like Moonshot.

AI/ML arXiv cs.AI

The Shared Discovery Paradox: How a One-Answer Rule Turns Better Information into Worse Search

Research paper introducing a benchmark to analyze the 'Shared Discovery Paradox,' where pooling information into a single recommendation can reduce overall group discovery.

AI/ML arXiv cs.AI

WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting

Introduction of WorldCupArena, a dynamic benchmark for evaluating LLMs and deep-research agents on football forecasting tasks.

AI/ML arXiv cs.AI

Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMs

Study on evidence-sufficiency prompting in clinical LLMs, finding that safety gains in reducing overconfidence are judge-dependent and can come with helpfulness costs.

AI/ML arXiv cs.AI

Can We Break LLMs Out of Self-Loops? Fine-Grained Reasoning Control with Activation Steering

Proposed SOPHIA, a method using hidden-state intervention and activation steering to prevent LLMs from getting stuck in reasoning self-loops.

AI/ML arXiv cs.AI

SGA: Plug&Play Geometric Verification for Educational Video Synthesis

Introduction of the Symbolic Geometric Agent (SGA) and Manim Visual Quality Score (MVQS) to improve the spatial correctness of LLM-generated educational animations.

AI/ML arXiv cs.AI

Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes

Research on how 'soft prefixes' can override correct logical judgments in LLMs, revealing vulnerabilities in model logical stability.

AI/ML arXiv cs.AI

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth

Introduction of DocOCR-Eval, an annotation-free framework for evaluating and selecting the best OCR tools for specific document collections.

AI/ML arXiv cs.AI

What Makes Linguistic Representations Good Models of High-Level Visual Perception in the Human Brain?

Study exploring how language model embeddings of image descriptions can predict human brain responses to visual perception.

AI/ML arXiv cs.AI

Stress Testing Concept Erasure with Large Language Model Agents

Researchers propose STACE, a framework that uses LLM agents to autonomously stress-test concept erasure in generative models to identify vulnerabilities and failure modes.