All Articles
16640 articles total
Comparing Spectrogram Front-Ends for Abnormal Heart-Sound Detection with a Convolutional Neural Network
A study comparing spectrogram front-ends for abnormal heart-sound detection using CNNs, finding PCEN and multi-resolution spectrograms outperform standard logmel.
Fully-sensorized smart-eyewear platform for on-device Machine Learning
ARGO is a smart eyewear platform utilizing the STM32N6 NPU for on-device ML and real-time urban obstacle recognition with an optimized YOLOv11 model.
International Agreements to Limit Frontier AI: Objectives and Exit
A research paper proposing frameworks and conditions for international agreements to limit the development of frontier AI to mitigate global risks.
RouteCost: A Production-Inspired Multi-Stage Framework for Pre-Order Shipping Cost Estimation in E-Commerce
RouteCost is introduced as a multi-stage framework for improving pre-order shipping cost estimation in e-commerce through demand forecasting and residual correction.
From Weights to Words: Expressing and Editing Preference Model Inferences in Natural Language
The 'weights to words' method allows users to inspect and edit preference model inferences by translating high-dimensional choice data into natural language dimensions.
Token-Level Cross-Modal Transformer with Contrastive Multi-Task Learning for Breast Cancer Subtype Classification and Survival Prediction
A paper proposing a token-level cross-modal transformer for joint breast cancer subtype classification and survival prediction using genomic and clinical data.
HantaWatch: Federated Learning for Hantavirus Genomic Surveillance
HantaWatch is a federated learning framework designed for collaborative Hantavirus genomic surveillance without sharing raw sensitive data.
OpenMHC: Accelerating the Science of Wearable Foundation Models
The OpenMHC project releases the largest open-access wearable health dataset and open-source foundation model implementations to democratize wearable AI research.
The Failures of Marginal Influence-Based Attribution Methods for Global Time Series Explanations
Research demonstrating that common time-series attribution methods like SHAP fail to be DAG-faithful, conflating direct and mediated temporal dependencies.
The Growing Compute Shortage
A Hacker News discussion regarding the increasing scarcity of compute resources for AI development.
America needs to stop getting shocked by Chinese AI
Discussion on the competitive landscape of AI, specifically the emergence of high-performing models from Chinese startups like Moonshot.
The Shared Discovery Paradox: How a One-Answer Rule Turns Better Information into Worse Search
Research paper introducing a benchmark to analyze the 'Shared Discovery Paradox,' where pooling information into a single recommendation can reduce overall group discovery.
WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting
Introduction of WorldCupArena, a dynamic benchmark for evaluating LLMs and deep-research agents on football forecasting tasks.
Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMs
Study on evidence-sufficiency prompting in clinical LLMs, finding that safety gains in reducing overconfidence are judge-dependent and can come with helpfulness costs.
Can We Break LLMs Out of Self-Loops? Fine-Grained Reasoning Control with Activation Steering
Proposed SOPHIA, a method using hidden-state intervention and activation steering to prevent LLMs from getting stuck in reasoning self-loops.
SGA: Plug&Play Geometric Verification for Educational Video Synthesis
Introduction of the Symbolic Geometric Agent (SGA) and Manim Visual Quality Score (MVQS) to improve the spatial correctness of LLM-generated educational animations.
Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes
Research on how 'soft prefixes' can override correct logical judgments in LLMs, revealing vulnerabilities in model logical stability.
DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth
Introduction of DocOCR-Eval, an annotation-free framework for evaluating and selecting the best OCR tools for specific document collections.
What Makes Linguistic Representations Good Models of High-Level Visual Perception in the Human Brain?
Study exploring how language model embeddings of image descriptions can predict human brain responses to visual perception.
Stress Testing Concept Erasure with Large Language Model Agents
Researchers propose STACE, a framework that uses LLM agents to autonomously stress-test concept erasure in generative models to identify vulnerabilities and failure modes.