All Articles
18326 articles total
QC-GAN: A Parameter-Efficient Quaternion Conformer GAN for High-Fidelity Speech Enhancement
QC-GAN is a parameter-efficient speech enhancement framework using Quaternion Conformers and MetricGAN to achieve high fidelity with significantly fewer parameters.
Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance
PhysAssistBench is a new benchmark for evaluating how LLMs assist physicians by coordinating clinical knowledge, EHR interaction, and patient communication.
AI-Driven Assessment of Human Tutors: Linking Training Performance to Real-Life Practice
A study demonstrates an AI-driven system using Gemini-2.5-pro to link human tutor training performance with real-life tutoring practice through transcription analysis.
Code-Augur: Agentic Vulnerability Detection via Specification Inference
Code-Augur introduces a security-specification-first paradigm for agentic vulnerability detection, using runtime falsification to ground LLM reasoning.
BCL: Bayesian In-Context Learning Framework for Information Extraction
BCL is a Bayesian In-Context Learning framework for information extraction that uses particle filtering to refine label representations.
EffiNav: Fusing Depth and Vision-Language for Efficient Object Goal Navigation
EffiNav fuses depth and vision-language models to improve the efficiency and robustness of object goal navigation for autonomous agents.
PEC-Home: Interpretation of Progressively Elliptical Commands in Smart Homes
PEC-Home introduces a simulated dataset to help LLM-based home assistants better interpret progressively elliptical commands in smart home dialogues.
Augmenting Dysarthric Speech Severity Assessment with MOS Supervision
This research proposes augmenting dysarthric speech assessment models using MOS supervision from speech synthesis evaluation data to overcome clinical data scarcity.
Vinyl Cache and Varnish Cache
A Hacker News discussion regarding the differences and use cases for Vinyl Cache and Varnish Cache.
Sparsity Curse: Understanding RLVR Model Parameter Space from Model Merging
Research introducing SAR-Merging, a new merging recipe for Reinforcement Learning with Verifiable Reward (RLVR) models to overcome the 'sparsity curse' and aggregate reasoning capabilities.
AI Sandboxes: A Threat Model, Taxonomy, and Measurement Framework
Proposed threat model and measurement framework for AI sandboxes to ensure safety, security, and regulatory assurance across digital and cyber-physical deployments.
Engagement Intensity as a Learner-Modeling Signal for Adaptive AI Ethics Instruction
A study evaluating how self-reported LLM usage frequency is a stronger indicator of AI perception than prior AI education when designing adaptive ethics instruction.
Correcting Sensor-Induced Distribution Drift with Wasserstein Adversarial Learning
A Wasserstein-GAN-inspired approach for unsupervised calibration of sensor-induced distribution drift to improve data-driven methods in sensor systems.
Multi-Modal Hyper-Graph Fusion for Low-Light Crowd Counting
Introduction of LCNet and the Multi-Modal Hyper-Graph Fusion module for robust crowd counting in low-light environments using RGB, depth, and edge cues.
APT: Atomic Physical Transitions for Causal Video-Language Understanding
Introduction of APT-Tune, a parameter-efficient recipe to teach VLMs to understand causal physical transitions (APTs) without losing event-level video understanding.
Dual Dimensionality for Local and Global Attention
Proposed Distance-Adaptive Representation (DAR) for Transformers, reducing KV cache dimensionality for distant tokens to maintain performance while increasing inference efficiency.
Benchmarking Action Spaces in Reinforcement Learning for Vision-based Robotic Manipulation
A benchmark study evaluating different action spaces in reinforcement learning for vision-based robotic manipulation, finding joint velocity to be most effective.
Better Adherence, Richer Context: A Field Evaluation of LLM-Powered Conversational Voice Diaries for Sleep
Field evaluation of LLM-powered conversational voice diaries for sleep, showing higher adherence and richer context compared to text-based diaries.
Hospitals and universities repurposing drugs at 90% lower cost
Hospitals and universities are finding ways to repurpose existing drugs to treat diseases at a significantly reduced cost.
TMR-GGNN: Credit Card Fraud Detection based on Time-Aware Multi-Relational Guided Graph Neural Network
Researchers propose TMR-GGNN, a time-aware multi-relational guided graph neural network for improved credit card fraud detection using heterogeneous interactions.