SpeechDx

★ 70

A multi-task benchmark for clinical speech AI spanning 12 datasets and 27 tasks.

AI framework integrating TE estimation and RL for treatment response optimization.

Multi-agent benchmark evaluating LLMs on strategic organizational decision-making.

A benchmark specification for reskilling-aware production-inventory control under disruptions.

MemTrace

★ 70

A benchmark for probing long-term memory in LLM agents focusing on knowledge points rather than questions.

Efficient reinforcement for visual-textual thinking using discrete diffusion models.

A multi-agent framework for feasibility-aware long-term action anticipation in video.

NeurMLLM

★ 70

A multimodal LLM framework for neurodegenerative disease screening using audio and text.

Scribby

★ 70

A multi-level LLM framework for structured semantic video analysis and summarization.

Lakebase

★ 70

A serverless cloud-based PostgreSQL database service used by Databricks for LTAP.

hts

★ 70

Agentic LLM framework for HTS code classification in maritime logistics.

A safety enhancement method for LRMs using SFT and DPO to trigger latent safety awareness.

Benchmark for evaluating LLM agents in a long-horizon multi-agent economy.

Code and datasets for studying the quality-utility paradox in mathematical reasoning distillation for SLMs.

A benchmark for evaluating multimodal spatial and temporal reasoning in urban wellbeing analytics.

A curated collection of research papers and resources regarding Embodied AI in healthcare.

OSGuard

★ 70

A benchmark for safety in computer-use agents evaluating local guardrails and end-to-end safety.

AdaTKG

★ 70

Adaptive memory implementation for temporal knowledge graph reasoning.

TimesFM

★ 70

A zero-shot foundation model for time-series forecasting.

Code for investigating the fragility of explanations in audio deepfake detection models.