All Dev Resources
3152 resources total
SpeechDx
★ 70A multi-task benchmark for clinical speech AI spanning 12 datasets and 27 tasks.
AI framework integrating TE estimation and RL for treatment response optimization.
CEO-Bench
★ 70Multi-agent benchmark evaluating LLMs on strategic organizational decision-making.
SkillChain-Gym
★ 70A benchmark specification for reskilling-aware production-inventory control under disruptions.
MemTrace
★ 70A benchmark for probing long-term memory in LLM agents focusing on knowledge points rather than questions.
Efficient reinforcement for visual-textual thinking using discrete diffusion models.
FactCheck
★ 70A multi-agent framework for feasibility-aware long-term action anticipation in video.
NeurMLLM
★ 70A multimodal LLM framework for neurodegenerative disease screening using audio and text.
Scribby
★ 70A multi-level LLM framework for structured semantic video analysis and summarization.
Lakebase
★ 70A serverless cloud-based PostgreSQL database service used by Databricks for LTAP.
hts
★ 70Agentic LLM framework for HTS code classification in maritime logistics.
Safe Trigger
★ 70A safety enhancement method for LRMs using SFT and DPO to trigger latent safety awareness.
CoffeeBench
★ 70Benchmark for evaluating LLM agents in a long-horizon multi-agent economy.
Code and datasets for studying the quality-utility paradox in mathematical reasoning distillation for SLMs.
UrbanWell-Benchmark
★ 70A benchmark for evaluating multimodal spatial and temporal reasoning in urban wellbeing analytics.
A curated collection of research papers and resources regarding Embodied AI in healthcare.
OSGuard
★ 70A benchmark for safety in computer-use agents evaluating local guardrails and end-to-end safety.
AdaTKG
★ 70Adaptive memory implementation for temporal knowledge graph reasoning.
TimesFM
★ 70A zero-shot foundation model for time-series forecasting.
Audio-XAI
★ 70Code for investigating the fragility of explanations in audio deepfake detection models.