A standardized MRI dataset and diagnostic benchmark for spine pathology diagnosis.

Project page for hybrid offline-online RL for Vision-Language-Action models.

SFgen

★ 70

Agentic recognition and generation flow for electronic component symbols and footprints.

ascdraw

★ 70

High-performance editor for ASCII/UTF-8 diagrams

A demonstration of a field-scoped epistemic grounding document for proteomics to guide AI coding.

QuArch

★ 70

A benchmark and leaderboard for evaluating LLM reasoning in computer architecture.

A comprehensive benchmark for multi-application navigation and cross-interface coherence in GUI agents.

MIRA-Ev

★ 70

A clinical argument mining benchmark for evidence detection and relational reasoning.

Diagnosis-decoupled evaluation framework for multi-turn medical consultation agents.

A method for concept-based counterfactual generation to improve explainability in time series AI models.

Implementation of a multi-stage marine species detection system using YOLO and DINOv3.

Code for interpreting how Transformer models encode causation and antithesis discourse relations.

A public dataset for evaluating VLM safety reasoning by distinguishing hazards from anomalies.

An LLM-guided multi-agent system for reasoning over genome-scale constraint-based metabolic models.

MAGE

★ 70

Multimodal multi-agent framework for macro placement refinement in industrial physical design flows.

OpenROAD

★ 70

Open-source digital flow for RTL-to-GDSII implementation.

A large-scale suite of realistic HVAC control environments for RL benchmark tasks

An optimal browser harness for AI agents to automate web interactions.

Source code and data for benchmarking optimisers for configurable systems.

PRISM

★ 70

Multimodal perception system for terrain mapping in unstructured environments using RGB-D-T imagery.