All Dev Resources
3143 resources total
DINOde
★ 70An ODE-based framework for continuous vision-text alignment in semantic segmentation.
CertiFOX
★ 70A framework for certified grounding in first-order logic, including the GroundFOX grounder and CheckFOX checker.
RL framework that learns value differences via an antisymmetric function to improve control.
CRAG-MM-Diagnostics
★ 70Diagnostic benchmark for stage-wise analysis of Knowledge-Intensive Visual Question Answering.
RL-based approach to improve the faithfulness of LLM internal decision-making disclosures.
TOUR
★ 70A Trajectory-Level Unlearning Benchmark for Offline Reinforcement Learning.
Marimo
★ 70A reactive notebook for Python that can run in the browser or as an IDE plugin.
IaC-Eval v2
★ 70A modernized AWS/Terraform benchmark with 186 tasks and Rego v1 intent policies for evaluating IaC generation.
Instruct-FD
★ 70An instruction-conditioned benchmark for evaluating controllable turn management in full-duplex dialogue systems.
ODA-Data
★ 70A high-quality paired multimodal geometry dataset for evaluating modality-dependent reasoning.
SeeExplainer
★ 70Parameter-free explainer for GNNs using granular-ball graph refinement.
HiMe
★ 70A locally deployable agent platform for real-time personal health insights from wearables.
CANN Bench
★ 70An open benchmark for evaluating AI-generated operator code on Huawei's Ascend NPU.
MKEvolve
★ 70A modular framework for co-evolving complex PyTorch modules and LLM-generated kernels for hardware efficiency.
EvoSQL
★ 70A co-evolution framework for memory-augmented Text-to-SQL synthesis via critic-generator interaction.
COSMIC Epoch
★ 70A Rust-based desktop environment being developed by System76.
Cryptsetup
★ 70A tool for configuring and managing encrypted storage devices on Linux.
SynSur
★ 70A generative pipeline for synthetic industrial surface defect generation and detection using diffusion models.
TRACE
★ 70A controlled method for evaluating LLM trust allocation across conflicting software artifacts.
SelectBench
★ 70A controlled benchmark and training set for selective evidence adoption in retrieval-augmented LLMs.