DINOde

★ 70

An ODE-based framework for continuous vision-text alignment in semantic segmentation.

CertiFOX

★ 70

A framework for certified grounding in first-order logic, including the GroundFOX grounder and CheckFOX checker.

RL framework that learns value differences via an antisymmetric function to improve control.

Diagnostic benchmark for stage-wise analysis of Knowledge-Intensive Visual Question Answering.

RL-based approach to improve the faithfulness of LLM internal decision-making disclosures.

TOUR

★ 70

A Trajectory-Level Unlearning Benchmark for Offline Reinforcement Learning.

Marimo

★ 70

A reactive notebook for Python that can run in the browser or as an IDE plugin.

A modernized AWS/Terraform benchmark with 186 tasks and Rego v1 intent policies for evaluating IaC generation.

An instruction-conditioned benchmark for evaluating controllable turn management in full-duplex dialogue systems.

ODA-Data

★ 70

A high-quality paired multimodal geometry dataset for evaluating modality-dependent reasoning.

Parameter-free explainer for GNNs using granular-ball graph refinement.

HiMe

★ 70

A locally deployable agent platform for real-time personal health insights from wearables.

An open benchmark for evaluating AI-generated operator code on Huawei's Ascend NPU.

MKEvolve

★ 70

A modular framework for co-evolving complex PyTorch modules and LLM-generated kernels for hardware efficiency.

EvoSQL

★ 70

A co-evolution framework for memory-augmented Text-to-SQL synthesis via critic-generator interaction.

A Rust-based desktop environment being developed by System76.

A tool for configuring and managing encrypted storage devices on Linux.

SynSur

★ 70

A generative pipeline for synthetic industrial surface defect generation and detection using diffusion models.

TRACE

★ 70

A controlled method for evaluating LLM trust allocation across conflicting software artifacts.

A controlled benchmark and training set for selective evidence adoption in retrieval-augmented LLMs.