Nimic

★ 80

Pure Python as a systems language with AOT compilation.

Feedforward model for real-time photorealistic volumetric rendering of CT scans.

Framework for generating natural language explanations of knowledge graph rules.

RV-Bench

★ 80

A novel evaluation methodology for benchmarking LLMs with Random Variables in mathematical reasoning.

HOLMES

★ 80

A benchmark for higher-order symbolic reasoning in LLMs with verifiable reasoning traces.

FlowPipe

★ 80

LLM-enhanced conditional generative flow networks for automated data pipeline construction.

Evaluation framework and datasets for measuring the interpretability of Sparse Autoencoders.

Project page for a geometric inductive bias framework for Vision-Language-Action Models.

Transformer-based multimodal framework for inferring binary contact states in robot manipulation.

Enterprise OCR model providing structured document representations and bounding boxes.

Evaluation protocol and data for studying LLM behavior under linguistic compression.

A model-agnostic orchestration harness to turn unreliable LLM solvers into reliable systems.

SurfBind

★ 80

Surface-centric learning framework for accurate molecular epitope prediction.

PV-TAM

★ 80

A metric for measuring the alignment between prompts and visual regions in VLMs.

MTO

★ 80

Framework for matching tasks to pre-training objectives to improve few-shot performance in PLMs.

Themis

★ 80

XAI-enabled testing and evaluation framework for RLHF with a cloud-based feedback collection platform.

MemClaw

★ 80

Production multi-tenant memory service for multi-agent LLM systems.

A reusable index for accelerating continuous subgraph matching via spectral pruning.

DigenRL

★ 80

Disaggregated RL framework for diffusion-based generative LLMs supporting flexible resource allocation.

NaviGen

★ 80

Framework for personalized multimodal generation from user interaction history.