A kernel-level LLM inference simulator for predicting token-level execution and latency.

A benchmark for evaluating general-purpose terminal-use agents across 120 real-world tasks.

Platform to access and prototype with the new Gemini 3.1 Flash-Lite Image model.

Environment-free agentic patch verifier for evaluating generated code patches without execution.

Reasoning-augmented framework for PET/CT lesion segmentation using LLMs and Reinforcement Learning.

A long-horizon coding benchmark for reimplementing software projects from behavior.

State-level experience system for on-demand retrieval to enhance LLM reasoning.

D2R-RAG

★ 80

Model-agnostic framework for diagnosing and repairing factual errors in RAG under budget constraints.

Kueue

★ 80

A Kubernetes-native job queueing controller for managing batch workloads.

GRACE

★ 80

Framework for unifying knowledge distillation and QAT for efficient Vision-Language Models.

Open-source Privileged Access Management system for secure server and network device access.

Benchmark for evaluating over-refusal and safe completion of medical AI assistants.

ADC-GNN

★ 80

Attention-guided Diffusion-Contrastive Graph Neural Network for fraud detection.

HO-FNO

★ 80

Higher-Order Fourier Neural Operator for explicit mode mixing in nonlinear PDEs.

HipNet

★ 80

A hippocampal memory network module for the DETR object detection architecture.

ProbeUE

★ 80

Code repository for probe-based uncertainty estimation to detect hallucinations in LLMs.

LaViD

★ 80

Language-to-Visual Knowledge Distillation framework for transfer of fine-grained conceptual knowledge.

Ante

★ 80

A language/system combining borrow checking and reference counting for memory safety.

A privacy-first, self-hosted alternative to Google Analytics.

Sinema

★ 80

An open-source Android TV client for Stash with PIN lock and D-pad navigation.