Oak

★ 85

Version control system specifically designed for AI agents

A benchmark of isomorphic cross-domain science problem pairs for evaluating LLM reasoning.

GPUAlert

★ 85

A CLI wrapper for diagnosing GPU training-job failures with zero instrumentation.

A task queue for AI coding agents utilizing the Model Context Protocol.

An attack tool to estimate hidden dimension, depth, and parameter count of LLMs via restrictive API access.

Implementation of multiprobe grid algorithms for approximate nearest neighbor search in high dimensions.

RAGP

★ 85

Redundancy-Aware Graph Pruning for prompt compression using Lévy walks on multiplex graphs.

An LLM harness for improving long-context reasoning via Recursive Evidence Replay.

Fault-tolerant LLM pipeline for autonomous scientific research in computational physics.

Token-efficient code memory for repository-level program repair.

A benchmark containing 1600 executable safety cases for evaluating agentic systems.

A tool to perform DFU restores on Apple Silicon Macs.

A benchmark for grounded Earth-science reasoning using a ReAct-style executable framework.

XSkill

★ 85

Framework for continual learning from experience and skills in multimodal agents.

Scalable agentic RL training framework for deep research agents.

A comprehensive benchmark for evaluating memory-induced sycophancy in agent systems.

EchoRisk

★ 85

Dataset and benchmark for cardio-oncology echocardiography risk stratification

Actor-based concurrency framework for C++11.

A framework for depth-aware 3D semantic scene graph generation via world-model priors.

A verification pipeline to resolve and audit bibliographic entries against multiple sources to detect hallucinations.