Framework for scaling test-time training for LLM reasoning via online policy evolution.

GCache

★ 88

A high-performance distributed cache infrastructure with RDMA-optimized networking for LLM serving.

PalmClaw

★ 88

A native on-device agent framework for mobile phones with direct device tool access.

A grammar-constrained decoding engine for syntactically valid and policy-compliant SQL generation.

A reproduce-intervene-mitigate workbench for LLM agents over MCP

SETA-Env

★ 88

The largest open-source verifiable terminal RL dataset for training terminal agents.

A training-free method for reading out the exact concept a hidden vector encodes at any layer in an LLM.

QAgent

★ 88

An autonomous multi-agent framework for end-to-end OpenQASM code generation.

GitLake

★ 88

A system implementing Git-like versioning (commits, branches, merges) for data lakehouses.

A forensic framework for attributing malicious code completions to backdoor fine-tuning data.

TACO

★ 88

Tail-Aware Credit calibration for LLM Reinforcement Learning to suppress undesirable positive updates.

MEMCoder

★ 88

A training-free self-evolving memory framework for improving code generation in private libraries.

An open-source interactive tool for auditing AI-mediated summaries in public consultations.

An autonomous LLM multi-agent framework for automated spatiotemporal trajectory analysis in biology.

RL framework that uses self-review and policy distillation to improve LLM behavioral corrections.

Open-source RL environment for learning in non-trivial two-player fighting game scenarios.

TOFFEE

★ 88

A learned system for synthesizing data agent trajectories at scale via MCTS.

A visual-symbolic reasoning paradigm for MLLMs using executable SVG code for mental visualization.

Radicle

★ 88

P2P Git replication with native issues and patches for decentralized code collaboration.

Plotix

★ 88

Self-hosted publish and subscribe metric dashboards for graphing data via curl.