Langfuse

★ 80

An open-source LLM observability and evaluation tool.

Echo

★ 80

A tool for achieving high-quality AI results using cost-effective open-weight models.

Whetuu

★ 80

A zero-config cross-shell prompt written in Zig

SciTrek

★ 80

A synthetic QA dataset for long-context numerical reasoning over scientific articles.

A response-set-conditioned evaluator-policy co-evolution framework for LLMs.

HIJACKKV

★ 80

Attack framework for exploiting position-independent KV cache reuse in LLMs.

An evidence-based evaluation framework for measuring LLM jailbreak attack success rates.

RL framework for decomposing teacher trajectories into replay-aligned prefixes and online continuations.

Diagnostic pipeline for evaluating reliability and silent failures in multimodal agentic search trajectories.

ITPEval

★ 80

A multi-ITP verification infrastructure and benchmark for formal proof translation.

Fly0

★ 80

A framework for zero-shot aerial vision-language navigation decoupling reasoning from planning.

Framework for dynamic expert clustering and structured compression in MoE LLMs.

A high-quality dataset for evaluating LLM table reasoning abilities across 26 sub-tasks.

An evaluation framework to robustly measure table reasoning capabilities in LLMs.

A leaderboard and evaluation framework for clinical LLM competence

Codebase for analyzing how LLMs organize action and predicate concepts during self-generated reasoning.

A tool to assist LLMs in identifying text placement on forms for better filling accuracy.

A benchmark for evaluating evidence-seeking vision-language models across pathology image analysis tasks.

A tool for validating, documenting, and profiling your data to ensure correctness.

DeepSQL

★ 80

A self-hostable AI DBA agent for Postgres and MySQL