All Dev Resources
3122 resources total
Langfuse
★ 80An open-source LLM observability and evaluation tool.
Echo
★ 80A tool for achieving high-quality AI results using cost-effective open-weight models.
Whetuu
★ 80A zero-config cross-shell prompt written in Zig
SciTrek
★ 80A synthetic QA dataset for long-context numerical reasoning over scientific articles.
DynamicRubric
★ 80A response-set-conditioned evaluator-policy co-evolution framework for LLMs.
HIJACKKV
★ 80Attack framework for exploiting position-independent KV cache reuse in LLMs.
JailMeter
★ 80An evidence-based evaluation framework for measuring LLM jailbreak attack success rates.
Prefix-GRPO
★ 80RL framework for decomposing teacher trajectories into replay-aligned prefixes and online continuations.
Diagnostic pipeline for evaluating reliability and silent failures in multimodal agentic search trajectories.
ITPEval
★ 80A multi-ITP verification infrastructure and benchmark for formal proof translation.
Fly0
★ 80A framework for zero-shot aerial vision-language navigation decoupling reasoning from planning.
Framework for dynamic expert clustering and structured compression in MoE LLMs.
JIUTIAN-TReB Dataset
★ 80A high-quality dataset for evaluating LLM table reasoning abilities across 26 sub-tasks.
TReB Framework
★ 80An evaluation framework to robustly measure table reasoning capabilities in LLMs.
MEDIC-Benchmark
★ 80A leaderboard and evaluation framework for clinical LLM competence
Codebase for analyzing how LLMs organize action and predicate concepts during self-generated reasoning.
FieldFinder
★ 80A tool to assist LLMs in identifying text placement on forms for better filling accuracy.
PathAgentBench
★ 80A benchmark for evaluating evidence-seeking vision-language models across pathology image analysis tasks.
Great Expectations
★ 80A tool for validating, documenting, and profiling your data to ensure correctness.
DeepSQL
★ 80A self-hostable AI DBA agent for Postgres and MySQL