All Dev Resources
3117 resources total
LLM-wiki
★ 85A tool for increasing performance in coding harnesses utilizing LLMs.
Pegasus WMS
★ 85A widely used scientific workflow management system for scalable and reproducible execution of complex pipelines.
SafeClawBench
★ 85A staged benchmark dataset for evaluating the security of tool-using LLM agents.
DSPy
★ 85A framework for algorithmically optimizing prompts and weights in language model pipelines.
LangGraph
★ 85A library for building stateful, multi-actor applications with LLMs using graphs.
DeFAb Dataset
★ 85A verifiable benchmark for defeasible abduction in foundation models.
LiveStarPro
★ 85A live streaming assistant for proactive video understanding over long-horizon streams.
TPOUR
★ 85Implementation of Temporal Preference Optimization for Unsupervised Retrieval.
MagicSim
★ 85Embodied interaction infrastructure for robot learning using a deterministic batched runtime.
NarrativeWorldBench
★ 85An open benchmark for measuring structural narrative metrics across long horizons and multiple languages.
AMPGAN v3
★ 85Multi-objective conditional GAN for discovering non-canonical antimicrobial peptides.
TrustErase
★ 85A verifiable, data-free unlearning framework for instant and auditable forgetting in AI
Research on editable and composable KV caches to optimize LLM inference latency.
ANEForge
★ 85Python package for direct computation on the Apple Neural Engine without CoreML.
A decentralized, prefix-cache-aware routing scheme for peer-to-peer LLM serving.
A framework for learning task-level knowledge and transferring it to heterogeneous agents.
PseudoBench
★ 85An adversarial benchmark for evaluating agentic auto-research systems' resistance to pseudoscience.
hott-nesy
★ 85Implementation of homotopy type theory for neurosymbolic inference.
llm-brewing
★ 85Diagnostic framework for tracing the internal lifecycle of code reasoning in LLMs.
DeepInsight
★ 85Unified evaluation infrastructure for the physical AI stack spanning model decoding to whole-body control.