LLM-wiki

★ 85

A tool for increasing performance in coding harnesses utilizing LLMs.

A widely used scientific workflow management system for scalable and reproducible execution of complex pipelines.

A staged benchmark dataset for evaluating the security of tool-using LLM agents.

DSPy

★ 85

A framework for algorithmically optimizing prompts and weights in language model pipelines.

A library for building stateful, multi-actor applications with LLMs using graphs.

A verifiable benchmark for defeasible abduction in foundation models.

A live streaming assistant for proactive video understanding over long-horizon streams.

TPOUR

★ 85

Implementation of Temporal Preference Optimization for Unsupervised Retrieval.

MagicSim

★ 85

Embodied interaction infrastructure for robot learning using a deterministic batched runtime.

An open benchmark for measuring structural narrative metrics across long horizons and multiple languages.

Multi-objective conditional GAN for discovering non-canonical antimicrobial peptides.

A verifiable, data-free unlearning framework for instant and auditable forgetting in AI

Research on editable and composable KV caches to optimize LLM inference latency.

ANEForge

★ 85

Python package for direct computation on the Apple Neural Engine without CoreML.

A decentralized, prefix-cache-aware routing scheme for peer-to-peer LLM serving.

A framework for learning task-level knowledge and transferring it to heterogeneous agents.

An adversarial benchmark for evaluating agentic auto-research systems' resistance to pseudoscience.

Implementation of homotopy type theory for neurosymbolic inference.

Diagnostic framework for tracing the internal lifecycle of code reasoning in LLMs.

Unified evaluation infrastructure for the physical AI stack spanning model decoding to whole-body control.