ProB

★ 80

A Prolog-based validation tool for Event-B formal methods.

Training-free probe for clinical efficacy and interpretability in 3D medical vision encoders.

A benchmark for evaluating AI coding agents against malicious issue requests.

A GPT-2-scale architecture with interpretable-by-design intermediate representations using Deep Parity Bottlenecks.

DuckPGQ

★ 80

A DuckDB community extension for executing graph workloads.

Full training and analysis code for investigating the mechanisms behind the Muon optimizer's grokking speed.

Moir

★ 80

A drop-in component for covariance-based knowledge editing in LLMs to prevent capability erosion.

Code and scripts for the winning argument mining system at UZH Shared Task 2026.

Converts GitHub Actions workflows to tangled workflows and back.

Buz

★ 80

A fork of Bun using modern Zig with sub-1s incremental builds

Simulation-Propose-then-OR-Dispose approach for industrial supply chain planning.

A benchmark for evaluating coding agents as interactive project builders starting from fuzzy requirements.

A three-layer reference implementation for traceable knowledge bases and agent gateways in humanistic research.

CMI-Mem

★ 80

RL-based lightweight memory manager for generalizable long-term memory in agent systems.

sign-c

★ 80

A server-side authentication scheme to protect execution bearing fields in LLM agent responses from tampering.

An intent-driven eBPF-based resource controller for managing AI agent OS resources.

Source code for PRefill attEntion STOpping (PRESTO) to improve LLM safety alignment.

HiSME

★ 80

A hierarchical skill meta-evolving framework for improving agentic system performance.

DeepEval

★ 80

An open-source testing framework for LLM and AI agent evaluation.

Langfuse

★ 80

An open-source LLM observability and evaluation tool.