SREGym

★ 90

A high-fidelity benchmark for AI Site Reliability Engineering agents with real-world failure scenarios.

Benchmark local LLMs for speed, memory, and quality

QwenLM

★ 90

The official repository for Qwen series models, including new coding and collaborative updates.

Experimental userspace to run macOS binaries on Linux ARM

Pelican

★ 90

A tool for managing LLM-based agents and prompt engineering workflows.

Open-source, ad-free Android companion client for self-hosted LLM servers like Ollama and LocalAI.

A systematic framework for technical documentation focusing on four distinct content types.

reelMind

★ 90

A CLI tool to turn Instagram reels into a local, searchable Markdown knowledge base for LLMs.

FilmOps

★ 90

A suite of Cinematic Language operators for evaluating text-to-video and reference-to-video generation.

MathNet

★ 90

A global multimodal benchmark and dataset for mathematical reasoning and retrieval.

Neural retrieval system for semantic code-to-code discovery across MediaWiki repositories.

Benchmark and evaluation harness for measuring operational stealth in AI security agents.

GPT-Red

★ 90

An automated red-teaming agent using self-play to discover prompt injection attacks.

PICOTTY

★ 90

An open-source serial console server based on Raspberry Pi Pico.

GPT-OSS

★ 90

An open-source model distilled from DeepSeek focusing on reducing censorship.

A public database for verifying the lineage, licensing, and scan status of open AI models.

RomM

★ 90

A self-hosted game library manager with save sync and server-side patching

A pipeline for identifying and using strategy-specific features to steer reasoning in LRMs.

Pictura

★ 90

A GPU-accelerated multi-agent driving simulator for perspective-view self-play.

Modus

★ 90

A decoder-only any-to-any multimodal model supporting arbitrary inputs and outputs.