DeepEval

★ 80

A testing framework for evaluating LLM and RAG applications.

Sentinel

★ 80

Open-source QA agent that reads code before performing interactive testing clicks.

A communication framework for collaborating AI coding agents.

A hybrid simulation framework combining spring-mass simulators with neural networks for robotic manipulation.

A coarse-to-fine framework for spatio-temporal video grounding in video streams.

Code for diagnosing visual dependency dissociation in video LLM benchmarks.

An open-source agent architecture for personalized health management.

AFIP

★ 80

An Attention-Focused Approach for Improved Image Perception to reduce hallucinations in MLLMs.

A benchmark of 1,100 primary care-to-specialist consultation cases to measure clinical safety in AI.

ReLope

★ 80

KL-Regularized LoRA Probes for Multimodal LLM Routing.

Multi-RM

★ 80

Code for evaluating reward model variants (dORM, dPRM, gORM, gPRM) across diverse domains.

A 218B parameter MoE model with an Apache 2.0 license for private enterprise deployment.

misa77

★ 80

A high-performance compression codec that decodes faster than LZ4.

Goku

★ 80

WASM (wllama)-powered LLM inference and model manager.

A tiny 18KB ls replacement written in no_std Rust.

A control framework for UAV target acquisition and tracking using NMPC and UKF.

A protocol to measure a model's internal representation of danger before response generation.

Efficient llama.cpp-based inference engine for onboard robot control.

SLEUTH

★ 80

A framework for structured epistemic working memory to scale multi-hop reasoning in language agents.

Dataset and code for nighttime agricultural visual navigation using unsupervised day-to-night translation.