Enterprise-grade evaluation and observability platform for AI agents.

Secure AI application deployment for individual devices.

UCB-HARE

★ 70

An algorithm using inverse-weighted harmonic rank exploration for fair bandit problems.

A co-evolutionary discrete diffusion framework for gene regulatory network inference.

A benchmark dataset for evaluating AI-driven STEM lesson generation agents.

A method for correcting reasoning errors in LLMs via direct CoT editing.

A 26-week supply-chain replenishment benchmark for evaluating LLM agents in partially observable environments.

A framework for agent failure recovery using action decision graphs.

Code for a framework explaining the mechanics of data mixing in neural scaling laws.

A tool for coding agents to prevent redundant grep operations by caching search results.

Coasty

★ 70

An API for automating computer interaction via AI agents.

Aict

★ 70

Unix coreutils that output XML/JSON, built for AI agents.

A benchmark modeling multiple farms coordinating battery storage under stochastic demand.

Grepathy

★ 70

A tool for analyzing and tracking AI-driven decision making and reasoning paths.

TRAIL

★ 70

A platform for configurable human-AI teaming experiments focusing on agent personas and coordination.

HGNP

★ 70

A human-inspired adaptive exploration-exploitation framework for Genetic Network Programming.

NameRank

★ 70

A recognition score and framework for measuring entity recall in LLM weights.

Clinical AI system for continuous musculoskeletal management using LLMs and hospital data.

TRACE

★ 70

A typed, versioned schema for recording reasoning traces and agentic commitments.

Implementation of Gaussian and Quantile uncertainty quantification for EO regression tasks.