MARS

★ 70

A modular multi-agent re-ranking framework for repeat-order food delivery recommendation.

Arcade

★ 70

A secure agent runtime providing authentication, authorization, and observability for AI agents.

RIDGE

★ 70

Autonomous validation framework for LLM-generated option pricing code

A pipeline for malware perturbation assessment using EMBER feature space and PINN-style latent-flow modules.

LAFP

★ 70

LLM As Forecasting Planner: training-free text conditioning for time-series foundation models.

Dataset of structured reasoning traces for medical question answering.

A lightweight diagnostic protocol for adaptive knowledge-graph retrieval in ARK-style retrievers.

A system implementing a calculus for fidelity-graded translations for trusted AI answers.

DocAnnot

★ 70

A framework for accelerating KIE dataset generation using GenAI and OCR.

A benchmark of 321 GUI task trajectories across 10 Ubuntu desktop application categories.

Cognivia

★ 70

An evidence-based AI therapist framework for cognitive distortion recognition and rational response generation.

Research and code for parameter-efficient fine-tuning trade-offs on small language models.

Code and data for the DebtBench benchmark and DebtGPT debt collection agent.

Framework for coordinated multi-view visualization generation using LLMs.

SQBench

★ 70

A benchmark for evaluating task delivery by language-model agents in production-oriented workflows.

TopoFE

★ 70

Topology-aware LLM-guided automated feature engineering for tabular learning.

An evolving safety benchmark for foundation models targeting AI risks.

DOSA

★ 70

Framework for reconstructing document-level semantic trees from visually-rich documents.

Hubbele

★ 70

Open-source notetaking app designed for human-AI collaboration.

Goose

★ 70

An open-source harness for AI coding agents