A training-free post-hoc score to measure grounding confidence in MLLMs and reduce hallucinations.

SSC-Loop

★ 85

A unified framework for signed social recommendation based on structural consistency maximization.

LLM powered and AST-based tool to visualize codebase

TLDraw

★ 85

A collaborative whiteboard library and canvas for building drawing applications.

Replication package for AI agent performance mutation analysis.

BaFCo

★ 85

A benchmark dataset for Bangla form comprehension focusing on Document Layout Analysis and Key Information Extraction.

Training-free pipeline for segmenting 3D objects into geometric primitives using VLMs.

cpc-3gpp

★ 85

Source code for Contrastive Predictive Coding implementation for 3GPP-compliant CSI feedback.

Danus

★ 85

Orchestration system for research-level mathematical reasoning agents using fact-graph memory.

A benchmark for evaluating LLM agents on multilingual long-horizon workplace workflows.

PCBWorld

★ 85

Open-source engine-grounded PCB routing environment built on the KiCad EDA engine.

NapMem

★ 85

A framework for learning to use long-term user memory as a structured action space

A first-principles research-question formation framework for scientific LLM agents.

A multi-agent adversarial benchmark for assessing rule adherence in semi-open textual sandboxes.

Microservices for high-performance AI deployment.

Benchmark for evaluating agent self-evolution via Ability-guided transfer across multiple agentic domains.

AssemCAD

★ 85

Axiom-grounded framework for production-ready CAD assembly generation from natural language.

Implementation of belief-rollout diagnostics and BIWM protocol for LLM agent evaluation.

Autonomous trading agent for prediction markets designed to bridge the gap between forecasting and trading.

REDI

★ 85

An open-source framework for automated data readiness for scientific AI.