A multi-agent-driven automated pipeline for continuously evolving multimodal benchmarks.

Paper on using sparse autoencoders to analyze internal activations of Neural Quantum States.

Zep

★ 75

A system for scalable oversight of coding agents using traditional software engineering constraints.

G-RRM

★ 75

A framework for guiding symbolic solvers with symbol-equivariant recurrent reasoning models.

Uncertainty-aware multi-agent collaboration framework for reliable software development.

Bounded-memory testbed for evaluating long-horizon LLM agents.

Inspect

★ 75

An evaluation framework for testing and analyzing LLM capabilities.

A benchmark for evaluating multimodal agentic capabilities through game development tasks.

Tool for capturing natural-language dependency evidence and modeling skill-package-service graphs.

DART-VLN

★ 75

Test-time memory decay and anti-loop regularization for discrete VLN

ctx

★ 75

CLI tool to search coding agent history on local machines.

A family of autonomous driving datasets designed through a strategic gap-identification framework.

ZeroFS

★ 75

A log-structured filesystem for S3

Framework for learning gait-aware quadruped locomotion with Temporal Logic Specifications.

UPADNet

★ 75

Unrolled Phase and Amplitude Decomposition Network for improved image deblurring.

An architecture for self-evolving agents using steering adapters and verifier-in-the-loop mechanisms to prevent regressions.

Graph-native reasoning models fine-tuned with GRPO for traceable scientific hypothesis generation.

RT-SFOD

★ 75

Real-Time Source-Free Object Detection framework using YOLOv10.

JL1-CD

★ 75

A multi-task benchmark for remote sensing change detection, captioning, and QA.