Monty

★ 75

An autoformalization framework for synthesizing executable assertions from natural language.

TSSM

★ 75

Triaxial State Space Model for global station weather forecasting with temporal-variable-historical modeling.

CAVA

★ 75

A runtime-semantics layer for converting heterogeneous agent activity into canonical action objects for governance.

A method for non-stationary RL that adjusts exploration intensity based on observable drift proxies.

A benchmark for evaluating procedural competence and 3D-aware synthesis in multimodal models.

A synchronized memory layer for AI coding assistants over SSH.

A research paper and public dataset analyzing LLM answer-choice conformity.

Hallo4D

★ 75

A framework for mitigating spatiotemporal hallucinations in 3D and 4D content generation.

Asynchronous algorithm for real-time robot control via VLA models.

Forgetting-resilient solution for updating fake speech detectors to new datasets.

Code-MUE

★ 75

A black-box framework for measuring uncertainty in Code LLMs using runtime behavior.

A human-validated benchmark for end-to-end retrieve-and-synthesize pipelines over data lakes.

RCWT

★ 75

A measurement protocol for evaluating task-budget displacement in multi-agent LLM context windows.

JoPMol

★ 75

A jointly controlled precision molecular generative model integrating biological states and molecular structures.

PyBaMM

★ 75

A Python battery mathematical modeling framework for simulating battery degradation.

pyDVL

★ 75

A toolkit for implementing and evaluating Data Valuation methods.

SeqGPT

★ 75

A constrained Transformer agent for the inverse design of multi-panel composite structures.

A counterfactual report-coordinate clamp to neutralize incentive pressure in LLM reports.

MemOps

★ 75

A benchmark for evaluating lifecycle memory operations in long-horizon LLM conversations.

Elenchos

★ 75

Generative evaluation framework for measuring abductive reasoning in LLMs.