ToolPro

★ 85

Represents agent tool intent as executable tool programs for flexible web services

A benchmark to measure the stamina of coding agents over 100+ interaction turns.

FlowFake

★ 85

Liquid Network architecture for efficient and robust audio deepfake detection.

SSProNet

★ 85

Secondary-structure-aware graph neural network for protein representation learning.

Source code and videos for human-like autonomy training via self-play and limited human data.

DeepSWIP

★ 85

Counterfactual reasoning for neural probabilistic logic programs via quotient-WMC.

ModSync

★ 85

Framework for modular-sparsity synchronization in conflict-averse training for generalized PINNs.

ENPIRE

★ 85

A harness framework for coding agents to automate real-world robotic policy self-improvement.

An agentic system for DeFi risk supervision using forecast-grounded LLM decisions.

A machine unlearning method for LLMs that removes private information without needing unlearning targets.

Code for evaluating token-optimized formats (TOON, TRON) in agentic AI

InfoPO

★ 85

Information-Driven Policy Optimization for User-Centric LLM Agents.

A pipeline and benchmark for capturing ripple effects of targeted interventions on language models.

An architecture for accountable behavior in LLM agents through modular separation of cognition and control.

TRAP

★ 85

Benchmark for Task-completion and Resistance to Active Privacy-extraction in AI agents.

Pipeline for synthesizing simulation-ready URDFs from RGB-D sequences.

Chezmoi

★ 85

A tool for managing dotfiles across multiple machines securely and efficiently.

A compact EEG classification architecture for low-power BCI deployment.

A manually annotated benchmark for evaluating PII redaction in LLMs across 11 domains.

fal.ai

★ 85

Developer-first, multi-model AI creative platform