Framework for evaluating and improving agents
-
Updated
Sep 1, 2026 - Python
Framework for evaluating and improving agents
A harness optimized to smaller LLMs
Long Horizon Terminal Benchmark with Dense Reward Grading
Official Implementation of "CLI-Gym: Scalable CLI Task Generation via Agentic Environment Inversion"
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Some tools help. Some tools assist. Villani Code intervenes.
A research framework for principled agent self-improvement under frozen evaluators and declared mutation boundaries, recording verifiable lineage to make it reproducible and auditable.
Strands-based agents and harnesses for agentic benchmarks.
Simple Long Horizon Agent - A simple yet effective AI agent for learning, experimentation, and long horizon work.
Spoox CLI - Terminal Agent - SPlit lOOp eXand agent
Fast, Multi-Cloud Sandbox Engine for AI Agents
A coding agent that builds it's own tools and mines it's own skills and always ships with receipts.
Evidence-first, crash-resumable self-evolution engine for DeepSeek Harness and Harbor.
Trajectories for running OpenHands on Terminal Bench
Backtest every model and provider switch on your own traffic. Local-first evals from production traces: money-safe by default, sealed-partition gates, verdicts with confidence intervals.
OpenCode TUI plugin for model benchmark scores, token costs, and token efficiency.
Self-owned TerminalBench 2.0 coding-agent harness with a Rust Worker, Harbor/verifier-grounded evaluation, and guarded Heuristic Learning.
Your model, many harnesses, many benchmarks.
Open-source desktop coding agent (Electron + React) with plan/execute modes, MCP, skills, and a terminal-bench 2.0 harness adapter
Central repository for Project Terminus task submissions. These environments and Oracle solutions are designed to evaluate state-of-the-art AI agents (like GPT-5.2 and Claude Opus 4.6) on complex, multi-step engineering challenges within a sandboxed terminal.
To associate your repository with the terminal-bench topic, visit your repo's landing page and select "manage topics."