Overview
Oqoqo is an evaluation and benchmarking platform for AI agents. Its official site describes a workflow for turning real work into task sets, rubrics, experiments, and repeatable runs, then measuring whether agents can complete those tasks in realistic environments.
The product is built for teams that need more than one-off manual testing. Oqoqo lets users define tasks, select agents, compare treatments such as raw agent versus MCP-assisted workflows, run each trial in a sandbox, and inspect pass/fail results with trajectories, token usage, frictions, and experiment-level lift.
For ToolWorthy readers, the most useful angle is agent product evaluation. If you are building MCP servers, CLIs, SDKs, docs, or internal tools for agents to use, Oqoqo provides a way to test whether AI agent workflows actually succeed before you ship changes.
Key Features
Private agent benchmarks - Build custom task sets and rubrics around your own product, workflow, API, CLI, or documentation.
Realistic sandbox runs - Each trial runs in its own environment with project state, files, tools, context, and credentials configured for the task.
Agent and model comparison - Run the same tasks across Claude Code, Codex, Cursor, GitHub Copilot, OpenCode, OpenClaw, Pi, Hermes, and supported model settings.
Treatment testing - Compare raw agents against changes such as MCP servers, skills, CLIs, SDKs, best-practice instructions, or product interface updates.
Trajectory inspection - Review commands, tool calls, file changes, errors, output, pass/fail reasons, token usage, and friction patterns for each run.
Web, CLI, and MCP access - Use the dashboard directly, or let agents launch and inspect experiments through CLI and MCP workflows.
How to Get Started
Start by choosing one task that a real user or agent must complete, such as creating an object in an API, editing a file, using an MCP server, or following a documented workflow. Define the instructions and rubric, then run a small experiment with one or two agents before expanding the benchmark.
Oqoqo is strongest when the task has an observable outcome. A good first benchmark should have clear success criteria, fixture data, and a repeatable environment. Avoid vague prompts such as "use this product well" until you have concrete tasks and rubrics.
Teams building AI workflow generator tools can also use Oqoqo as a regression harness. When an MCP server, CLI, SDK, or documentation update changes the agent interface, re-run the same experiment and compare pass rates, token usage, and friction notes.
Pricing & Plans
Oqoqo has a free plan with 100 plan runs per month. The Pro plan starts at $20 per month for 300 plan runs, and the Ultra plan is $60 per month for 1,000 plan runs.
The pricing page says every plan includes all features, unlimited team members and projects, unlimited experiments, tasks, treatments, and assets, plus web app, CLI, and MCP access. Runs cover Oqoqo platform fees and cloud infrastructure; model inference is separate, so users bring their own model keys or subscriptions. Top-ups start at $5 for 25 runs.
Best For
- Product teams testing whether agents can use their product
- Developers building MCP servers, CLIs, SDKs, and agent-readable docs
- Engineering teams comparing Claude Code, Codex, Cursor, GitHub Copilot, and related agents
- AI infrastructure teams that need private benchmarks rather than public leaderboard-only evals
- Operators who want evidence before rolling out agent workflow changes
FAQ
What is Oqoqo?
Oqoqo is a platform for building evals and private benchmarks for real-world AI agent tasks.
What can Oqoqo evaluate?
The official docs say Oqoqo can evaluate skills, MCP servers, CLIs, SDKs, APIs, docs, files, and workflows that agents need to use.
Which agents does Oqoqo support?
The official site lists Claude Code, Codex, Cursor, GitHub Copilot, OpenCode, OpenClaw, Pi, and Hermes, with more agents planned.
Is there a free Oqoqo plan?
Yes. The free plan includes 100 plan runs per month.
How much does Oqoqo Pro cost?
The Pro plan starts at $20 per month for 300 plan runs.
Do Oqoqo runs include model inference costs?
No. Oqoqo says runs cover platform and cloud infrastructure fees, while model inference is separate. Users bring their own model keys or subscriptions.




