ClawWork: The Benchmark That Makes AI Agents Pay Their Own Way
HKUDS turns agent evaluation into a simulated economy, with task bounties, inference costs, occupation-specific judging, and a survival rule that punishes waste.

Introducing ClawWork ๐: Transform your openclaw/nanobot from AI assistant into a money-earning AI coworker. Watch it earn ๐ฐ$10K+ in just 7 hours by completing real professional tasks across 44+ industries โ from Technology & Engineering to Business & Finance, Healthcare & Social Services, and Legal & Operations. Finally, an AI that doesn't just assist โ it works as your true coworker and makes money.
- ClawWork reframes agent evaluation around solvency, not just correctness.
- Its central move is to price thinking, work, and learning inside one loop.
- The repo combines browser automation, occupation-specific judging, and a dashboard into a single economic system.
- The result is less a chatbot benchmark than a labor market simulator for models.
A benchmark with a balance sheet
ClawWork is easiest to understand as a benchmark with cash flow. An agent starts with finite money, earns bounties for completed work, and can fail if it burns too much inference on the way to an answer.
That changes the question. Instead of asking whether a model can produce the right output, ClawWork asks whether it can survive long enough to produce it without going broke.
Why HKUDS built a labor market instead of a quiz
The project comes out of HKUDS and sits in the same broader lineage as OpenClaw, but the point here is sharper than browser automation. ClawWork uses real occupational tasks from GDPVal, so the unit of evaluation is not trivia correctness, it is useful work in a job-shaped context.
That matters because most agent benchmarks flatten the world into isolated prompts. ClawWork keeps the mess: different professions, different wages, different rubrics, and different costs to get from intent to deliverable.
The real invention is the work or learn fork
The cleverest part of ClawWork is not the dashboard or the browser wrapper. It is the forced choice between work and learn: earn now, or spend budget improving the odds of earning later.
That is a career problem, not a chatbot problem. It gives the agent a small version of the trade-off humans face all the time, where every hour spent leveling up is an hour not spent billing.
- Classify the task into one of the supported occupations.
- Estimate the task value from hours and hourly wage.
- Track every provider call so token usage becomes a real cost line.
- Let the agent choose whether to work or learn.
- Grade the resulting artifact with occupation-specific criteria and update the balance sheet.
How ClawWork prices, tracks, and judges output
The codebase is split into a few clear layers. livebench/ holds the simulation logic, clawmode_integration/ bridges the economic rules into Nanobot-style agent loops, frontend/ renders the live dashboard, and eval/meta_prompts/ stores profession-specific rubrics for the judge model.
That structure matters because it separates the job economy from the browser mechanics. One layer figures out what the task is worth, another layer counts the cost of model calls, and another layer decides whether the artifact deserves to stay in business.
The result is a clean accounting loop. Task classification turns a user request into a profession, provider wrappers capture actual inference charges, artifact tools create deliverables, and the judge compares the output against standards for that occupation.
What the system is actually measuring
The judge is not asking a generic question like, "Was this helpful?" It is asking whether the artifact meets the standard for the job it claims to do, which is a much harder and more realistic test.
| System | What it measures | What it misses |
|---|---|---|
| ClawWork | Quality per dollar, plus survival under budget pressure | It still evaluates in a simulated setting, not a live labor market |
| Playwright automation | Whether scripted browser steps succeed | It is brittle when UI structure changes and it does not price computation |
| OpenClaw or similar agent runners | Whether an agent can act in the browser | It usually does not make solvency the core metric |
| Static benchmarks | Answer accuracy on fixed prompts | They ignore occupational context, long-horizon cost, and failure from waste |
What ClawWork beats, and what it does not try to be
This is not trying to be a universal browser automation framework that wins by DOM cleverness alone. It is trying to be a measuring system for economically constrained agents, which means the bar is higher and the lens is narrower.
That distinction is the whole point. If a model can complete a task but only by spending too much to get there, ClawWork treats that as a real failure, not a cosmetic one.
The new metric is solvency
ClawWork points toward a different future for agent evaluation. The most important number may not be accuracy, latency, or even raw task success, but whether the system can keep earning while it thinks, learns, and works.
That is a sharper standard for agents meant to do real work. If the model cannot stay solvent, it is not ready to behave like a coworker, no matter how fluent it sounds.