ClawWork: The Benchmark That Makes AI Agents Pay Their Own Way

HKUDS turns agent evaluation into a simulated economy, with task bounties, inference costs, occupation-specific judging, and a survival rule that punishes waste.

10 min read • View on GitHub • More from HKUDS

An AI worker sits at a desk with a browser window open, a calculator nearby, and a ledger spread across the table. The scene explains ClawWork's core idea: an agent is not just asked to finish work, it is asked to afford the computation it consumes while doing it.
ClawWork treats every task like a job with accounting attached. The agent's output matters, but so does the bill it leaves behind.

Introducing ClawWork ๐Ÿš€: Transform your openclaw/nanobot from AI assistant into a money-earning AI coworker. Watch it earn ๐Ÿ’ฐ$10K+ in just 7 hours by completing real professional tasks across 44+ industries โ€” from Technology & Engineering to Business & Finance, Healthcare & Social Services, and Legal & Operations. Finally, an AI that doesn't just assist โ€” it works as your true coworker and makes money.

Chao Huang, The University of Hong Kong ยท Chao Huang on LinkedIn
Key Takeaways

A benchmark with a balance sheet

ClawWork is easiest to understand as a benchmark with cash flow. An agent starts with finite money, earns bounties for completed work, and can fail if it burns too much inference on the way to an answer.

That changes the question. Instead of asking whether a model can produce the right output, ClawWork asks whether it can survive long enough to produce it without going broke.

WSJ-style hedcut portrait of Chao Huang based on his GitHub avatar. The portrait supports the origin story by showing the public face behind the project's shift from assistant-style automation to economically accountable agent work.

Why HKUDS built a labor market instead of a quiz

The project comes out of HKUDS and sits in the same broader lineage as OpenClaw, but the point here is sharper than browser automation. ClawWork uses real occupational tasks from GDPVal, so the unit of evaluation is not trivia correctness, it is useful work in a job-shaped context.

That matters because most agent benchmarks flatten the world into isolated prompts. ClawWork keeps the mess: different professions, different wages, different rubrics, and different costs to get from intent to deliverable.

A close-up of a balance scale and ledger where one side drains while the other fills. The scene shows that in ClawWork, the cost of thinking is not hidden in the background, it is part of the work itself.
The project makes inference expense visible. A model can be smart and still lose if it spends too much to get there.

The real invention is the work or learn fork

The cleverest part of ClawWork is not the dashboard or the browser wrapper. It is the forced choice between work and learn: earn now, or spend budget improving the odds of earning later.

That is a career problem, not a chatbot problem. It gives the agent a small version of the trade-off humans face all the time, where every hour spent leveling up is an hour not spent billing.

One loop ties together pricing, action, and judgment. ClawWork is interesting because it makes the economy visible at the same moment it makes the agent act.

  1. Classify the task into one of the supported occupations.
  2. Estimate the task value from hours and hourly wage.
  3. Track every provider call so token usage becomes a real cost line.
  4. Let the agent choose whether to work or learn.
  5. Grade the resulting artifact with occupation-specific criteria and update the balance sheet.

How ClawWork prices, tracks, and judges output

The codebase is split into a few clear layers. livebench/ holds the simulation logic, clawmode_integration/ bridges the economic rules into Nanobot-style agent loops, frontend/ renders the live dashboard, and eval/meta_prompts/ stores profession-specific rubrics for the judge model.

That structure matters because it separates the job economy from the browser mechanics. One layer figures out what the task is worth, another layer counts the cost of model calls, and another layer decides whether the artifact deserves to stay in business.

The result is a clean accounting loop. Task classification turns a user request into a profession, provider wrappers capture actual inference charges, artifact tools create deliverables, and the judge compares the output against standards for that occupation.

What the system is actually measuring

The judge is not asking a generic question like, "Was this helpful?" It is asking whether the artifact meets the standard for the job it claims to do, which is a much harder and more realistic test.

SystemWhat it measuresWhat it misses
ClawWorkQuality per dollar, plus survival under budget pressureIt still evaluates in a simulated setting, not a live labor market
Playwright automationWhether scripted browser steps succeedIt is brittle when UI structure changes and it does not price computation
OpenClaw or similar agent runnersWhether an agent can act in the browserIt usually does not make solvency the core metric
Static benchmarksAnswer accuracy on fixed promptsThey ignore occupational context, long-horizon cost, and failure from waste
A split scene shows two ways of evaluating agents. One side is a treadmill with score counters, while the other is a real desk with invoices, a browser, and a survival gauge. The image explains that ClawWork is not a better quiz, it is a different machine.
Static benchmarks reward output. ClawWork rewards output that can survive its own cost structure.

What ClawWork beats, and what it does not try to be

This is not trying to be a universal browser automation framework that wins by DOM cleverness alone. It is trying to be a measuring system for economically constrained agents, which means the bar is higher and the lens is narrower.

That distinction is the whole point. If a model can complete a task but only by spending too much to get there, ClawWork treats that as a real failure, not a cosmetic one.

The new metric is solvency

ClawWork points toward a different future for agent evaluation. The most important number may not be accuracy, latency, or even raw task success, but whether the system can keep earning while it thinks, learns, and works.

That is a sharper standard for agents meant to do real work. If the model cannot stay solvent, it is not ready to behave like a coworker, no matter how fluent it sounds.