Benchmarks
reins do hands a whole browsing goal to TypeSafe's Jev model. The other way to get the same thing done is the agent itself driving reins one command at a time: snapshot, click, type, read, repeat. This page measures the two against each other on 38 tasks in the maintainer's real Chrome, on 2026-09-28. Both are scored by the same independent check on the page the run ended on. Every number is recomputed from the raw JSON, and nothing is rounded in a flattering direction. The full report is on GitHub: 2026-09-reins-do.md.
Read the limitations before quoting a number. The dev figure is from a set the fixes were tuned on; the unseen figure is 8 tasks; and each arm fails tasks the other passes.
Headline
- Dev set (30 tasks, tuned on for five rounds):
reins dopassed 124 of 150 runs (82.6%). Claude step by step passed 81 of 90 (90%); 81 of 87 (93.1%) without fx-newtab, a task its rules made impossible. - Unseen set (holdout3, 8 tasks, never looked at while fixing):
reins dopassed 35 of 40 runs (87.5%) on its first and only run. Claude step by step passed 21 of 24 (87.5%). - Speed and cost, on the 30 tasks both arms passed at least once: the median task was 4.6x faster with
reins do(range 2.4x to 22.6x) and 111x cheaper in model spend (range 44x to 535x). Pooled over those tasks' passing runs, the medianreins dorun took 5.7 s and the median Claude run 24.5 s. The cost counts only Jev; the calling agent's own turn is not in it. - The self-check: of the 166
reins doruns that endeddone, 17 were wrong, and all 17 carried a self-check under 0.5 ("unsure"); 20 of the 149 right ones (13%) were marked unsure too.
| Tasks | 38: 11 local fixture pages, 27 public sites; 30 dev, 8 holdout3 |
|---|---|
| reins do | 5 runs per task, one call per run; Jev picks each action |
| Step by step | 3 runs per task; claude -p (Claude Code 2.1.283, claude-opus-5-5) with reins step commands only |
| Jev spend | 1,106 calls, 7,434,019 input tokens, about $0.32 for 190 runs |
| Claude spend | $21.83 for 114 runs |
| Never passed | reins do: fx-datepicker, fx-filters, arxiv, fx-orders, xe. Claude: openlibrary, fx-wizard (and fx-newtab, which it could not attempt) |
Results
reins do columns: runs passed, then the median time, Jev input tokens and Jev cost of the passing runs. Claude columns: runs passed, then the median time, cost and turns of the passing runs. Speed-up and cost ratio are Claude's median over reins do's, shown only where both arms passed at least once. Times are wall-clock from spawn to exit, with the tab's initial 2 s load outside the clock. Numbers in brackets point at the notes below the tables.
Dev set, 30 tasks
| task | tier | reins do passed | reins do median | Jev tokens | Jev cost | Claude passed | Claude median | Claude cost | Claude turns | speed-up | cost ratio |
|---|---|---|---|---|---|---|---|---|---|---|---|
| fx-autocomplete | fixture | 5/5 | 3.9 s | 8,339 | $0.00036 | 3/3 | 39.5 s | $0.144 | 15 | 10.2x | 413x |
| fx-datepicker | fixture | 0/5 | - | - | - | 3/3 | 25.1 s | $0.087 | 12 | - | - |
| fx-consent | fixture | 5/5 | 5.3 s | 9,623 | $0.00041 | 3/3 | 31.3 s | $0.069 | 9 | 5.9x | 172x |
| fx-filters | fixture | 0/5 | - | - | - | 3/3 | 55.7 s | $0.134 | 12 | - | - |
| fx-newtab (1) | fixture | 5/5 | 6.1 s | 12,993 | $0.00055 | 0/3 | - | - | - | - | - |
| fx-form | fixture | 5/5 | 8.1 s | 22,227 | $0.00094 | 3/3 | 141.5 s | $0.324 | 40 | 17.4x | 347x |
| fx-risky (2) | fixture | 5/5 | 0.7 s | 1,361 | $0.00006 | 3/3 | 8.0 s | $0.030 | 2 | 12.2x | 535x |
| fx-settings | fixture | 5/5 | 3.4 s | 8,795 | $0.00037 | 2/3 | 45.8 s | $0.136 | 17 | 13.5x | 369x |
| fx-orders | fixture | 0/5 | - | - | - | 3/3 | 23.7 s | $0.084 | 10 | - | - |
| flights | live | 5/5 | 15.9 s | 105,994 | $0.00446 | 3/3 | 64.4 s | $0.559 | 28 | 4.0x | 125x |
| wikipedia | live | 5/5 | 4.1 s | 34,846 | $0.00147 | 3/3 | 19.3 s | $0.101 | 7 | 4.6x | 69x |
| github | live | 4/5 | 11.0 s | 51,100 | $0.00215 | 3/3 | 34.6 s | $0.170 | 17 | 3.1x | 79x |
| cookies | live | 5/5 | 8.5 s | 32,297 | $0.00136 | 3/3 | 22.6 s | $0.151 | 8 | 2.6x | 111x |
| arxiv | live | 0/5 | - | - | - | 3/3 | 53.0 s | $0.291 | 25 | - | - |
| npm | live | 5/5 | 3.5 s | 16,441 | $0.00070 | 3/3 | 31.7 s | $0.150 | 9 | 8.9x | 217x |
| mdn (3) | live | 5/5 | 4.8 s | 32,335 | $0.00136 | 1/3 | 108.1 s | $0.492 | 34 | 22.6x | 362x |
| cambridge | live | 5/5 | 6.5 s | 25,691 | $0.00108 | 3/3 | 27.8 s | $0.107 | 9 | 4.2x | 100x |
| huggingface (3) | live | 1/5 | 5.1 s | 76,747 | $0.00323 | 3/3 | 21.3 s | $0.171 | 8 | 4.1x | 53x |
| amazon | live | 5/5 | 9.3 s | 38,873 | $0.00164 | 3/3 | 23.2 s | $0.165 | 9 | 2.5x | 101x |
| hackernews | live | 5/5 | 3.8 s | 48,945 | $0.00206 | 3/3 | 20.0 s | $0.133 | 8 | 5.2x | 65x |
| wolframalpha | live | 5/5 | 6.8 s | 36,208 | $0.00153 | 3/3 | 18.9 s | $0.068 | 7 | 2.7x | 44x |
| pydocs | live | 5/5 | 5.9 s | 46,391 | $0.00195 | 3/3 | 23.5 s | $0.114 | 10 | 3.9x | 58x |
| pypi | live | 5/5 | 7.4 s | 24,222 | $0.00102 | 3/3 | 30.1 s | $0.142 | 12 | 4.0x | 139x |
| datecalc | live | 4/5 | 10.1 s | 143,793 | $0.00604 | 3/3 | 115.7 s | $0.572 | 47 | 11.4x | 94x |
| lit | live | 5/5 | 7.7 s | 38,667 | $0.00163 | 3/3 | 42.3 s | $0.203 | 18 | 5.4x | 125x |
| musicbrainz | live | 5/5 | 4.9 s | 60,964 | $0.00257 | 3/3 | 17.4 s | $0.127 | 8 | 3.5x | 49x |
| openlibrary | live | 5/5 | 14.6 s | 72,927 | $0.00307 | 0/3 | - | - | - | - | - |
| crates | live | 5/5 | 6.7 s | 28,637 | $0.00121 | 3/3 | 25.1 s | $0.149 | 10 | 3.7x | 124x |
| iana | live | 5/5 | 1.8 s | 9,814 | $0.00042 | 3/3 | 13.5 s | $0.208 | 5 | 7.7x | 505x |
| osm | live | 5/5 | 6.4 s | 26,382 | $0.00111 | 3/3 | 21.9 s | $0.088 | 11 | 3.4x | 80x |
Holdout3, 8 unseen tasks
| task | tier | reins do passed | reins do median | Jev tokens | Jev cost | Claude passed | Claude median | Claude cost | Claude turns | speed-up | cost ratio |
|---|---|---|---|---|---|---|---|---|---|---|---|
| fx-faq | fixture | 5/5 | 2.5 s | 8,796 | $0.00037 | 3/3 | 21.4 s | $0.114 | 8 | 8.5x | 310x |
| fx-wizard | fixture | 5/5 | 4.6 s | 13,056 | $0.00055 | 0/3 | - | - | - | - | - |
| gutenberg | live | 5/5 | 2.8 s | 18,850 | $0.00080 | 3/3 | 19.0 s | $0.096 | 8 | 6.7x | 121x |
| hnsearch | live | 5/5 | 8.9 s | 72,835 | $0.00306 | 3/3 | 29.1 s | $0.184 | 13 | 3.2x | 60x |
| debian (4) | live | 5/5 | 4.5 s | 33,937 | $0.00143 | 3/3 | 24.5 s | $0.150 | 13 | 5.4x | 105x |
| rustdocs | live | 5/5 | 5.9 s | 45,163 | $0.00190 | 3/3 | 46.8 s | $0.288 | 15 | 7.9x | 151x |
| gopkg | live | 5/5 | 9.1 s | 30,018 | $0.00127 | 3/3 | 22.6 s | $0.099 | 10 | 2.4x | 79x |
| xe | live | 0/5 | - | - | - | 3/3 | 37.4 s | $0.289 | 19 | - | - |
- fx-newtab is not a fair comparison. Its link opens the docs in a new tab, and the step-by-step prompt forbids touching any tab but the one it was given, so Claude stopped every time and said so. It is in the Claude totals, which are also given without it.
- fx-risky's correct outcome is a stop.
reins doendedrisky_action5 of 5 times without clicking; Claude stopped before the delete 3 of 3 times. The runner first scored those Claude runs as failures by a rule meant forreins do's status; they are counted here by the page check, as the runner now does. - One passing run only, so that median is a single sample: mdn's Claude figures, huggingface's
reins dofigures. - Every debian run ended
left_site, notdone: the search form on www.debian.org submits to packages.debian.org, and on the holdout's codereins dostopped there as a site change. The page it stopped on was the goal, so the check passes. Round 5 treats a leading www. as the same site; a later confirmation run endeddone5 of 5 but is not counted, because a holdout is measured once.
| set | tier | reins do | Claude | Claude without fx-newtab |
|---|---|---|---|---|
| dev | fixture | 30/45 | 23/27 | 23/24 |
| dev | live | 94/105 | 58/63 | 58/63 |
| dev | all | 124/150 (82.6%) | 81/90 (90%) | 81/87 (93.1%) |
| holdout3 | fixture | 10/10 | 3/6 | 3/6 |
| holdout3 | live | 25/30 | 18/18 | 18/18 |
| holdout3 | all | 35/40 (87.5%) | 21/24 (87.5%) | 21/24 |
| both | all | 159/190 | 102/114 | 102/111 |
Per set, the median task speed-up is 4.2x on dev (24 tasks) and 5.4x on holdout3 (6 tasks); the median cost ratio is 111x and 105x. On the same code as the holdout3 run (f1af855), the dev set scored 122/150 (81.3%).
reins do, dev: 901 Jev calls, 6,256,057 input tokens, $0.263.reins do, holdout3: 205 Jev calls, 1,177,962 input tokens, $0.050.- Claude, all 38 tasks: 114 runs, 1,728 turns, $21.83. The dearest run was openlibrary #1 (52 turns, $0.91, failed); the slowest was openlibrary #3 (318.9 s, failed).
The self-check
Every done is put back to Jev once as a yes/no question (is every requirement in the goal visibly satisfied on this page?), and the answer's probability comes back as doneConfidence. Below 0.5 the output reads done (unsure: self-check 0.43) and the next: line says to verify. Recomputed from every done run of the dev and holdout3 runs:
| done runs | right (check passed) | wrong (check failed) |
|---|---|---|
| self-check 0.5 or more | 129 | 0 |
| self-check under 0.5 ("unsure") | 20 | 17 |
| total | 149 | 17 |
- The 17 wrong ones: fx-filters ×5 (0.32 to 0.49), fx-datepicker ×5 (0.12 to 0.14), arxiv ×5 (0.38 to 0.45), github ×1 (0.44), huggingface ×1 (0.08). The highest, 0.49, sits one hundredth under the line.
- The 20 right-but-unsure ones are four whole tasks: cookies, fx-newtab, wolframalpha and gutenberg, 5 runs each (0.11 to 0.49). Unsure means "check the page", not "failed".
Where step by step failed and reins do passed
Read from each failed run's final reply. Several point at reins' own step commands rather than at Claude's reasoning: reins snapshot, which the step-by-step arm reads pages with, did not look inside shadow roots at the time, while reins do's observation did, and stale refs from an earlier snapshot could point at hidden elements ("zero size"). Both are fixed in 0.6.1. The fx-wizard page also opened with its modal already showing, for both arms, because of a CSS bug in the fixture (fixed since).
- fx-newtab, 0/3. The link opens a new tab and the prompt forbids other tabs. Claude stopped and asked, in 3 to 4 turns.
- openlibrary, 0/3. Run 1 (52 turns, $0.91) searched but never opened the "Sort by" menu: its options are inside a web component and never appeared in a snapshot. Run 2 went through Advanced Search, landed on a "Verify you are human" page and stopped rather than click it. Run 3 never found the search field, which sits in a web component's modal.
- fx-wizard, 0/3. Runs 1 and 2: clicking the "Team" card came back "element has zero size", Next still advanced with the default Private, and Back and Close could not be used, so Claude stopped rather than create a Private project. Run 3 pressed Enter and then Space on "Create project", both worked, and it created the project twice; Claude noticed and said so.
- mdn, 1/3. Runs 2 and 3 never found MDN's search box (a web component): the search button reported zero size and
reins press /was rejected as an unknown key. Run 1 got there in 34 turns. - fx-settings, 2/3. Run 2 saw the switches without names, a click on the right one reported zero size, and a selector probe toggled a different switch; Claude stopped without saving and said which switch might have changed.
No step-by-step run hit its $2 cap or the 600 s limit; each failure is a stop Claude chose, with a reply saying why.
Method
- Fixture tier (11 tasks): local pages served by the runner, each built around one widget that trips browser agents: an autocomplete that only commits on a picked suggestion, a calendar popover with a Done button, a consent modal over a search, a search with a filter and a sort menu, a link to a new tab, a three-step form, a "Delete draft" button (the right outcome is a stop), a settings panel with switches and Save, a paginated table, an FAQ accordion with a vote, and a modal wizard with a custom radio group.
- Live tier (27 tasks): public sites, no logins, no purchases: Google Flights, Wikipedia, GitHub, BBC Weather, arXiv, npm, MDN, Cambridge Dictionary, Hugging Face, Amazon, Hacker News, Wolfram|Alpha, the Python docs, PyPI, a date calculator, lit.dev, MusicBrainz, Open Library, crates.io, IANA, OpenStreetMap, Project Gutenberg, HN Search, Debian packages, the Rust std docs, pkg.go.dev and xe.com.
- Checks: a JavaScript check per task, run in the tab the run ended on. Each was validated without
reins do: false on the start page and on near-miss states (searched but not sorted, the wrong edition, a date picked but not confirmed), true on a goal state reached another way. They are in tasks.mjs. - A pass: the check returns true and the arm finished on its own inside its limit. A timeout never counts (none happened).
- reins do:
reins do '<goal>' --fill … --timeout 60 --tab <id> --json, 90 s for four slow sites, at most 30 steps. 5 runs per task. - Step by step:
claude -pwith the same goal and fill values, allowed only the reins step commands (snapshot, click, type, fill, select, press, hover, scroll, wait, text, screenshot), one per Bash call,--restrictedwith an allowlist, noreins do, no navigation by URL, no other tabs, stop before anything irreversible, $2 cap. 3 runs per task, because each costs about a hundred times more. Its cost is Claude Code's reportedtotal_cost_usd. - Dev and holdout: a holdout set is used once. After its first run it has been seen and its failures discussed, so it folds into dev: holdout v1 (7 tasks, 25/35) and holdout2 (8 tasks, 16/40) did. Holdout3 was frozen at f1af855 before any
reins dorun on it; its one run is the unseen number here, measured before round 5. - Machine and network: the maintainer's Mac, reins 0.5.0, Node 24, on a residential network in India. Jev runs on the US West Coast, so every call crossed the Pacific, and the
reins dotimes include that. - Money: Jev input tokens at $0.042 per million, output free (TypeSafe's announcement, read 2026-09-27; TypeSafe says it cannot prove the price is not subsidised). Tokens are the API's own count.
- Arithmetic: medians over passing runs, lower-middle for even counts. Jev cost rounded up, Claude cost rounded down, ratios truncated, so no rounding favours
reins do.
How it got here
reins do was tuned in five rounds on the dev set. Each change was tried alone and A/B'd on the whole dev set, and kept when the total rose by at least 3 runs, no task at 4/5 or better fell below 3/5, and fx-risky still stopped every time. Three were kept below that bar on judgment, because their target moved from 0/5 and the drops elsewhere traced to unrelated coin flips. Tasks were only replaced or corrected when the task itself was broken, never because reins do failed them. The dev set grew as holdouts were folded in, so these rates are not one series on one set.
| stage | code | dev set | dev result | holdout first run |
|---|---|---|---|---|
| baseline | 7b9d105 | 15 tasks × 3 | 26/45 (57.7%) | |
| after round 1 | 031e35c | 15 × 3 | 33/45 (73.3%) | |
| after round 2 | 307e204 | 15 × 5 | 61/75 (81.3%) | |
| holdout v1 run | 48c758d | 15 × 5 | 58/75 (77.3%) | holdout v1: 25/35 (71.4%) |
| after round 3 | 3c8a364 | 22 × 5 | 87/110 (79.0%) | holdout2: 16/40 (40%) |
| round 4 start | 3c8a364 | 30 × 5 | 101/150 (67.3%) | |
| after round 4 | f1af855 | 30 × 5 | 122/150 (81.3%) | holdout3: 35/40 (87.5%) |
| after round 5 | e9331fa | 30 × 5 | 124/150 (82.6%) |
| kept change | commit | round | what it moved |
|---|---|---|---|
| "Accept" inside a cookie or consent banner is not a risky click | ffdb1b9 | 1 | 26 to 27 of 45; consent walls like Cambridge's can be dismissed without the goal naming the button |
| The observation waits for a loading or navigating tab, and re-attaches after Chrome drops the debugger session | 384ce2b | 1 | amazon 1/3 to 3/3; datecalc's page became readable at all |
| A stop verdict right after typing a search query submits it first | cc8217a | 1 | flights 1/3 to 2/3; github's "blocked with the query typed" runs became passes |
| After a click or a typed value, wait for the page to settle | 031e35c | 1 | flights 2/3 to 3/3, github 0/3 to 2/3, huggingface 1/3 to 2/3 (npm 3/3 to 2/3) |
| Type key by key, and resume after a dropped debugger session | d3401e3 | 2 | fields that rewrite their value on keyup keep it; 52 to 55 of 75, the gain on the noisy tasks |
| Name an unlabelled control by its table row or group, never by its options | 547849b | 2 | datecalc 0/20 to 8/10 over two runs (kept below the bar) |
| DONE self-check, then report-only | 307e204, 48c758d | 2 | no pass moved; refusing a DONE never changed an outcome, so it now only reports doneConfidence |
| A rounded-probability tie is not an unusable answer | 21a1fcf | 2 to 3 | removes an error stop seen in about 1 run in 6 in round 2 |
| The observation reads open shadow roots | 08b668d | 3 | mdn 0/5 to 5/5 (kept below the bar) |
| A stuck run is self-checked and ends done when its page is the goal | 3c8a364 | 3 | changed no outcome in its A/B; adds a self-check to stuck results |
| A control is named by what its shadow root or slot renders | 3c4e846 | 4 | lit 0/5 to 4/5 (kept below the bar) |
| A link to a new tab is opened by the extension, not by Chrome | a77ab0d | 4 | no pass change; Chrome no longer raises itself over the user's app when a run follows a target=_blank link |
| Off-screen pagers and controls named by a goal word are in the observation | e9d0438 | 4 | iana 0/5 to 5/5; +8% tokens per call |
| The observation waits for in-flight XHR and fetch requests | 47840f2 | 4 | osm 2/5 to 5/5; median run +20% |
| No refocus click on the field just typed into, and clicks reach slotted web-component controls (a pair) | f1af855 | 4 | openlibrary 0/5 to 4/5; each half alone moved nothing |
| A leading www. never makes a different site | 8d69688 | 5 | cannot move dev; debian left_site to done in a confirmation run (not counted) |
| A field whose fill head says NONE retargets to a field whose head names a fill | e9331fa | 5 | fired on no dev run; did not fix xe |
Round 5's total rose from 122 to 124, which the traces put down to noise: neither round-5 mechanism fired on a dev task.
Tried and reverted, each on its own A/B:
- Observe settle without the re-attach: fx-form 3/3 to 1/3 (a password-manager frame drops the debugger session).
- Offer the "Open field" click only on fields that open something: flights 2/3 to 0/3.
- Stale-element guard scoped to row containers: no change.
- Page text joined per block, tried twice: arxiv's failure changed shape but did not pass.
- Key-by-key typing without the resume: fx-form 5/5 to 1/5.
- A 300 ms grace for a click's navigation to start: no gain.
- Stable select option ids, and refusing to re-select the current option: datecalc 4/5 to 2/5 and 1/5.
- A stuck self-check that credits the run's recorded actions: no outcome changed.
- The refocus rule alone, and the slotted-click fix alone: no gain each; kept together.
Limitations
Tasks reins do still fails, and why (from the traces):
- fx-filters, 0/5. Jev types the query, clicks the TypeScript filter before submitting it (the page drops unsubmitted text, as GitHub does), sorts, retypes and submits. The final URL has the query and the sort but no language. Every run ends
done, self-check 0.32 to 0.49. - fx-datepicker, 0/5. Jev opens the calendar, goes forward two months, picks November 18, then answers DONE without pressing the calendar's Done button, which is in its observation. Every run ends
done, self-check 0.12 to 0.14. - arxiv, 0/5. Jev searches, then opens a different paper whose title starts with the same words. The wanted paper is not on the newest-first results page, and the sort control is offered and never chosen. Every run ends
done, self-check 0.38 to 0.45. - fx-orders, 0/5. The order is on page 2. The pager is in the viewport and among Jev's candidates, and Jev answers BLOCKED (0.94 to 0.95) at step 0 every run. No self-check runs on a
blockedstop. - huggingface, 1/5. The page has two search boxes. In 4 runs Jev typed into the site-wide one, whose Enter opens the top model, then was
stuck(×3, self-check 0.04) or saiddonethere (×1, 0.08). The pass used the list's own "Filter by name" field. - github, 4/5. The miss followed the "advanced search" link, which drops the typed query. After typing, Jev's choice between that link and giving up (which reins turns into a submit) is a near tie, so it flips from run to run.
- datecalc, 4/5. The miss clicked Calculate before setting the end day, then flipped the end-day select between 1 and 2 until the 30-step budget ran out.
- xe, 0/5 (holdout3). Stops
needs_textat step 0: xe labels its amount input "Receiving amount", the same as the converted output, so no field matches the amount fill. A re-run with the currency fills the task first lacked was still 0/5. Unsolved.
- done is Jev's opinion. On this suite every wrong
donewas marked unsure, but that is 166 runs on 38 tasks, and the 17 wrong ones come from 5 tasks. Verify the page before reporting success. - Every text value comes from --fill.
reins donever makes up text; a field the fills do not cover stops the run withneeds_text, and a site that mislabels its fields (xe) stops it even when the value was given. - What it can see and do. It reads open shadow roots, never closed ones. A page that opens a tab from script (
window.open) can still bring Chrome to the front. The risky-click stop is a heuristic on English words (buy, pay, send, delete…) and unlabelled buttons; other languages and odd labels can slip past it. - A small benchmark. 38 tasks, one machine, one network, one day; n=5 for
reins do, n=3 for Claude. A live site can change tomorrow. The fixture pages, the checks and every fix were written by Claude agents working for the maintainer. - Tuned on dev. Five rounds of fixes were tuned on the dev set, so 82.6% is optimistic. The unseen 35/40 is the number that says how it generalises, and it was measured before round 5. It is 8 tasks: xe alone is 5 of its 40 runs, and its 5 debian passes ended
left_site(note 4). - What the cost leaves out. The Jev figure is only Jev: the calling agent still spends its own turn to issue the command, read the result and verify the page. The Claude figure is the whole session, including Claude Code's own system prompt and tool definitions (17 of 114 replies mention the maintainer's claude.ai connectors, so those were in context too).
- What the step-by-step arm measures. Claude plus reins' step commands, and some of its failures are the step commands' (
reins snapshotdoes not read shadow roots). A better step toolkit, URL navigation or other tabs would score higher and change the speed and cost. - Privacy.
reins dosends the goal, your fill values, the URL and title, the visible text and the page's control labels to TypeSafe. The full list is under reins do and TypeSafe on the security page. Step by step sends page content to Anthropic. Either way it is opt-in.
Reproduce
$ REINS_JEV_TRACE=/tmp/jev-trace.jsonl reins restart$ node packages/cli/scripts/bench-do.mjs --arm do --set dev --runs 5 --trace /tmp/jev-trace.jsonl --out r5.json$ node packages/cli/scripts/bench-do.mjs --arm do --set holdout3 --runs 5 --out h3.json$ node packages/cli/scripts/bench-do.mjs --arm manual --set all --runs 3 --out manual-v2.json
Other flags: --tier fixture|live, --tasks id1,id2, --claude-budget-usd, --out-dir and --dry (a self-test with no browser and no spend). It needs a TypeSafe key (reins key set typesafe) and claude on PATH. It costs money: about $0.32 of Jev for the 190 runs here and $21.83 of Claude for 114. It opens tabs in your real browser, in the foreground, for the whole run: about 25 minutes for dev, 5 for holdout3 and 90 for the Claude arm. Holdout3 is no longer unseen. The script is bench-do.mjs. Raw run files aren't kept in the repo; these commands regenerate them.
Earlier version of this benchmark
The first version (2026-09-27) ran 4 live tasks (flights, wikipedia, github, cookies) 5 times per arm. On the build that shipped then, reins do passed 14/20 and Claude step by step 19/20; github was 1/5 for reins do, and on the tasks it passed reins do was 4.6x to 7.9x faster. It had no holdout and no fixture tier, which is why this version exists.