2026-10-06, updated 2026-10-07 with the changes we then made (E4). Every number here comes from run logs; the list with sources is in facts.json, the full record in docs/verification/mcp-efficiency.

We wanted to say: "Codex or Claude Code can make a similar model on their own, but building it in Laydyne over MCP takes far fewer tokens, and it is easier to manage." Before saying it, we measured it. For the tasks below, it is not true: Codex alone used fewer tokens. This article shows what we compared, the numbers, why it came out this way, what we tried to make it lighter (and why we did not ship most of it), what we then changed to cut the requests and the result sizes (input tokens fell by 26–53 %, but Codex alone still uses fewer), and where Laydyne is still the better tool.

What we compared

Four tasks, each with the same deliverables for both setups:

TaskWhat was askedDelivered
T1A 12 × 9 m shop: entrance in the south wall, a counter and a display table near it, two shelf rows at least 6 m long with a 1.2 m aisle, every passage at least 0.9 mA dimensioned plan.svg, a 3D view, a short report
T2The same shop, changed: a second counter touching the first, the display table moved 2 m north (or as far as the 0.9 m passages allow)The list of changes, a new plan.svg with the changes marked
T3A 30 × 20 m warehouse: 4 dock doors, a 4 m staging area, a 6 × 4 m office, back-to-back pallet racking (2.7 × 1.1 × 6 m bays) with 3 m aisles, as much as fitsA dimensioned plan.svg, a 3D view, bay and pallet counts
T4Hand-over of the warehouseA scaled drawings.pdf with a title block, a quantities.csv
  • A — Codex alone worked in an empty folder: it kept the layout in a model.json of a given format and made plan.svg and a three.js page itself, checking overlaps and passage widths with its own code.
  • B — Codex + Laydyne MCP built the same thing in an open Laydyne Studio tab, checked it with check_space, and saved the drawing that get_view_image makes.

Both ran Codex CLI 0.160.1 with gpt-6-astra (reasoning low), twice per task. The user's other MCP servers and plugins were switched off for both. A script, not our eyes, judged every result: room size, overlaps, narrow gaps, aisle widths, counts, the change applied exactly, dimensions on the drawing (in the PDF by its text or by OCR), the CSV against the model.

The numbers

Every run of both setups passed the same checks (16 of 16 each). The difference is in tokens:

TaskCodex alone, input tokensCodex + Laydyne, input tokensEstimated API cost, alone → with Laydyne
T1 shop94,698481,018$0.45 → $1.04
T2 change89,478453,216$0.37 → $1.03
T3 warehouse107,223533,375$0.52 → $1.24
T4 hand-over246,915655,818$0.70 → $1.40
Input tokens per task
Input tokens per task

Means of two runs. With Laydyne, Codex sent 2.7–5.1 times as many input tokens; most were cached, and it wrote fewer output tokens than Codex alone (0.48–0.76 times). Priced at the gpt-6-astra list price ($10 input, $1 cached input, $50 output per million tokens; the runs themselves used a ChatGPT sign-in, not API billing), that is 2.0–2.8 times the cost. Time was about the same (T1: 125 s alone, 114 s with Laydyne).

Estimated API cost per task
Estimated API cost per task

Why

More model requests. Codex alone wrote one generator script, ran it, fixed it and answered: about 4 requests. With Laydyne it looked up the tools, read the space, built, checked, drew and saved: about 11. Codex re-sends its own context (about 21,000 tokens: its instructions, its list of skills, its permissions) and everything gathered so far with every request, so each extra request costs tens of thousands of input tokens.

Reading the tool definitions. Codex 0.160 does not send MCP tool definitions with every request; the model looks them up in a list (ALL_TOOLS) when it needs them. It puts the server's instructions in front of each tool's description. Laydyne's instructions were 9,006 characters long, repeated in front of all 67 tools: 161,345 tokens of tool text, 81 % of it the same instructions again. Codex cuts such a listing to about 10,000 tokens, but that cut listing then rides along with every later request. We estimate the tool lookups at about a third of the input tokens in the Laydyne runs.

Large results printed whole. The agent often printed a whole tool result. Laydyne's results carry the same JSON twice (as text and as structured content), and the first read of the space carries about 7 KB of fixed explanation.

What we tried, and what we kept

For clients that send every tool definition with every request (the Claude app, ChatGPT, Cursor) the tool list alone is about 89,000 tokens per request. We tried a lighter list, keeping every tool callable:

  • The tools of nearly every session listed first and in full; the others in brief (what the tool is for, its top-level inputs), with a new describe_tool that returns any tool's full description and input from the server.
  • A brief schema only relaxed the full one, so a client that checks arguments before calling never refused a valid call; Studio still checked every call against the full schema.
  • Server instructions cut from 9,006 to 1,192 characters.
What the Laydyne MCP costs before any work
What the Laydyne MCP costs before any work

On paper that is 39,878 instead of 89,162 tokens per request for a client that loads every tool (−55 %, counted with the o200k_base tokenizer). In Codex it did not pay off:

  • On the four tasks above, re-running with the lighter list changed little (T1 −3.6 %, T2 −10.5 %, T3 +56 %, T4 +13 %; two runs each, within the spread between runs).
  • On flows that need the tools listed in brief — the process run and the option comparison of our clinic use case, and a 2-storey office with a stair and a lift — every run reached its outcome with every list, and the agent never called a tool with wrong arguments because it had seen only the brief schema. But Codex spent more: summed over the three flows, 2.63 million input tokens with the full list and long instructions, 3.77 million with the lighter list (+43 %), 3.78 million with an adjusted version that listed those flows' tools in full (+44 %), and 3.25 million with the full list and only the short instructions (+24 %).
Flows that use the tools listed in brief
Flows that use the tools listed in brief

Why: Codex reads tool definitions as short TypeScript declarations. A tool listed in brief sends the agent to describe_tool, which answers with the full JSON schema — heavier than the declaration it replaces. And the long instructions, which Codex repeats in front of every tool, also work as a manual: with the short ones the agent explored and re-read more.

So we did not ship the lighter list or the shorter instructions; the code stays on a branch for a test with a client that loads every tool. We shipped the fix for the most frequent failed call: agents sent part colours as "#dbeee7" and were refused (4 of 6 failed calls in the first runs). Colour strings and empty labels are now accepted.

Then we cut the requests and the result sizes

The numbers pointed at two things on our side: how many requests the agent needs, and how much each tool result adds to what Codex re-sends with every later request. On October 7 we changed the following and measured the same four tasks again, twice each, together with Codex alone on the same day.

  • One copy of every result. Laydyne sent each result twice: as JSON text and as structuredContent. Clients read one or the other. - Codex and Claude Code read only structuredContent (Codex then also drops pictures). - Cursor reads only the text; ChatGPT reads both. - A Codex script that printed a result printed both copies.

Results now carry their JSON once, as text, which every client reads.

  • Fixed explanations once. The first read of a space (get_space with detail: "summary") repeated about 6.6 KB of capabilities, coordinate rules and scope. That text now comes from get_space {detail: "about"}, when needed. The summary of an empty floor went from 9,001 to 2,596 characters.
  • Short results on request. create_space takes result: "brief" (the part count and IDs instead of every part), and so does a preview of an edit.
  • A quick start at the top of the instructions. Codex shows nested tool inputs as unknown, so agents re-read tool declarations and guessed shapes. - The instructions now open with the five usual steps and the shapes of their inputs. - They ask the agent to chain one step's calls in one script and to print only what it needs. - A hall of racks is pointed to addArray rather than the authoring kit, which the agent could not download in Codex's sandbox anyway.
  • Calls sent together wait their turn. The link used to answer busy to calls that arrived in parallel. It now lets them wait up to 8 seconds and runs one change at a time. No run in either measurement hit busy, so this one does not show in the numbers.
TaskInput tokens, beforeafterRequests, before → afterCodex alone, same day: input (requests)Estimated API cost, before → after (Codex alone)
T1 shop520,728244,564 (−53 %)12 → 7135,326 (5.5)$1.09 → $0.60 ($0.60)
T2 change613,148391,216 (−36 %)13 → 11117,899 (4.5)$1.35 → $0.75 ($0.41)
T3 warehouse581,063334,425 (−42 %)11.5 → 9121,202 (5)$1.38 → $0.78 ($0.61)
T4 hand-over506,980375,160 (−26 %)11 → 9.5257,262 (7.5)$1.20 → $0.80 ($0.70)

Means of two runs. "Before" was measured again on October 7 on the code before the changes, so that before and after differ only in the code. Day to day, the same setup varied by −23 % to +35 % per task against the October 6 numbers above (Codex alone too: T1 took 135,326 input tokens on October 7 and 94,698 on October 6). All 32 runs of October 7 passed the same scripted checks.

  • The tool output the agent saw fell from 140,456 to 76,432 characters a run. None of it was printed twice any more (17,650 characters a run before).
  • A longer flow, the 2-storey office with a stair and a lift, fell from 986,278 to 576,901 input tokens (−42 %) and from 19.5 to 14 requests. Both runs reached the outcome before and after.

Where Laydyne still costs more:

  • Input tokens. With Laydyne, Codex still sends 1.5–3.3 times the input tokens of Codex alone. The estimated cost is about the same for T1 and 1.1–1.8 times for the others, because Codex alone writes more output tokens, which cost more.
  • The tool list. About 27 % of Laydyne's input is the tool list that Codex prints at the start, cut to about 10,000 tokens and re-sent with every later request.
  • Codex's own context. Another 57 % is Codex's own context (about 21,000 tokens) times the number of requests. Neither part depends on the shape of our results; only fewer requests reduce them.
  • The hand-over PDF. Laydyne draws it as a 150 dpi picture, so agents rebuilt it from the SVG with their own tools. That is most of T4's spread between runs (238,334 to 511,987 tokens). A vector PDF is next on our list.
  • Not changed: a short-lived download link for files. Codex's default sandbox has no network, and an agent connected with OAuth holds no key for its shell. A link that carries its own key would put a secret in a URL.

Bigger and longer work: we measured where Laydyne should win

Measured 2026-10-07. Full record: docs/verification/laydyne-advantage; numbers with sources in its facts-e5.json.

The four tasks above were small. So we measured the kinds of work where we expected Laydyne to earn its tokens, with the same method: same brief and deliverables for both setups, scripted scoring, Codex CLI 0.160.1, now with gpt-6.1-sol at medium reasoning (the current default; the runs above used gpt-6-astra, so compare ratios, not absolute numbers). Three runs per setup, two for the scan and the revision cycle.

TaskCodex alone: input tokens · est. cost · timeCodex + LaydyneLaydyne ÷ alone (tokens · cost · time)
A small change to a big model: in a 20-floor, 2,420-part building, move one partition 2 m, extend a wall, fix two rooms' areas, add 8 seats, change nothing else, deliver the floor plan437,329 · $0.23 · 245 s1,058,005 · $0.32 · 180 s2.42 · 1.40 · 0.73
Simulation with a textbook answer: one vs two counter staff (M/M/1 vs M/M/2), 20 seeded replications, 95 % intervals, checked against queueing theory89,427 · $0.07 · 76 s889,368 · $0.24 · 190 s9.95 · 3.39 · 2.50
From a drawing: walls, doors, 15 named rooms, lifts, stairs and fittings of a 64 × 44 m floor286,692 · $0.16 · 269 s1,106,197 · $0.42 · 247 s3.86 · 2.59 · 0.92
From a scan of that drawing (turned, blurred, JPEG)760,691 · $0.28 · 374 s1,118,345 · $0.39 · 285 s1.47 · 1.43 · 0.76
A revision cycle over three separate sessions: issue revision A; change it and issue B with the changes marked; reopen A exactly and list every difference926,352 · $0.41 · 342 s1,958,886 · $0.61 · 308 s2.11 · 1.51 · 0.90

Medians (the revision cycle: the three sessions together). Cost is estimated at the gpt-6.1-sol list price ($2 input, $0.10 cached input, $10 output per million tokens up to 272K per request); the runs used a ChatGPT sign-in.

Quality was the same. All 13 valid runs of each setup passed their checks: the change exact with nothing else touched on any of the 20 floors, the simulated means within the theory's band, the traced floor matching the published model (scores of 0.93–0.96 out of 1 for both). Laydyne's own model calls (read_underlay_with_ai and the like) were not used in any run.

Tokens: Codex alone still wins every task. The gap shrinks as the work grows — from 2.7–5.1× on the small tasks above (with the other model) to 2.4× on the big model and 1.5× on the scan — because much of Laydyne's cost is fixed: Codex re-sends its own context with every request (about 28–41 % of Laydyne's input), reads the tool list (16–37 %) and reads the space (11–15 %). Codex alone pays more as the job gets harder: on the scan it made 23 requests instead of 10. Even with the tool list costing nothing, the big-model change would still have used more tokens with Laydyne (about 770,000 against 437,000, estimated).

Time and the parts nobody asked for: Laydyne wins. It was faster on the big model (180 s against 245 s) and on both drawings. Its floor plan came out as a drawing sheet — scale 1:250, dimensions, an area schedule, a legend of what was added, moved and changed, a title block — where Codex alone drew its own plan with the changes coloured but no scale or dimensions. The brief did not ask for those, so they did not count.

Simulation with a textbook answer: Codex alone, by far. A short Python script with exponential draws put the theoretical waits of M/M/1 and M/M/2 inside its 95 % intervals in four model requests; Laydyne's engine got the same answer but cost 3.4 times as much. It also showed a mismatch: after a warm-up, Laydyne counts the waits that end in the observed window (a usual steady-state convention), while the brief selected customers by when they arrived, so the agent ran all 40 seed-and-option combinations one by one and computed the means from the per-customer records. Laydyne earns its place in simulations where the floor matters — walking routes, shared resources, many steps — which this task did not test.

Managing revisions across sessions: Codex alone, again. Asked in three separate sessions to issue a revision, change it and issue the next one, then reopen the first exactly, Codex alone kept each revision as a full copy of the project file with a checksum and reopened it by copying the file back (52–57 s for that session). With Laydyne the agent kept revisions as separate projects in the browser's storage instead of as design options inside the project, so it had no built-in diff: it re-read whole saved projects to compare them, and reopening took 117–121 s, about 2.2 times as long and 4.1 times the input tokens (medians). Both setups got every session right.

So: "far more token-efficient" is not true for anything we measured, small or large. Laydyne was faster on big models and drawings, and its drawings and saved options come for free, but an agent working alone managed a 2.6 MB, 20-floor project file without breaking anything. Untested still: clients that load every tool with every request, people editing the model alongside the agent, and simulations that depend on the floor.

When Laydyne is the better tool, and when it is not

Across the nine tasks we measured, Codex alone used fewer tokens every time, at the same scripted quality. After our changes the gap is small on some (the shop: about the same estimated cost; the change to the 20-floor model: 1.4 times the cost) and large on others (a textbook queue: 3.4 times the cost, because writing that simulation yourself is the easiest kind of task). If tokens are all that count, use the agent alone.

What Laydyne gave in the same runs:

  • Time on bigger work. 0.73 times the time on the change to the 20-floor model, 0.76 on the scanned drawing, 0.92 on the clean drawing.
  • Deliverables nobody asked for. Its plans come out at a scale, with dimensions, room areas, a legend of what changed and a title block; the plans the agent wrote alone had none of these.
  • A model people look at and correct, and checks you do not write. The model is live in a browser tab: a planner can move a shelf by hand, the agent reads it back, and Undo works for both. check_space finds narrow passages, blocked doors and unreachable areas.

Where we are not there yet:

  • Revisions across sessions. The agent kept the revisions as separate browser projects, where the diff tools did not reach, and reopening one took longer than restoring a file. Issues are now recorded in a register with the sheets and files they sent (drawing revisions).
  • The fixed cost per request. Most of Laydyne's extra tokens come from the agent re-sending its own context with every request and reading the tool list. Fewer, larger steps help more than smaller results.

We will keep measuring with the same method and publish the numbers either way.

Try it yourself

Run the same task in two empty folders and compare the turn.completed usage lines of codex exec --json:

Design a small shop: a 12 m × 9 m room (12 m west–east, 9 m north–south). The front is the south side, with a 1.8 m wide entrance in the middle of the south wall.
Put in:
- a checkout counter near the entrance;
- two rows of shelving, each row at least 6 m long, shelves 0.5–0.6 m deep, with a 1.2 m clear aisle between the rows;
- a display table near the entrance.
Every passage must be at least 0.9 m wide, nothing may overlap, and the entrance must stay clear.
Check your result (overlaps, passage widths, dimensions) before you finish.
Deliver:
1. plan.svg in the current folder: a scaled 2D plan with the overall dimensions and the main aisle width dimensioned.
2. A 3D view of the shop.
3. In your final answer: what you placed, the narrowest passage, and what you assumed.

Add for Codex alone: "Work with files in the current folder only. Keep the design in model.json, make plan.svg from it, and for the 3D view make view3d.html with three.js from a CDN." Add for Laydyne: "Use the Laydyne MCP tools: build the shop in the Laydyne Studio I have open. Save the 2D plan as plan.svg (get_view_image with format "svg") and look at a 3D view with get_view_image."

codex exec --json --skip-git-repo-check -C ./alone - < prompt-alone.txt | grep turn.completed
codex exec --json --skip-git-repo-check -C ./laydyne - < prompt-laydyne.txt | grep turn.completed

The steps of both setups, side by side, are in flow.json.