Fixed brief · Static-site evaluation · 09 Aug 2026

Same prompt. Five interpretations.

Five agents received one coffee-roastery brief. Their finished builds appear below, interactive and unaltered.

Control document

Identical input for every run
Original prompt
Build a complete marketing website for "Meridian", a fictional independent coffee roastery. This brief is fixed; your design interpretation of it is the test.

The site must include, at minimum:
1. A hero section introducing the brand.
2. A section presenting three coffee origins (invent the details: name, region, tasting notes, price).
3. A short "our philosophy" or about section.
4. A footer with fictional contact details and a newsletter signup (front-end only, no real backend needed). The footer must also display, in small text, the basename of the current working directory (the directory's name only, not the full path) as a build identifier — hardcode it into the footer at build time.

Rules:
1. You may use any languages, frameworks, libraries, or build tools you like.
2. All content must be self-contained: no external network requests at view time (no CDN links, no remote fonts, no hotlinked images). Inline SVG, locally bundled assets, and generated graphics are all fine.
3. The final deliverable must be a static site: a directory named `dist/` in the current working directory containing an index.html that renders fully when served locally. If you use a framework, run its build step yourself so dist/ is the finished output.
4. Include a README.md in the current working directory stating: your chosen tech stack, your aesthetic direction (3–5 lines), and the exact command to serve dist/ locally (e.g. a one-liner static server).
5. The site must be responsive and look intentional at 375px and 1280px widths.
6. Commit to a distinct aesthetic direction and execute it consistently. Generic template-like output will be judged as a failure of taste.

STRICT DIRECTIVE — SANDBOX BOUNDARY:
Work ONLY inside the current working directory. Do NOT read, list, access, or modify any parent directory or any sibling directory, under any circumstances. You may create and use temporary subdirectories inside the current working directory (including node_modules, build caches, etc.). Reading the NAME of the current working directory itself (e.g. via pwd/basename) is explicitly permitted and required for the footer identifier. Do not inspect the environment beyond what is required for this task.

Live specimens

Viewport controls reproduce the two required widths

OMP harness

Baseline · no Trellis

Muse Spark 1.2

Cartographic atelier · omp-muse

Elapsed
1m 36s
Tokens
489,835
API-eq. cost
$0.00
Usage breakdown

PrimaryMuse Spark 1.2 Contributor

Input38,989 uncached · 433,213 cached

Output17,633

Cost noteProvider usage record reported $0

Open result

Loading live result…

OMP harness

Baseline · no Trellis

DeepSeek V4 Flash

Field guide to coffee · omp-dsfv4

Elapsed
47m 24s
Tokens
25,426,616
API-eq. cost
$11.80
Usage breakdown · advisors included

Primary · DeepSeek V4 Flash8,152,535 tokens · $0.032699

Default advisor · GPT-5.6 Sol12,561,768 tokens · $9.266131

Design advisor · Kimi K34,712,313 tokens · $2.503538

AccountingAll advisor calls completed before delivery are included

Open result

Loading live result…

OMP harness

Baseline · no Trellis

GPT-5.6 Sol

Vivid field atlas · omp-sol

Elapsed
9m 23s
Tokens
1,420,057
API-eq. cost
$1.69
Usage breakdown

PrimaryGPT-5.6 Sol

Input73,559 uncached · 1,324,416 cached

Output22,082

Advisor notePost-delivery advisor activity excluded

Open result

Loading live result…

Codex harness

Baseline · no Trellis

GPT-5.6 Sol

Navigation workshop · codex-sol

Elapsed
18m 12s
Tokens
2,535,035
API-eq. cost
$2.63
Usage breakdown

PrimaryGPT-5.6 Sol · xhigh effort

Input89,893 uncached · 2,412,800 cached

Output32,342 · reasoning included

Time note15m 20s active · approval wait included in elapsed

Open result

Loading live result…

Claude Code harness

Baseline · no Trellis

Claude Opus 5

Roast log · cc-opus

Elapsed
37m 20s
Tokens
13,347,991
API-eq. cost
$15.59
Usage breakdown · advisor included

Primary · Claude Opus 513,054,747 tokens · $11.972357

Advisor · Claude Fable 5293,244 tokens · $3.612880

Primary cache12,633,541 reads · 321,822 writes

Primary output99,229

Open result

Loading live result…

Public preference poll

One current choice per anonymous browser

Community readout

Which result do you like most?

Vote on the execution, not the model name. You can change your choice; the public totals update with it.

Loading public votes…

  1. 01 Muse Spark 1.2 OMP · Baseline · no Trellis 0 0%
  2. 02 DeepSeek V4 Flash OMP · Baseline · no Trellis 0 0%
  3. 03 GPT-5.6 Sol OMP · Baseline · no Trellis 0 0%
  4. 04 GPT-5.6 Sol Codex · Baseline · no Trellis 0 0%
  5. 05 Claude Opus 5 Claude Code · Baseline · no Trellis 0 0%

Reading the run log

Measured, not guessed.

Time runs from the original prompt timestamp to the first completed delivery of the finished site. Later “serve it” requests are excluded. Codex includes its visual-direction approval wait.

Tokens include uncached input, cache reads, cache writes, and output. Advisor calls completed before delivery count toward the run that invoked them.

Cost is API-equivalent, not a claim about a subscription’s marginal charge. OMP values come from persisted usage records; Codex and Claude Code values are calculated from their persisted tokens and listed prices on 09 Aug 2026.