---
{
  "n": 70,
  "title": "The demo's three scenarios",
  "abstract": "",
  "refs": [
    "todo:2285",
    "message:13532",
    "doc:69"
  ],
  "seen": [
    "agent",
    "user"
  ],
  "data": {
    "written": true,
    "environment": "main",
    "revisions": 1,
    "open_until": 1790894178.418787
  },
  "created": 1790892378.40346,
  "updated": 1791140105.8721302,
  "deleted": 0.0,
  "completed": 0.0,
  "outcome": "",
  "type": "doc"
}
---
Every scenario is a real session: I play the agent and the user through the real journal while it is recorded. In the demo the visitor plays the user: each message waits in the box until they press Send, each question waits until they click an answer, and each plan waits until they press Approve the plan. A question's answers each lead to their own recorded ending; nothing else branches.

## 1. Crumb & Co. — a bakery's website
Shows: a plan approved with its button, to-dos worked one by one, a question that changes what gets built, a fact the agent keeps.

Shared start:
1. Visitor sends: "Hi! Can you build a simple website for Crumb & Co., our neighbourhood bakery? A home page, our menu and a contact page."
2. The agent plans three phases (The pages, Visiting, Phones), files five to-dos under them, and marks the plan ready.
3. Visitor presses **Approve the plan**. The agent starts it and builds the stylesheet and the home page.
4. Before the menu the agent asks: **Should the menu show prices?**

Answer A, "Yes, show prices":
- The menu gets a price column beside each item.
- Phone check: the price column squeezes the names on a phone. The agent notices, logs it, and stacks the price under each name below 600px.
- Visitor sends: "Remember: we're closed on Mondays." The agent keeps it as a fact and puts it on the contact page.
- The plan finishes; the agent sums up, naming the phone fix.

Answer B, "No, just the names":
- The menu lists names with a one-line description, and a note: "Ask at the counter for today's prices."
- Phone check passes the first time.
- Visitor sends: "Remember: we're closed on Mondays." Same fact and contact page.
- The plan finishes; the agent sums up, naming the counter note.

## 2. Pebble Pantry — helpers beside Claude
Shows: helpers on Codex and Claude in their own environments, orchestrator mode, a helper going off track, bringing a worktree back.

Shared start:
1. Visitor sends: "Pebble Pantry scales recipes badly: 2.5 guests gives 3 eggs and 0.33333 cups. Plan the fix, and tell me which parts could go to helpers."
2. The agent plans the rounding, the unit tests and the README, and says which suit helpers.
3. Visitor presses **Approve the plan**.
4. Visitor sends: "Send the unit tests to Codex in its own worktree, and the README to a Claude helper. You take the rounding." Quill (Codex) and Wren (Claude) start; the agent works the rounding.
5. Visitor sends: "From now on you only plan and review." Orchestrator mode turns on.
6. Wren's README section jokes instead of showing an example. The agent asks: **Wren's README is off track. What should happen?**

Answer A, "Stop Wren, send a fresh helper":
- Wren is stopped, and Moss (Claude) writes the README again and reports.
- Quill reports; the agent brings Quill's worktree into main.
- The review finds grams aren't rounded; the agent files a suggestion and sums up who did what.

Answer B, "Let Wren fix it":
- The agent tells Wren what is wrong; Wren rewrites the section with a real example and reports.
- Quill reports; the agent brings Quill's worktree into main.
- The review finds grams aren't rounded; the agent files a suggestion and sums up who did what.

## 3. Ledgerly — a customer's bug report
Replaces Away from the desk. Shows: a bug traced with a failing test, the work log, a question that decides the fix, a rule or a fact kept for next time, a report, and a check that guards it.

Shared start:
1. Visitor sends: "A customer says invoice 2041 is a cent off: three lines at €9.99 with 21% VAT. Can you find out why?"
2. The agent opens work, reads the invoice code, and writes a failing test that reproduces it: the code rounds the VAT on each line, the customer's accountant expects it on the total.
3. The agent asks: **Round the VAT per line or on the invoice total?**

Answer A, "On the total":
- The agent changes the code to add the lines first and round the VAT once; the test passes.
- It keeps a rule: VAT is rounded once, on the invoice total.
- It writes a report: what went wrong, which invoices since March are affected (four), and the fix.
- Visitor sends: "Make sure this never comes back." The agent adds a check that runs the invoice tests, and it passes.

Answer B, "Per line, as we do now":
- The code stays; the agent fixes the customer's expectation instead: the PDF now shows the VAT on each line so the sum is visible.
- It keeps a fact: VAT is rounded per line, as the tax office allows, and the PDF shows it.
- It writes a report: no amounts were wrong, the invoice was unclear, and what changed on the PDF.
- Visitor sends: "Make sure this never comes back." The agent adds a check that runs the invoice tests, and it passes.
