---
{
  "n": 76,
  "title": "One repository for rows, kept in memory",
  "abstract": "",
  "refs": [
    "todo:4042",
    "todo:4039",
    "message:24590",
    "message:24593"
  ],
  "seen": [
    "agent",
    "system",
    "user"
  ],
  "data": {
    "written": true,
    "status": "final",
    "environment": "main",
    "revisions": 2,
    "open_until": 1791638805.371768,
    "buttons": [
      {
        "label": "Make a thorough plan of option B",
        "say": "Make a thorough plan of option B",
        "choice": "answer",
        "ask": "How far do we go now?"
      },
      {
        "label": "Make a thorough plan of option C (all phases)",
        "say": "Make a thorough plan of option C, all phases",
        "choice": "answer"
      },
      {
        "label": "Do option A now, plan the rest",
        "say": "Do option A now and plan the rest",
        "choice": "answer"
      },
      {
        "label": "Change it first",
        "say": "I want changes to the proposal first",
        "choice": "answer"
      }
    ],
    "pick": 1,
    "pressed": [
      "Make a thorough plan of option C (all phases)"
    ]
  },
  "created": 1791624826.760297,
  "updated": 1791637005.372098,
  "deleted": 0.0,
  "completed": 0.0,
  "outcome": "",
  "type": "doc"
}
---
Written by Mr. Babbage (Fable), read only, for Alfred, 2026-10-10.

## Where rows live today, and every path that reads them
A row is a Markdown file with a JSON head: one per row, under `<home>/<type>/NNN.md`. Beside the rows sit:

- an index of summaries (`.index/index.json`)
- a `changes.log` of row numbers
- packed archives of old closed rows
- an `events.jsonl` per environment, holding every write as a numbered event that carries its type, n, environment and pid

`RowStore` (src/controllers/stored.py:144) is already the one place that reads row files, and every controller holds one as `rows`. It keeps process-wide caches: stamps, summaries, derived lists, running totals, and parsed rows. Loaded rows already stay in memory. But reads still touch disk:

- every `summaries()` call does two stats
- every `standing()` call stats every row it returns
- every `load` deep-copies the row

The four lazy types (doc, report, browser, nudge) are parsed on every read.

Places that reach past the funnel or defeat it:

- **Rows loaded in loops.** `agent_session` loads the ticket it was just handed (tickets/controller.py:248), and `_at` loads one row per dependency per `_waiting_on` call (:448).
- **A fresh `Record` per call.** About eight places build a new `Record(root, ticket.work_environment)` (tickets/controller.py:138, :175, cards.py:94, :115, :199, :204), and a fresh Record has no request memo.
- **Session lookups.** `Sessions.holder` walks every session and checks each pid, with no memo (engine/sessions.py:234), once per ticket.

**Why one board load costs 158 Records and 73 plan checks for 39 tickets.**

- `board()` computes every card's state.
- `_slots` then computes every card's state a second time.
- `_actions` asks the plan status a second time.
- `_running()` is computed three times.

Each pass builds new Records and walks all sessions with pid checks. That comes to 2.3 s in a quiet process and 69 s on the loaded machine. The cost is repeated work, stats and pid probes, not the number of rows.

## The repository design
Keep `RowStore` and the name `rows`; it already is the funnel. Change what it promises:

- **Reads never touch disk once a folder is held.**
  - `peek` checks the held row against its in-memory stamp instead of stat-ing the file.
  - `summaries` uses a per-folder version number, bumped by every write and refresh, instead of two stats.
  - `exists` answers from the summaries.
- **Writes go to disk and patch memory in the same call**, as they do today.
- **A per-type memory policy** declared on each resource:
  - `held`: how many parsed rows stay in memory, as a rolling cache with the oldest evicted;
  - `eager`: parsed at server start.
  - Suggested defaults:
    - tickets, boards, environments, plans, agents, works and to-dos are eager;
    - messages and comments keep 500;
    - docs and reports stay lazy with 32.
- **Other processes' writes** reach memory through the event log. The server's watch thread already ticks every second; it tails each environment's events and refreshes just the one row a foreign event names. The periodic restat stays as a safety net for writes made without an event.
- **Counters.** Declared per type, adjusted on every create, delete and refresh, and kept in the index. They are never rebuilt from rows.
- **Pagination.** `page(where, order, after, size)` with a keyset cursor, so a write between pages never shifts a page. Boards page per lane, and card state is computed only for the cards on the page.
- **The guard in the tests.** The test kit already counts file opens and folder scans per call. Add two channels:
  - rows parsed;
  - Records built.

  The tests then fail when:
  - a warm call parses anything;
  - one call parses the same row twice;
  - a parse recurses into the same folder;
  - a board builds more than one Record per environment it touches.

## Migration, smallest risk first
0. **Measure.** Add the parsed and Records counters and a board benchmark to `journal speed`. No behaviour change.
1. **The board stops repeating itself.**
   - Compute sessions, running agents, each environment's Record and the board once per call, and pass them down.
   - Memoise plan status, waits and the board within the call.
   - A `Record.sibling(env)` shares the request's memo.
   - Session liveness is checked once per sessions stamp.

   Plan checks go from 73 to 39 or fewer, Records from 158 to about one per environment, and pid probes from sessions × tickets to sessions.
2. **Reads never stat.** Folder version, in-memory stamp check, `exists` from summaries. Proven by a warm second board load opening and stat-ing nothing.
3. **Event-driven refresh.** The watch thread tails the event logs and refreshes rows written by other processes. Proven by a write from another pid becoming visible within one tick with no folder scan.
4. **Per-type policy and counters.** `held` and `eager` on each resource, rolling caches, and persisted counters. Proven by a held=2 type parsing a third row again, and by counters after a restart matching a fresh count.
5. **Pagination.** Board lanes in pages of 50 in the server and the viewer.

Each phase ships alone and leaves everything working. Phases 1 and 2 give back most of the time.

## Options
- **A. Phases 0-2.** The board stops repeating itself and reads stop stat-ing. Expected: the 2.3 s board drops under 0.3 s, and the 69 s case collapses. Lowest risk.
- **B. Phases 0-4 (recommended).** Adds event-driven refresh, the per-type memory policy and real counters, which meets every requirement except pagination. Medium risk: the refresh is new cross-process behaviour, but the periodic restat stays beside it.
- **C. Phases 0-5.** Everything, including pagination in the board API and the viewer.

## Benchmark of plan 31, main against release-2690
Measured with `journal speed --board`: a journal built for 40 tickets, 90 sessions and 7,800 messages, two rounds each, on a loaded machine (milliseconds swing by a factor of two or more between rounds; the counts are exact).

| Measure | main 2.268.5 (round 1 / 2) | release-2690 (round 1 / 2) |
| --- | --- | --- |
| Ticket board of 40 tickets, ms | 559 / 800 | 982 / 1,156 |
| Ticket board, pid probes | 320 / 320 | 14,560 / 14,560 |
| Status of 40 tickets, ms | 5,416 / 4,357 | 13,978 / 8,621 |
| Status of 40 tickets, pid probes | 3,280 / 3,280 | 288,080 / 288,080 |
| List of 7,800 messages, ms | 2,330 / 1,944 | 24,322 / 5,479 |
| Claude PreToolUse hook, ms | 33 / 13 | 239 / 19 |
| Record reads, board / status | 204 / 160 | 204 / 160 |
| Memory after timing, MB | 66 / 62 | 74 / 77 |

Finding: release-2690 makes 45 to 88 times more pid probes. `Tickets.agent_session` calls `Sessions.holding()`, a pass over every session with a liveness check on each, once per ticket; the memo that should make it one pass per request is absent when a record has none (this benchmark, and any command run outside the server), so 40 tickets times 90 sessions is 3,600 probes per status. Main asked `holder(env)`, which probes only the sessions of that one environment. The list of messages is slower on release-2690 and varies a lot between rounds; it needs its own profile.

## Benchmark after the two fixes on release-2690
Commits 945ff91ac (a ticket's agent is found with holder(env) when the record keeps no memo) and bc1539b69 (a folder keeps 10,000 parsed rows instead of 500). One round: ticket board 1,018 ms with 320 pid probes (main 320); status of 40 tickets 4,653 ms with 3,280 probes (main 3,280); list of 7,800 messages 2,243 ms (main about 1,900-2,300); hook 97 ms under load (main 13-33); memory after timing 63 MB (main 62-66). The list was slow because a listing of more rows than the folder keeps (500) read every row again each time; main kept all of them (second listing of 700 messages: main 0 rows read, release 700).
