---
type: plan
---

# Brief: design the review process we go live with

**For a deep-thinking session, which will produce a stream that implements what it recommends.**
Written 18 Sep 2026 by rian and Claude, at the end of a long session rebuilding the review screen.

---

## The mission

**Produce the best possible catalogue review process**, then write the stream that builds it.

The sequence after that is fixed: the stream runs → a **full collection sweep** → then **a full
series of review passes until no decisions remain** on the set of listings mosiah holds. What comes
out the other side is the clean catalogue we go live with. So this is not an incremental improvement
to a screen. It is the one chance to get the machinery right before it is pointed at everything.

## What this is NOT

**We are building the room, not doing the work in the room.**

Nothing here should decide a catalogue question. Gendered bottles, eau de parfum against eau de
toilette, whether a tube is packaging — those are decisions rian makes *at the review*, on the case
in front of him, and the guidelines record what he decided rather than deciding first. A draft of
this brief made that mistake: it turned three examples rian happened to be looking at into standing
rules, and he pulled them. **There are probably a hundred more factors like those, and almost all of
them are unknown today. The deliverable is a process that reliably notices one, stops, and gets it
written down properly — not a ruling on the handful we have seen.**

## Read these first; this brief deliberately does not repeat them

| | |
|---|---|
| `main/docs/REVIEW-PROCESS.md` | the certain boundary, the file format, every grouping default |
| `main/docs/AI-REVIEW-GUIDELINES.md` | what a pass follows, and the precedents answers have set |
| `.claude/commands/review-pass.md` | the pass skill as it stands: one layer, one question per precedent |
| `.logs/planning/review-simulation-2026-09-18.md` | **the evidence base.** A full simulation, seven decisions across five brands on a throwaway copy, all undone afterwards. Read the section headed "The structural finding" |
| `main/docs/COLLECTORS.md`, `main/docs/DATA-MODEL.md` | how listings are read and what a listing holds in three layers |
| `main/docs/VOCABULARY.md`, `agents.md` | one term per concept; the rules that are failure-backed |

---

## Strand 1 — make the raw data worth reviewing, before the sweep

The single largest finding of the simulation, and it is not a review problem at all.

**The retailers split into two populations that cannot be joined to each other.**

| | listed wording | barcodes |
|---|---|---|
| Bordershop, Dubai Duty Free, Heinemann ×3 | **none** | **98–99%** |
| Attenza, ARI, The Loop, Shilla, iShopChangi | **100%** | **6–18%** |
| Extime, Avolta | 100% | 79% / 37% |

3,847 product variants live only in the first group; 4,253 only in the second; **200 in both**. And
**90% of the catalogue — 15,035 of 16,719 product variants — is stocked at exactly one retailer.**
The first group publishes internal shorthand for names (*Joh.Walk Bl L PET*, *Glenfid 15 V3 GP*), so
a fresh collection fills their wording with that shorthand. The names are not missing; they are
hostile. Proven on one bottle in the simulation log.

Known and settled already, so do not re-derive:
- **The collectors are consistent.** All nine attach a raw fragment on their single construction;
  one reader turns it into the listed columns for all eight platforms; name and brand come back at
  100% on every platform on fresh data.
- **The gaps are staleness, not capability.** Fragment keeping began 5 September; it is 100% either
  side of that date. A sweep closes it.
- **Retailer markup is no longer stored** (124 MB of tile HTML removed, gated by a test).

**So the question for this strand is narrower than it looks:** given a sweep is coming, what else
should change in what we keep or derive, so the passes have more to work with? Candidates the
simulation raised, none of them decided: whether to derive a normalised name on the barcode-rich
side; whether the parser version should move when the kept payload shape changes (it did not);
whether 302 empty product lines and the 43 that undo leaves behind should exist at all.

## Strand 2 — better recommendations, in the best order

The pass currently works one layer at a time, outermost first: brands, then product lines, then
product variants, then attributes. The simulation showed that order is **wrong in one specific way**:
cross-divide product variant matching (joining a barcode-rich row to a name-rich row) has to happen
*before* the product line work, because it changes what the product lines contain. That layer does
not exist today.

The human cost is the constraint. A person cannot make thousands of decisions carefully and
consistently. What worked in the simulation and should survive: **one question per precedent, with
the held-back ones named**, so the reviewer can say "that one is different" before answering. Three
questions covered 57 product lines on one brand.

## Strand 3 — the skill must get better with every decision

Each answer should make the next pass smarter, and **the improvement has to live in written
artefacts, not in a model's memory** (see Constraints). Today a pass is instructed to read the
decisions and notes and rewrite the guidelines before proposing anything; whether that is enough, and
whether it should be mechanical rather than a matter of the pass remembering to, is open.

One elegant property to consider: if a note is required on every critical decision (see the ladder
below), then **the notes ARE the precedent record**, captured by construction rather than by
discipline. That suggests generating the guidelines from the notes rather than maintaining them by
hand.

## Strand 4 — notice the unknown, and this is the sharpest requirement

**Rian will approve in bulk.** Nobody reads a thousand cards one at a time. So anything the pass
gets subtly wrong sails straight through — and that is as true of the routine-looking rows as the
flagged ones.

His worked example, which arrived within the first few decisions of the very first brand:

> A product line called *"for men"*. The ordinary process folds it somewhere or gives it its own
> line, rates it low, and it is waved through. But underneath is a question nobody has answered:
> **what do we do with gendered bottles at all?** And a trap inside that — there may be a "for men",
> a "for women", **and an unmarked one that is actually the women's**, because the women's is the
> default. Anything assuming "unmarked = the parent" gets it backwards.

**So being accurate on average is cheap; being loud about novelty is the job.** A missed new
category does not cost one wrong row, it costs every row of that kind, silently, for ever. The pass
must recognise when a question is *a new kind*, refuse to treat it as routine, put it up as a
precedent to set — and then watch for that factor in everything afterwards.

*The unmarked member of a set* is worth naming as a general pattern. It will recur far outside
fragrance: a whisky with no age statement beside aged siblings, a "Classic" beside named variants.

---

## The approval ladder — rian's starting proposal, open to a better one

**Attention today is a label that gates nothing.** This turns it into a control, which means the
pass's rating becomes load-bearing.

| level | bulk? | what answering takes |
|---|---|---|
| **critical** | never | confirmation **and** a note, always |
| **important** | never | one click, note optional — **except "something else", where the note is the answer** |
| **medium / low** | yes, as a group | one confirmation for the group, with an optional note applying to the whole group |

Rian: *"That's just my idea, but Fabel may have a better idea of how to handle this, and I think more
importantly, a good way to determine the level of attention needed per decision."*

**The bar must be adjustable.** If the first sweep produces hundreds of criticals, the level has to
be able to move without rewriting the process.

**The harder half is rating, not gating.** Attention should fall as precedent accumulates and the
skill grows confident — a question that is critical the first time is routine the fiftieth. The
mechanism for that decay, and what makes it trustworthy, is the real design problem.

---

## Hard constraints

1. **Model independence.** This round runs on staging with Claude as the AI, driven by hand. Live,
   a user presses **"run a pass"** or it is scheduled, through an API, **not tied to a Claude session
   on mosiah**. The skill must therefore be a specification any model can execute: explicit inputs,
   explicit outputs, explicit stopping conditions, no reliance on conversational context. This
   interacts directly with Strand 3 — "self-improving" must mean *the written artefacts improve*.
2. **The legal posture is not negotiable.** Facts, never expression; no login; a block is a refusal;
   robots.txt re-read every run. `agents.md`.
3. **Terminal condition.** The passes must converge: "until no decisions remain" needs each pass to
   leave the open set smaller, and a defined done.
4. **Every decision stays undoable**, and the data comes back. Proven: 12 batches, 126 decisions, 0
   skipped, every product variant and listing restored.

## Open questions this brief does not answer

1. **Naming.** The ladder says *important*; the code's levels are `critical / high / medium / low /
   none`. One term per concept — which word wins?
2. **Who checks the rating?** If the pass both assigns attention and detects novelty, nothing
   independently catches it under-rating something new as medium, where bulk approval then hides it.
3. **Does bulk approval need a scope smaller than "the rest"?** Rian's ladder bulk-approves a *group*
   of medium and low. What defines the group — a precedent, a product line, a brand?
4. **Convergence.** If answering a question can raise new ones, what guarantees the passes end?
5. **Should the guidelines be generated from the decision notes** rather than hand-maintained?
6. **How does a person correct a precedent later**, once fifty decisions have been made under it?

## The shape of the work over time

**This first batch is the largest headache the human reader will ever face**, and that is expected.
Once a wide range of decision types, brands and product lines has been covered, new listings and
changes to existing ones produce fewer decisions needing less attention. It becomes a big job again
only when a new category is added or the collection sources expand dramatically.

Design for that curve: the process should be tuned for a hard first pass and a light steady state,
not for a constant load.
