# Issues & ideas register — internal running list; not client-facing
last swept 2026-09-05 (Stream A2)
<!-- Role: the topical register. handoff.md is the chronological story; this is
     "what is open, watched, or parked" at a glance. Update in place: move items
     between sections rather than appending duplicates. Client-facing framing
     lives on /discuss and is deliberately narrower than this file. Every figure
     here is a dated measurement, not a current count; ask /api/stats for those. -->

Last swept 2026-09-05 (Stream A2). `main/docs-check.sh` warns when this line is older than a week.

## P1 — blocks the launch. Decide or fix before anything else.

- **[P1] The four Shopify shops (YUL, PTY, BOG, SAL) refuse our identity at robots.txt.**
  Measured 2026-09-04: HTTP 403 to the declared UA, 200 to a plain curl, all four hosts. Last
  good runs 22 Aug. Under the ratified robots policy (401/403 is a refusal) `robots.py` raises
  `SourceBlocked` before the first feed request, so their next run records `blocked` and they
  fall under the BLOCKED rule (dated prices stay visible). Whether `/products.json` is also
  refused is unchecked. *Decision for rian:* `.logs/decisions-for-rian.md` #10 (accept, check
  the feed, or ask the operators for a `permission_record`).

- **[P1] Dubai is refusing our bot by NAME, not by IP.** Diagnosed 2026-09-03.
  Same host, same second: our declared UA -> HTTP 403, `Mozilla/5.0` -> 200.
  Their `/robots.txt` is refused too and the error page is
  `errors.edgesuite.net`, so this is Akamai bot management at the edge, not the
  application. Last good run Aug 22 00:19 (1,527 products) with the SAME UA, so
  their rules changed, not ours. Swept the other 15 launch sources with our real
  identity: all 200. Dubai is the only refusal.
  *Why it is P1:* DXB is 1,489 products (39% of the launch 3,771) and 1,477 of
  1,915 barcodes (77%).
  *Not bypassed, deliberately.* Evasion is one line of code and it is the wrong
  move while rian is building retailer relationships.
  *Decided 2026-09-04 (Decision 3):* the bot identity page becomes its own static
  project, never an alias on the app; the UA is repointed once it serves; the ask
  to Dubai when rian judges the relationship ready; Dubai launches on dated 21 Aug
  prices. Still open: the hostname (§10 #3) and the BLOCKED rule ratification (§10 #5).

- **[P1] Every launch price is stale.** Newest observation anywhere in the
  sixteen is 2026-08-25; DXB is 2026-08-21. A full recollection is required
  before the airports are switched on. The last-checked date on every price
  makes this visible to visitors, which is the point of showing it.
  *Order (Decision 4):* targeted beauty first, full crawls last; the drinks refresh
  of already-collected sources is exempt and runs once `robots.py` has landed (it did,
  2026-09-04) and migration #1 is deployed.

- **[P1 → P2] Seoul (Shilla): the browser passes; the page is server-rendered.** Probed
  2026-09-05 (Stream A2, `.logs/runs/shilla-probe-2026-09-05.log`): one product page rendered
  through the sidecar with our declared UA answered HTTP 200, no `cf-mitigated`, a full product
  document (name, brand, list/discount/online-member USD prices, a UPC in "REF.NO", size, stock,
  category). The 4 Sep 403 challenge was Cloudflare judging the text client, not our identity.
  Caveats: robots.txt (`*`, Crawl-delay 5) disallows `/files/`, `/medias/` and `/estore/_ui/`,
  where the page's own scripts and imagery live, so the sidecar refuses those requests and the
  collector reads the document at `domcontentloaded`; the sitemap is empty, so discovery must
  walk category pages (needs its own announced probe); three price tiers exist and the
  collector must record which one it publishes (`price_type`). Decision 5's replacement-airport
  probe (§10 #9) is not needed unless rian prefers it. *Built 2026-09-05:* `shilla.py`
  (location ICN, USD as displayed, discount tier published, list crossed out). Discovery is
  the honest limit: the category grids are drawn by robots-disallowed scripts, so the walk is
  home page + related rails under a 200-render budget; a full catalogue needs the partnership
  ask (their robots.txt now names many AI/rendering crawlers; we are not named).
- **[P1 → P2] Singapore (iShopChangi): collector built on the rendered fetch; first run
  waits for the sidecar deploy.** Probed 2026-09-05 (`.logs/runs/changi-probe-2026-09-05.log`):
  four renders of one product page. The shell draws nothing until the Adobe Launch script has
  loaded (the bundle dereferences a global it defines), so that host is declared as this
  source's one asset host; the page then fetches its own `pdp/inventory.json`, which the
  collector reads as received. Three channel prices per offer (Departure = Arrival = S$271.10,
  non-traveller S$333.87 on the probed bottle); we publish Departure as `price_type`
  "departure". **No barcode anywhere** (page, inventory, tiles): SIN matches on brand, name and
  size only. Marketplace: each offer names a seller (kept in raw). Targeted set seeded (622
  URLs, `.logs/runs/changi-targets-2026-09-05.md`). Open: rian's hand check of the tier
  (decisions #14), the tag-manager assumption (#13), listing-page harvest via category pages
  for the full catalogue (category sitemaps exist; not probed).

## P2 — silently corrupting data. Fix before the hub pages ship.

- **[P2] A number in a beauty name is read as a size, and two Extime rows carry the 7-litre
  misread.** Found by the first `app.cli audit` (2026-09-05, this morning's dump): Extime
  "DELICIA DRENCH 59" stored as 5,924 ml and "CHEIROSA 76" as 7,624 ml (perfume names, no
  size in them); "15yo Sherry Cask" and "18yo Port Cask" at 7,000 ml (the "0.70cl" reading
  the size veto now refuses at ingest, but these rows predate it). The audit's
  `oversize_singles` metric lists them; threshold pinned at today's six so a new one is red.
  *Owner:* Stream A (`parse_size_ml` on beauty names; a rederive for the four rows).

- **[P2] The Shopify drinks filter is English-only.** `DRINK_HINTS` contains
  "liquor" but not "licor", so Panama's own `Licores` shelf is invisible: only
  181 of 395 pass `_is_drink`. Spanish shelf names (Vinos, Rones, Champanas,
  Cordiales) fail everywhere. We are dropping liquor at PTY and BOG right now.
  *Owner:* Stream A, Wave 2 (Shopify Spanish and beauty shelves).

- **[P2] 110 product rows are duplicates of another row** (52 match_key groups, measured
  2026-09-03). A bottle appears as two products with two prices instead of one product with
  a comparison. This corrupts the AggregateOffer markup the SEO plan is built on, so it has
  to be fixed before that markup is published.
  *Owner:* Stream A, Wave 2, as recorded merges (Decision 6: `merged_into_id`, never deletes).

- **[P2] Brand names are not normalised, and brand pages depend on it.** Measured
  2026-09-03 at 16-airport scale: "Don Julio" / "Don Julio Tequila" / "Don Julio®" are three
  brands; "Moet & Chandon" and "Moët & Chandon" are two; six pairs differ only in case or
  accent. Each would generate a competing brand page. The same family: **name-matching
  splits when there is no barcode** ("Don Julio 1942 Tequila 750ml" vs "Don Julio 1942 (R)
  Tequila 750 ml"; the Dublin gap). Safe direction (splits never create fake savings, they
  undercount comparisons) but it thins the comparable set and looks sloppy.
  *Owner:* Stream A, Wave 2 (Decision 6: brands become a table, migration #3; the brand fold).

- **[P2] Launch products with no category fall out of the category structure** (about 260
  measured 2026-09-03; some are the already-collected beauty products at Athens).
  *Owner:* Stream A, Wave 2 (the uncategorised backfill, an idempotent CLI command).

## Open — data quality

- **Extime is Paris, not CDG alone.** The storefront is one "paris" tenant
  covering Charles de Gaulle AND Orly; the payload's only terminal signal is
  the combined value `["CDG_ORY"]` and there is no per-airport sitemap. The
  location is therefore named "Paris Charles de Gaulle and Orly" with iata
  CDG. If Adam needs CDG-only pricing, that is an unmet requirement, not a
  bug to fix later.
- **Extime `services.xml` returns HTTP 500** (reproduced twice). Some service
  "products" are invisible to sitemap discovery. Out of scope for drinks, but
  it means sitemap coverage is not catalogue coverage.
- **Never verified on Extime:** an out-of-stock product's rendering, a genuine
  multi-variant product, and the /fr/ tree. The collector handles all three
  defensively (skip rather than guess) but the paths are untested against
  reality.

- **Gratien & Meyer JFK $77 vs Europe ~$21** (found 8/25, rian's human check).
  Both prices verified real on Avolta's own pages (JFK microdata says 77.00).
  Ratio analysis across all 304 JFK/Europe shared products: median 1.07x,
  p90 1.52x, max 3.74x = this bottle. Verdict: almost certainly Avolta's own
  pricing-entry error ($17 -> $77 typo would land on the European price).
  Standing: kept on its product page (it IS the published price); excluded
  from featuring by the corroboration rule (spread >3x needs >=3 shops).
  Watch: does it correct on a future JFK crawl?

- **Per-variant listings carry no barcode.** A size row emitted from a
  configurable tile (sku::size) has gtin=None, so (a) image enrichment
  (keyed on barcode) can never attach a photo, and (b) cross-shop matching
  falls back to name matching, which is what splits (P2 above). Possible fix:
  pull per-variant GTINs from the tile's spConfig if present.

- **Ghost products at hidden stores.** The 8/25 purge removed superseded
  plain-sku listings only where a per-size sibling exists at the same store.
  Stores whose tiles could not be re-read keep old listings (deliberate), so
  a few barcode-less ghosts survive at hidden locations (e.g. "1942 Tequila
  70cl"). Self-heals when those stores next crawl clean; re-run the purge
  pattern if not.

- **Reverse ghost edge (not yet seen).** If a tile ever goes multi-size ->
  single-size, its ::size listings would linger the same way the plain ones
  did. Same purge pattern applies; watch for it once crawls are scheduled.

- **Wine-competition silver/bronze medal artwork missing** (aiwc/biwc/miwc/
  nyiwc). Cards fall back to the "Award winner" text tag — works, but
  sourcing the artwork would finish the set.

## Open — build

- **Coverage claims carry a verify date and are re-checked before client conversations;
  sites change under us.** The 8/31 audit (all fetched fresh) stands: DOH legacy site with no
  catalogue; IST a shop directory with no product prices; the Heinemann family robots
  `/*/search/` still present; Lotte robots `*: Disallow /`; DFA historic robots-no kept;
  AMS has a real catalogue on schiphol.nl but 403s every automated fetch including a
  browser UA (fingerprinting), so anti-bot is a technical no and the lane is partnership.
  Corrected that day: CDG Extime openable (now collected); ICN and SIN as recorded under P1.
  The posture-per-source table lives in `main/docs/COLLECTORS.md` once Stream A fills it.

- **Award aging / re-entry incentive** (rian, 8/27, from his board comments):
  demote or hide awards older than a year on listings to encourage brands to
  re-enter the competitions each season. Business-model-relevant; belongs in
  the deeper awards-sync rebuild (Stream F).
- **Board tier-set nuance** (rian, 8/27): for part-built rows he wanted both a
  timing answer AND an effort answer ("a bit before launch, but with a heavy
  option"). Card copy carries the nuance for now; revisit if the board gets
  reused for phase 2.

- **Scheduler + freshness alarm.** The big one before Cannes: cron the
  collections, and alert when a store returns far less than its last run
  (the Heathrow lesson). Prices must be current in late September.
  *Owner:* Stream E (host cron, nightly audit, verify after any collect) with Q's
  `verify` as the tripwire (Decision 10).

## Parked — business/decisions (rian's court)

- Hosting move to cloud: question brief in `notes/hosting-brief.md`; the crawler
  moves with the app; no IP-rotation services ever (conduct posture). Now build plan
  §10 #5 (figure and message to Adam) and #7 (where collectors run, after E's egress test).
- Revenue model: two-layer (customers pay Adam, Adam pays for the engine);
  IP license-back clause for the next SOW (`notes/revenue-thinking.md`).

## Parked — for the back office

- **AI-assisted QA (rian, 2026-09-04).** Local small model proposes, humans confirm in batch:
  brand-merge suggestions before publish; draft airport/brand descriptions; "likely
  inconsistent, verify" lists. Needs the candidates queue + `decided_by` (plan §4). Not before
  the back office exists; Claude-in-session covers it until then.
- **Per-source identity mode with consent record (rian, 2026-09-04).** See plan §3b design
  note. Needs `sources.identity_mode` + `permission_record` (migration #1 carries the columns);
  owner-area toggle; collector guard.
- **Human data-entry fallback + likely-stale mechanism (rian, 2026-09-04).** Back office
  surfaces high-value comparisons automation could not refresh; human enters price + URL + date
  as an observation with `decided_by=human`; stale rows flagged, demoted from featured, dated.
  Needs `price_observations.source_kind` (collector|human|api) from the start (migration #1).
- **Merge review queue + recorded merges (back office).** `merge_candidates`, `product_merges`,
  old-id redirects in seo.py. AI proposes, humans confirm in batch.
- **Verified-field semantics (back office).** Decision 7: `overrides` is the single record of
  human decision, with `collector_disagrees_since`; no per-field verified columns.
- **Reverification queue with scoring and human-fed re-ranking (rian, 2026-09-04).**
  `reverifications` table; meaningfulness scorer (rules first); humans clear a batch; re-rank
  the rest from their verdicts (active-learning shape). Belongs with the back office; the table
  and scorer fields exist from the identity migration so nothing is thrown away.

## Resolved — kept for the reasoning

- **The 2026-09-04 audit defect list (landed 2026-09-04, Stream A).** Barcode/size
  contradictions are rejected rather than re-routed (two 7,000 ml Glenfiddichs); `observed_at`
  per fetch and the FX rate per observation; eight stuck runs and MEX's MXN row fixed by
  backfill and prevented at the source; Dubai stock unknown unless stated; image attribution
  says barcode or name and name matches must agree on ages; 25 orphan parent tiles purged.
  Each is a test in `tests/test_audit_defects.py`; mechanism in `main/docs/COLLECTORS.md`.
- **Raw-record retention (landed 2026-09-04, Stream A).** `raw_records` JSONB per listing per
  run + `parser_version`, migration #1; until it existed every identity-rule change cost a
  recrawl. `app.cli rederive` follows.
- **`avolta.py` used the stdlib robots parser** (resolved 2026-09-04, Stream A). It predates
  wildcards, so Heathrow's `Disallow: /*?` read as permission and we fetched `?p=` pagination
  the host refuses. One policy module now (`collectors/robots.py`), matched on the bot's own
  name, applied by every collector on every run; the misread is a test.
- **Rendering is in-house, not paid** (decided 2026-09-04, Decision 5; framed 8/31). SIN-class
  sites (JS-only, crawl-welcoming) are read with our own headless Chromium in an isolated
  sidecar; no paid fetcher, no IP rotation, no anti-bot evasion (which also rules out
  "stealth" browsing AMS past its 403s). Supersedes the 8/31 SIN scouting note.
- **Adam has the review surfaces.** He approved and submitted the quote on `/quote` on
  2026-09-03; the walkthrough video sits on `/discuss`.
- **Invented barcodes (fixed 2026-09-03).** `clean_gtin` padded any 6-to-12 digit
  number to 13 and check-digit tested it, so a bare SKU passed about one time in
  ten; `avolta.py` and `shopify.py` both fed it SKUs. 119 stored barcodes were
  padded short numbers. A blanket ban was wrong: 149 of Avolta's 489 barcodes
  agree exactly with Dubai/Heinemann, so 30% of those SKUs are real EANs. The
  separator is width. `gtin_from_sku()` accepts a supplier code only if it is
  ALREADY 8/12/13/14 digits and never pads one up: rejects 61 codes of which 1
  was corroborated, keeps 148 of the 149 that were. 88 stored fakes nulled; the
  ambiguous 11-significant-digit group left alone (10 of 75 corroborated).
- Review page rebuilt to rian's pitch with the demo picker + positioning sweep
  (8/26, v0.24.0): issues internal, page carries features + decisions only.
- Armand de Brignac cross-size comparison (8/25): dropped-decimal slug parse
  + per-store variant defaults. Fixed structurally (per-variant emission,
  parser rule, purge migration); families verified size-pure post-recrawl.
- Superseded-listing ghosts from the per-variant switch (165 purged, 8/25).
- Medal artwork white tiles + Cloudflare cache (8/25): flood-filled to
  transparency; medal URLs versioned (MEDAL_ART_VERSION — bump on artwork
  change).
- 19 airports in the picker (8/25): trip service missed the visibility
  filter; all query paths now covered.
