Fable 5 — Pilot Output & Scale-Readiness Review (2026-07-11)
# Fable 5 — Pilot Output & Scale-Readiness Review (2026-07-11)
**Reviewer:** Fable 5 (adversarial review round 4 for this project: v1 plan → 12 blocking findings; v2 → final check; family genomes 7-14 → targeted fixes; this round → the actual pilot output + the workflow that would carry a scale-up).
**Scope reviewed:** all 15 JSONL files in `/opt/oimy-engine/docs/synthetic-user-testing/gold-standard-output/` (547 data rows), `review-feedback.jsonl`, `gold-review-server.js` + service, `/root/diag_gate/` (the §2 diagnostic-gate artifacts), the v4 plan + addenda v4.1-v4.4 (git `09c8f6f`), the architecture briefing (Hetzner copy, 2026-07-11-updated), and the overnight session log (`58b9f1e`).
**Method:** every mechanical claim below was verified by scripted checks run on the server against all 547 rows (scripts left at `/root/gold_audit*.py`), plus manual reads of ~60 gold responses across all 5 pilot families and 3 overnight families, plus a full manual classification of all 49 diagnostic-gate generations.
---
## 0. Verdict, in three lines
- **(a) Content quality of the 5 pilot families:** GOOD ENOUGH TO BE THE QUALITY BAR, with three specific fixes (F-1, S-1, S-2 below). The gold responses are genuinely differentiated, the history format is production-correct, and the register work in the transliterated families is real. This is not the problem.
- **(b) Workflow/infrastructure readiness for scale:** NOT READY. None of the plan's quality gates exist as runnable infrastructure (§7.3 mechanical QC: not implemented; §9 cultural rubric: never ran; review tooling covers 5 of 15 files; human review coverage is 3 of 547 rows). The overnight non-pilot batch already shows the drift these gates exist to catch.
- **(c) "Hundreds of families by July 13":** NO — not with this method, and not with any method that preserves the quality bar. What IS realistic by July 13, and a real path to hundreds after, is in §5.
---
## 1. The §2 diagnostic gate — the missing half, completed here
**Correcting the record first:** an earlier draft of this review's brief said the gate "never ran." That is false. The expensive half ran and ran well: `/root/diag_gate_assemble_run.py` (Jul 10 16:47) built 49 regime-3 (day-28-42) production-faithful contexts across 3 families (Whitaker, Okonkwo, Delgado) — within the plan's 40-60 target — and collected real generations from stock `gemma4:12b-it-qat` (Jul 10 16:52, all 49 error-free). What never happened was the cheap half: no §2.3 judge classification and no §2.4 stop/proceed verdict exists anywhere on the box, in git, or in the plan doc. The pilot proceeded without the gate's decision ever being computed.
**I have now computed it, for real, by reading all 49 (context, generation) pairs myself** against §2.3's binary rubric (behavior-shaped = ignores/misuses available context, generic despite rich info, personalization degrades while fluent; capacity-shaped = raw coherence degrades with length — repetition, drift, syntactic breakdown, attention collapse).
| Bucket | Count | % of 49 |
|---|---|---|
| No meaningful failure (usable or near-usable response) | 21 | 43% |
| Behavior-shaped failure | 16 | 33% |
| Minor behavior-class imperfection (canned phrasing, mild stale echo, over-length) | 12 | 24% |
| **Capacity-shaped failure** | **0** | **0%** |
Of the 28 rows with any failure: **100% behavior-shaped, 0% capacity-shaped.** Against §2.4's thresholds (≥70% behavior-shaped → proceed; ≥30% capacity-shaped → stop): **PROCEED, decisively.** Coherence held perfectly at history depth 16 and ~30-37K-char system prompts — no repetition, no syntactic breakdown, no length-driven collapse anywhere in the set.
The three dominant behavior failures, all squarely in the dataset's fixable target class:
1. **Stale-turn re-answering** (~13 clear instances): the model re-answers a *previous* user turn (sometimes two turns back) before or instead of the current one — e.g. `diag_whitaker_35_0` responds about the anniversary when the user is talking about the truck; `diag_delgado_39_1` ignores the user's sign-off entirely and re-answers the prior turn. Note: these contexts were built with **user-only history** (correct at build time; reversed by the same-night session-boundary fix, plan addendum v4.4) — this failure mode is exactly what the production fix addressed, so a share of it should disappear at inference now. It is also exactly what alternating-history training rows teach against.
2. **Register mismatch:** Whitaker's terse profile gets 5-8-line empathy-first therapist prose anyway ("That feeling ... is heavy" opens 9 of 17 Whitaker generations).
3. **Canned BCT surface phrases:** "You just gave your brain new evidence: that is what someone who X does" recited near-verbatim across families — technique visible instead of internalized.
**Caveats on this verdict, stated honestly:** (i) contexts are user-only-history shaped, which no longer matches production — the capacity conclusion (0% capacity-shaped at full length) stands regardless, but a small re-run (n≈20) with alternating history is cheap and worth doing before the full-scale training commit; (ii) §2.1 asked for "Whitaker + one Indian family + one other" — no Indian family was included, so nothing about long-context transliterated input was probed; (iii) this classification was done by one reviewer (me), not the plan's recommended Opus judge — the calls are conservative and the margin (0/49 capacity) is not close, but the classification file deserves a second pair of eyes if anyone doubts it.
---
## 2. MUST-FIX before any further generation or any training run
**F-1. Family-11's tone fix corrupted its own histories — 24 of 24 rows now internally contradictory.**
The 03:20 tone-fix (correctly responding to Bharath's review) rewrote all 25 `target`s from analytical English into colloquial Tanglish — but every row's `history` still embeds the **old English assistant turns** (verified: 24/24 rows' last-assistant history entry matches the superseded `.bak` target, 0/24 match the new one). As shipped, every family-11 row after day 1 trains: "Oi has been answering this user in long analytical English for weeks; now answer in casual Tanglish." That teaches voice inconsistency inside the exact multi-day-coherence dataset built to teach the opposite. Fix: rebuild each row's history from the revised targets (scriptable in minutes since history is derived data). **General lesson: any future target revision must trigger downstream-history regeneration — this must be a script, not a memory.**
**F-2. The §7.3 mechanical QC gate does not exist as code, has never run, and 20 minutes of scripted checking found real violations it would have caught.**
No gate script exists anywhere on the box (confirmed by filesystem search; the only rubric file, `judge-rubric-v1.md`, is the loop-phase0 eval rubric, not this). Found by my scripted pass over all 547 rows:
- **Markdown/formatting violations** (a plan hard rule) in at least 8 targets, all in the overnight non-pilot batch: bulleted/numbered-list coaching scripts in `md_reeves_grandparent_10_0`/`_33_0`, `md_okonkwo_15_0`/`_25_0`, `md_delgado_13_0`/`_18_0`, `md_castellano_reyes_14_0`, `md_okafor_bianchi_34_0`. Zero such violations in the 5 pilot files — the drift is specific to the batch that ran without review.
- **Plan-mandated sidecar fields absent from all 547 rows:** `difficulty_tier` (§8 — the 40/40/20 ratio was "resolved" per the session log but is recorded nowhere per-row, so it is unauditable), `writer_rotation_cell` (§7.5 — moot-but-undocumented since a single author wrote everything, itself a deviation from §7.5/§7.6), `cultural_authenticity_qc_result` (§9 — absent because the pass never ran, see F-3), `signal_availability_state`.
- **Coverage-matrix ceiling (§3.3) near-misses in the full-arc files:** untagged-day share is 45% (family 5) and 48% (family 4) against the ≥50-60% floor — borderline, fixable by tag audit rather than regeneration. The other nine full arcs comply (50-60%).
- The good news inside this finding: clinical-jargon, framework-citation, and follow-up-promise scans came back **zero violations in all 547 targets**, and the signal-availability rule is honored everywhere (see §4). The gates are cheap to build — my audit scripts are half of one already — and they demonstrably catch real problems. There is no excuse to generate one more row before they exist and run.
**F-3. The §9 cultural-authenticity QC has never run in any form.**
No rubric pass (the mechanical/scalable first layer, hard-fail on caricature, regenerate-once-then-discard) has touched families 10, 11, or 12. No external Telugu review is arranged. The **only** quality check any Indian-family content has ever received is Bharath's manual review of 3 rows of family 11 — which immediately found real problems (wrong language register, over-analysis, bookish vocabulary), i.e., the layer that hasn't run is precisely the layer that catches things. Additionally, the family-11 fix itself now needs Bharath's re-review: 25 Tanglish targets machine-rewritten in response to feedback about machine-written Tamil sounding bookish must not be assumed fixed without native eyes on at least a sample.
**F-4. Both of the plan's execution gates were overrun by the overnight scale-up.**
Per plan §15.1, nothing downstream starts before §2's verdict — which was never computed (now supplied, §1 above, and it happens to say proceed; that is luck, not process). Per §1.5, the pilot exists to be reviewed *before* the remaining families generate — instead, 10 additional full 42-day families (1, 2, 3, 5, 6, 7, 8, 15, 16, 17 — 420 rows, 77% of the current corpus) were generated 05:04-05:15, while human review stood at 3 rows of one family. The F-2 formatting violations cluster exactly in that ungated batch. Nothing here is unrecoverable — the batch is decent on spot-read and can be gated retroactively — but the pattern (gates defined with care, then skipped under time pressure) is the single biggest process risk to a 10x-50x scale-up, and it already fired at 3x.
**F-5. The training rows' `system` field is ~10x smaller than production's real system prompt — this must become an explicit, documented decision before training, or it is a train/serve mismatch.**
Gold rows carry bespoke 1.5-3.9K-char per-family system prompts. Production's real assembled `system_content` (verified directly from the diagnostic-gate contexts, which replicate production assembly) is **30,837-37,304 chars** — the combined-prompt + memory + anchors + threads stack. The plan's own foundational rule (§6, opening line) is that context shape must match inference-time shape. Two readings: (i) if the fine-tuned Gemma is to be served behind production's full system assembly, the corpus as built teaches against the wrong system shape; (ii) if the LoRA serving path will use compact systems (consistent with `train_v31.jsonl`'s ~1,798-token/row mean — the prior corpus plainly did not carry 30K systems), it's fine — but that is exactly plan §14.3 item 4 ("which serving path carries real Gemma-4-LoRA traffic"), still open, and the compact-system choice is documented nowhere (the briefing's checklist covers `history` shape in detail and never mentions `system`). Resolve and write it down before `train_v4` assembles; do not let this be discovered after a training run.
---
## 3. SHOULD-FIX before or during roster completion
**S-1. Bharath's family-11 register feedback is a partially systemic pattern — audit families 10 and 12 against it before scaling.** The worst component (targets in English at all) was family-11-specific: families 10 and 12 were authored in transliterated Hindi/Telugu from the start (verified by direct read). But the second component — the analytic-lecture register ("Pehle ek baat seedhi: ... Inhe alag rakhna. Ab aapne ek kaam maanga, toh ek hi deta hoon" — family 10, day 9) versus the "good friend who lightens the mood" voice Bharath described — appears in family 10's heavier days and, more mildly, family 12. These need his sampled read with the family-11 notes in hand, not silent inheritance of the old register.
**S-2. Multi-turn days are missing almost entirely.** Plan: ~1.3 avg turns/day (~25-30% two-turn days) → ~928 raw rows across 17 families. Actual: 1.0 turns/day everywhere except family 12 (2 two-turn days in 547 rows). Cost: ~130 missing rows and, more importantly, the entire "short second-turn follow-up" behavior class — a real production shape — is untrained.
**S-3. The roster is incomplete in ways that matter more than more-families would:** family 13 (the held-out eval family) does not exist — **without it there is no `eval_v4_holdout` and no §10 generalization measurement at all**, making any training run unevaluable by the plan's own design; family 14 (Tamil-US ordinary) is absent; family 12 has only days 1-10 (by design so far, but the remaining 32 days are gated on a Telugu review that hasn't been arranged); families 9/10/11 are tagged-days-first slices (76-83% of their generated days are tagged) — legitimate as front-loading, but **untrainable as-is** without the ordinary-day backfill or they blow the §3.3 ceiling from the other direction.
**S-4. Review tooling covers a third of the corpus and has no scale path.** `gold-review-server.js` hardcodes the 5 pilot families; the 10 overnight files are invisible to the only review mechanism that exists. There is no sampling mode, no regenerate-from-feedback loop, no reviewed/unreviewed bookkeeping beyond the feedback JSONL. (A stratified-sampling review plan reportedly exists from a prior discussion — it is not documented anywhere I searched: not in git, not in the plan, not on the box. Pending Bharath's confirmation; it needs to be written down and implemented either way.)
**S-5. Diagnostic-gate residue:** re-run a small alternating-history round (n≈20, include one Indian family) before the training commit; land the classification (§1) as a durable artifact next to the generations file.
**S-6. `gold-standard-output/` is not under version control.** The family-11 tone fix survives only as a `.bak` convention; F-1 was found by diffing that one lucky backup. Same standing ask as session-log §6 — a repo (or the planned Hetzner repo) should cover the dataset dir before any regeneration wave rewrites files in place.
---
## 4. LOOKS GOOD — verified, not assumed
- **History format is production-correct across all 547 rows** (the review's #1 pre-registered worry): alternating `{role, content}` dicts, user-first, assistant-last, zero non-alternating sequences, zero inverted Q/A pairs found; condensation rule implemented for real (8 verbatim / 200-char+"…" truncation observed on 5,000+ older entries / hard 24-message cap hit repeatedly, never exceeded); day-N history correctly embeds day-N-1's actual turns, including correct interleaving of skeleton days that were never emitted as training rows (what looked like continuity breaks in family 10 is correct intervening-day content).
- **Pilot composition matches §1.5 exactly** (4, 9, 10, 11 + family 12 days 1-10).
- **Signal-availability discipline is right everywhere it exists:** families 9 and 12 carry the `[VOICE MOOD SIGNAL — SPECULATIVE FORMAT, NOT PRODUCTION-VERIFIED]` block with `speculative_format: true` on every row (25/25, 12/12); all other families' profiles explicitly say no signal available; zero targets reference a mood signal without one in context.
- **Zero clinical jargon, zero framework citations, zero check-in/follow-up promises** in any of 547 targets (scanned).
- **`content_type: "transliterated_multilingual"` present on all 60 Indian-family rows.**
- **The content is genuinely differentiated where I read it closely.** Whitaker's targets hold the terse register with discipline (day 1's "Set those three first and see if it earns the space on your phone. If it doesn't, you're right to ditch it" is exactly the anti-generic bar the plan describes); untagged ordinary days spot-read as genuinely ordinary (inverse §3.3 check holds on samples); family 12's Telugu rows carry the plan-required honest-uncertainty annotations for the native reviewer.
- **The feedback loop works when exercised:** review UI → Bharath's 3 notes at 02:37-02:53 → all 25 family-11 targets rewritten by 03:20. Fast and real (its one bug is F-1).
- **The diagnostic gate's generation half was built well** — production-faithful 30K+ system assembly, ramped history depth, clean 49/49 run.
---
## 5. The real question: "hundreds of families by July 13"
**Direct answer: no. Not with the current method, and not honestly with any method in 2 days.** Not softened, and for four independent reasons — any one of which is sufficient:
1. **Authoring is founder-judgment-bound, not compute-bound.** The 17 families exist because Bharath personally named each gap (families 14-17 were each added by a specific founder instruction). Bespoke genomes at that rigor run a handful per day *with* the founder in the loop. "Hundreds" via this path is weeks-months, not days — and the judgment step is precisely what makes the current content good.
2. **Review bandwidth is the hard wall.** Bharath's measured pace: 3 rows / 16 min. The current 547 rows ≈ 48 hours of line-by-line review. 300 families ≈ 12,600 rows ≈ 1,100+ hours. Even an aggressive 5% stratified sample is ~630 rows ≈ 56 hours. Without automated QC running first (which doesn't exist yet, F-2) plus a written sampling design (S-4, pending), "hundreds of families" means either shipping ungated rows into the training corpus or an impossible review queue. There is no third option.
3. **The plan's own scaling logic forbids it.** §1.2 (unchanged through every version, re-verified by two Fable rounds): generate ~655-874, train, measure target-token-share effects, and only then consider Fable's suggested ~2,500 — *with eval evidence in hand*. No training run has yet shown that even 547 of these rows help. Hundreds of families (≈12,600+ rows, several times the current corpus's target-token share) is a bet placed before the first result is known, against the explicit design.
4. **The infrastructure drift already visible at 3x** (F-2's violations clustering in the one ungated batch, F-1's silent history corruption) says exactly what would happen at 20x with no gates.
**What IS realistic by July 13** (in order; 1-2 focused days, mostly automatable):
1. Build and run the §7.3 mechanical QC suite over all 547 rows (half exists in `/root/gold_audit*.py`); fix the 8 formatting violations; backfill the missing metadata fields.
2. Fix F-1 (regenerate family-11 histories from revised targets — scriptable).
3. Run the §9 cultural rubric pass over families 10/11/12; get Bharath's sampled re-review of post-fix family 11 and a family-10/12 register read (S-1); arrange the external Telugu check.
4. Complete the roster to the plan's own 17: family 13 (unblocks the entire eval design), family 14, family 12 days 11-42, ordinary-day backfill for 9/10/11, and the two-turn days (S-2). End state: the full ~928-row corpus, gated, reviewable, evaluable.
5. Resolve F-5 (system-field decision) and re-run the small alternating-history diag round (S-5).
6. Extend the review server to all families + implement the stratified-sampling review flow once Bharath's intended design is confirmed.
**The honest path to hundreds, after July 13, if still wanted:** the Phase-0 seed-extrapolation idea, which the genome schema was explicitly built to support — the `vary-along axes` and `reject-lists` in every genome are extrapolation controls that nothing has ever consumed programmatically. Requirements before it can run without quality collapse: (a) a variant-generator harness (seed genome + axis vector → variant genome + arc skeleton) with the 17 hand-authored families as seeds and calibration data; (b) the F-2 mechanical QC suite as a hard gate on every generated row; (c) the §9 rubric pass hard-gating all Indian-language variants; (d) a dedup/diversity check across variants of the same seed (near-duplicate arcs from one seed are oversampling wearing a costume, not coverage); (e) stratified human sampling with written per-cell coverage floors; (f) **the §1.2 evidence gate**: train on the 17-family corpus first, measure on `eval_v4_holdout` (family 13 — which is why S-3 matters), and let those results size the expansion. Realistic build time for (a)-(e): 3-5 focused days. Then hundreds of variant families in 1-2 weeks, gated — not 2 days, but real, and the only version of "hundreds" that wouldn't quietly destroy the quality bar the pilot just set.
---
## 6. Evidence appendix (verify me)
- Mechanical audits: `/root/gold_audit.py` (schema/roles/tags/jargon/markdown/promises/signal scan, all files), `/root/gold_audit2.py` (continuity + flagged rows), `/root/gold_audit3.py` (F-1 staleness proof: 24/24 stale, 0/24 fresh), `/root/gold_audit4.py` (intervening-day continuity + untagged-day sample), `/root/dump_diag.py 0 49` (the 49 diag pairs as I read them).
- F-1: compare `family-11-tamil-us-teen.jsonl` rows' `history[-1]` against `.bak-pre-tone-fix-2026-07-11` targets.
- F-5: `system_chars` field in `/root/diag_gate/diag_gate_generations_gemma4_12b-it-qat.jsonl` (30,837-37,304) vs `len(system)` in any gold row (1,501-3,862).
- Overnight-batch provenance: file mtimes 05:04-05:15 vs `review-feedback.jsonl` timestamps 02:37-02:53 vs session-log §4.3 ("remaining 11 families ... not yet started" as of the log).
- Row inventory: 11 full-arc files × 42 + 25 + 23 + 25 + 12 = 547 rows; missing: families 13, 14; family 12 days 11-42.