Insights F01: quoted but not booked

The 98-question target is NOT ESTABLISHED; overall family acceptance is FAIL. Only five original source IDs have direct corrected-oracle mappings. The measured real-family sample has eight attempts, including an authored three-turn conversation; it cannot establish correctness for all 98 source questions.

Observed correct completions in repeats 1 / 2 / 3: fresh candidate 0 / 0 / 0 versus control 1 / 3 / 3 per 122 attempts; real-sample candidate 2 / 1 / 2 versus control 1 / 1 / 2 per 101 attempts. Within the eight-attempt real-family sample, candidate 1 / 0 / 1 versus control 0 / 0 / 1. Coverage limits and measured outcomes are separate findings; neither comparison-unverified answers nor missing references count as correct.

For the five directly mapped original source questions, control correct counts are 0 / 0 / 1 and candidate counts are 1 / 0 / 0. Their pooled correct count does not increase. The eight-attempt sample’s additional candidate success comes from its authored conversation; it is not an additional correct original source question.

The application change gives the current generator a mandatory, question-conditioned family definition: client-requested quotes, client identity, booking existence, the user anti-join and the ratified booking window. Older conflicting guidance is reconciled. A fixed refusal is implemented for recognized unsupported variations. Measured refusal coverage is incomplete: one outside-envelope case remained comparison-unverified in all three candidate repeats; family obligations remain explicitly unproved. Verifier selection receives the same definition; the measured verifier is off. No generator/model selection or verifier repair change was made.

Goal F01 maps to source family F10. Source F01 is acquisition-source signups. All correctness labels below mean correctness under the unchanged corrected D2 oracle. The correspondence of every reference to every family variation has not been independently established.

Populations and method

Three repeats per arm/population; 122 fresh attempts (31 families), 101 real-sample attempts (45 families). D3 Arm A supplies three historical fresh controls and real repeat 1; base-tip real repeats 2–3 are new. New runs use two workers, frozen clock 2026-09-06 12:00 UTC, London periods, zai requested glm-5.2, temperature 0, seed 42, thinking disabled and a 25-second provider timeout. Provider response identity is retained separately. Historical D3 measured 2513a718; its final-tip metadata/telemetry corrections are qualified by D3’s 934-attempt non-interference replay, not a live latency equivalence test. New controls and candidates share the F01-isolated harness. Historical controls lack flock evidence. Sequential host/provider conditions remain confounded.

Every new repeat and restart acquires /tmp/lore-eval.lock; wait and host observations are retained. Primary latency is nearest-rank p95 of feature_completion_ms (submission through terminal-result retrieval), measured on a shared host. Timed failures are included; missing timings are excluded and counted. Precision is correct/(correct+silent wrong); comparison-unverified is excluded and undefined precision stays n/a.

First-turn error-rate denominators are supported first turns: fresh 116, real 86, fresh target 5, real target 6, fresh excluding target 111, real excluding target 80. Whole fresh contains three conversations; whole real contains five.

The fresh slice was previously exposed in D2. The recorded design process used question shapes and ratified definitions; candidate outputs were not used to tune application code. Application and harness were frozen before measurement. These records support consistency, but cannot independently prove non-exposure or the absence of influence from previously observed answers. Full-98 missing oracle/context observations are coverage gaps, not incorrect answers.

The previous session lost its controller during candidate fresh repeat 1. Its 122 submitted attempts survive, but its terminal summary, original exit status and original cleanup receipt are missing. Recovery verified the orphan server identity, stopped it under the shared lock and checked unchanged frozen inputs. Recovery does not establish uninterrupted lock ownership through the original shutdown. The run remains incomplete and excluded from primary metrics; its compact attempts and a clearly labelled recovery receipt are retained under aborted/. The single allowed clean restart supplies scored repeat 1. The earlier zero-question control admission is retained separately. No answer quality selected either restart.

Whole fresh slice

Mean [minimum, maximum] from three repeats. Undefined values are excluded from means and remain n/a in the per-repeat table; the number of defined values is retained in results.json.

Metric Control Candidate
First-turn execution errors 28.00 [27.00, 30.00] 25.67 [24.00, 27.00]
Correct completed 2.33 [1.00, 3.00] 0.00 [0.00, 0.00]
Silent wrong 30.00 [27.00, 34.00] 26.33 [24.00, 30.00]
First-turn error rate 24.14% [23.28%, 25.86%] 22.13% [20.69%, 23.28%]
Precision among scored answers 7.15% [3.33%, 10.00%] 0.00% [0.00%, 0.00%]
Refusals 39.00 [38.00, 41.00] 42.00 [40.00, 45.00]
Guard refusals 39.00 [38.00, 41.00] 42.00 [40.00, 45.00]
Verifier refusals 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
Comparison unverified 21.33 [19.00, 23.00] 26.33 [22.00, 31.00]
Fully correct conversations 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
p95 application latency, shared host (s) 30.42 [27.29, 34.54] 28.14 [25.14, 32.17]
Task errors 29.33 [28.00, 32.00] 27.33 [26.00, 29.00]
Infrastructure failures 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
Blocked by parent 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
Arm / repeat First-turn errors Correct Silent wrong Refused Unverified comparison Precision p95 shared-host seconds Timed/total Refusal guards
base-fresh-r1 30 1 29 38 22 3.33% 29.43 122/122 {"CalendarSemanticsError": 33, "NativeDealUnitError": 4, "UNKNOWN-GOLD-COLUMN": 1}
base-fresh-r2 27 3 27 41 23 10.00% 34.54 122/122 {"AdditiveGrainError": 1, "CalendarSemanticsError": 39, "NativeDealUnitError": 1}
base-fresh-r3 27 3 34 38 19 8.11% 27.29 122/122 {"CalendarSemanticsError": 36, "NativeDealUnitError": 2}
cand-fresh-r1 26 0 24 40 31 0.00% 25.14 122/122 {"CalendarSemanticsError": 35, "LossReasonSourceError": 1, "NativeDealUnitError": 4}
cand-fresh-r2 24 0 25 45 26 0.00% 32.17 122/122 {"CalendarSemanticsError": 42, "NativeDealUnitError": 3}
cand-fresh-r3 27 0 30 41 22 0.00% 27.12 122/122 {"CalendarSemanticsError": 37, "NativeDealUnitError": 4}
Candidate − control Difference [95% family interval] Control spread Candidate spread
Precision among scored answers -7.22% [-14.77%, -2.04%] 6.67% 0.00%
Comparison unverified 5.00 [-1.37, 11.75] 4.00 9.00
Correct completed -2.33 [-4.23, -0.68] 2.00 0.00
First-turn execution errors -2.33 [-7.06, 2.74] 3.00 3.00
Fully correct conversations 0.00 [0.00, 0.00] 0.00 0.00
p95 application latency, shared host (s) -2.87 [-7.83, 1.03] 7.25 7.03
Refusals 3.00 [-4.55, 10.49] 3.00 5.00
Guard refusals 3.00 [-4.55, 10.49] 3.00 5.00
Verifier refusals 0.00 [0.00, 0.00] 0.00 0.00
Silent wrong -3.67 [-7.81, 0.33] 7.00 6.00

The mean correct-completion difference is -2.33 per repeat, beyond the control’s observed range of 2.00. The range is an observed noise floor, not a confidence bound.

Whole real sample

Mean [minimum, maximum] from three repeats. Undefined values are excluded from means and remain n/a in the per-repeat table; the number of defined values is retained in results.json.

Metric Control Candidate
First-turn execution errors 11.00 [10.00, 13.00] 11.67 [10.00, 14.00]
Correct completed 1.33 [1.00, 2.00] 1.67 [1.00, 2.00]
Silent wrong 23.00 [19.00, 28.00] 27.00 [25.00, 29.00]
First-turn error rate 12.79% [11.63%, 15.12%] 13.57% [11.63%, 16.28%]
Precision among scored answers 5.34% [4.35%, 6.67%] 5.81% [3.57%, 7.41%]
Refusals 31.67 [27.00, 36.00] 26.67 [25.00, 30.00]
Guard refusals 31.67 [27.00, 36.00] 26.67 [25.00, 30.00]
Verifier refusals 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
Comparison unverified 27.33 [25.00, 30.00] 27.67 [26.00, 30.00]
Fully correct conversations 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
p95 application latency, shared host (s) 26.55 [22.16, 31.20] 31.33 [28.77, 35.62]
Task errors 17.67 [17.00, 19.00] 18.00 [17.00, 19.00]
Infrastructure failures 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
Blocked by parent 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
Arm / repeat First-turn errors Correct Silent wrong Refused Unverified comparison Precision p95 shared-host seconds Timed/total Refusal guards
base-real-r1 10 1 22 36 25 4.35% 31.20 101/101 {"-": 1, "CalendarSemanticsError": 32, "NativeDealUnitError": 2, "WindowBindingError": 1}
base-real-r2 13 1 19 32 30 5.00% 22.16 101/101 {"-": 1, "AdditiveGrainError": 1, "CalendarSemanticsError": 29, "NativeDealUnitError": 1}
base-real-r3 10 2 28 27 27 6.67% 26.30 101/101 {"-": 1, "CalendarSemanticsError": 24, "NativeDealUnitError": 2}
cand-real-r1 14 2 29 25 26 6.45% 35.62 101/101 {"-": 1, "CalendarSemanticsError": 23, "NativeDealUnitError": 1}
cand-real-r2 10 1 27 25 30 3.57% 28.77 101/101 {"-": 1, "CalendarSemanticsError": 20, "FanOutSumError": 1, "NativeDealUnitError": 2, "WindowBindingError": 1}
cand-real-r3 11 2 25 30 27 7.41% 29.59 101/101 {"-": 1, "CalendarSemanticsError": 28, "NativeDealUnitError": 1}
Candidate − control Difference [95% family interval] Control spread Candidate spread
Precision among scored answers 0.33% [-2.15%, 3.04%] 2.32% 3.84%
Comparison unverified 0.33 [-3.27, 4.11] 5.00 4.00
Correct completed 0.33 [0.00, 0.99] 1.00 1.00
First-turn execution errors 0.67 [-2.36, 3.44] 3.00 4.00
Fully correct conversations 0.00 [0.00, 0.00] 0.00 0.00
p95 application latency, shared host (s) 2.45 [-0.76, 7.91] 9.05 6.85
Refusals -5.00 [-8.88, -1.31] 9.00 5.00
Guard refusals -5.00 [-8.88, -1.31] 9.00 5.00
Verifier refusals 0.00 [0.00, 0.00] 0.00 0.00
Silent wrong 4.00 [0.93, 7.48] 9.00 4.00

The mean correct-completion difference is 0.33 per repeat, inside or equal to the control’s observed range of 1.00. The range is an observed noise floor, not a confidence bound.

Fresh target sample (seven attempts)

Mean [minimum, maximum] from three repeats. Undefined values are excluded from means and remain n/a in the per-repeat table; the number of defined values is retained in results.json.

Metric Control Candidate
First-turn execution errors 1.00 [1.00, 1.00] 0.00 [0.00, 0.00]
Correct completed 0.33 [0.00, 1.00] 0.00 [0.00, 0.00]
Silent wrong 0.33 [0.00, 1.00] 0.33 [0.00, 1.00]
First-turn error rate 20.00% [20.00%, 20.00%] 0.00% [0.00%, 0.00%]
Precision among scored answers 50.00% [0.00%, 100.00%] 0.00% [0.00%, 0.00%]
Refusals 2.00 [1.00, 3.00] 0.67 [0.00, 1.00]
Guard refusals 2.00 [1.00, 3.00] 0.67 [0.00, 1.00]
Verifier refusals 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
Comparison unverified 2.33 [1.00, 3.00] 5.00 [5.00, 5.00]
Fully correct conversations 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
p95 application latency, shared host (s) 23.49 [21.01, 26.02] 20.86 [14.71, 28.61]
Task errors 2.00 [1.00, 3.00] 1.00 [1.00, 1.00]
Infrastructure failures 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
Blocked by parent 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
Arm / repeat First-turn errors Correct Silent wrong Refused Unverified comparison Precision p95 shared-host seconds Timed/total Refusal guards
base-fresh-r1 1 0 0 1 3 n/a 23.43 7/7 {"CalendarSemanticsError": 1}
base-fresh-r2 1 0 1 2 3 0.00% 21.01 7/7 {"CalendarSemanticsError": 2}
base-fresh-r3 1 1 0 3 1 100.00% 26.02 7/7 {"CalendarSemanticsError": 3}
cand-fresh-r1 0 0 0 1 5 n/a 14.71 7/7 {"CalendarSemanticsError": 1}
cand-fresh-r2 0 0 0 1 5 n/a 19.25 7/7 {"CalendarSemanticsError": 1}
cand-fresh-r3 0 0 1 0 5 0.00% 28.61 7/7 {}
Candidate − control Difference [95% family interval] Control spread Candidate spread
Precision among scored answers -50.00% [-50.00%, -50.00%] 100.00% 0.00%
Comparison unverified 2.67 [2.67, 2.67] 2.00 0.00
Correct completed -0.33 [-0.33, -0.33] 1.00 0.00
First-turn execution errors -1.00 [-1.00, -1.00] 0.00 0.00
Fully correct conversations 0.00 [0.00, 0.00] 0.00 0.00
p95 application latency, shared host (s) -5.34 [-5.34, -5.34] 5.01 13.91
Refusals -1.33 [-1.33, -1.33] 2.00 1.00
Guard refusals -1.33 [-1.33, -1.33] 2.00 1.00
Verifier refusals 0.00 [0.00, 0.00] 0.00 0.00
Silent wrong 0.00 [0.00, 0.00] 1.00 1.00

One family cluster: this interval is degenerate and does not quantify within-family or all-98 uncertainty. It cannot establish the inferential acceptance target.

The mean correct-completion difference is -0.33 per repeat, inside or equal to the control’s observed range of 1.00. The range is an observed noise floor, not a confidence bound.

Real target sample (eight attempts; five direct source IDs)

Mean [minimum, maximum] from three repeats. Undefined values are excluded from means and remain n/a in the per-repeat table; the number of defined values is retained in results.json.

Metric Control Candidate
First-turn execution errors 0.00 [0.00, 0.00] 0.33 [0.00, 1.00]
Correct completed 0.33 [0.00, 1.00] 0.67 [0.00, 1.00]
Silent wrong 0.33 [0.00, 1.00] 0.00 [0.00, 0.00]
First-turn error rate 0.00% [0.00%, 0.00%] 5.56% [0.00%, 16.67%]
Precision among scored answers 50.00% [50.00%, 50.00%] 100.00% [100.00%, 100.00%]
Refusals 1.33 [1.00, 2.00] 0.00 [0.00, 0.00]
Guard refusals 1.33 [1.00, 2.00] 0.00 [0.00, 0.00]
Verifier refusals 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
Comparison unverified 6.00 [4.00, 7.00] 6.33 [6.00, 7.00]
Fully correct conversations 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
p95 application latency, shared host (s) 20.21 [17.07, 22.16] 17.86 [15.39, 21.57]
Task errors 0.00 [0.00, 0.00] 1.00 [0.00, 2.00]
Infrastructure failures 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
Blocked by parent 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
Arm / repeat First-turn errors Correct Silent wrong Refused Unverified comparison Precision p95 shared-host seconds Timed/total Refusal guards
base-real-r1 0 0 0 1 7 n/a 21.41 8/8 {"CalendarSemanticsError": 1}
base-real-r2 0 0 0 1 7 n/a 22.16 8/8 {"CalendarSemanticsError": 1}
base-real-r3 0 1 1 2 4 50.00% 17.07 8/8 {"CalendarSemanticsError": 2}
cand-real-r1 0 1 0 0 7 100.00% 15.39 8/8 {}
cand-real-r2 0 0 0 0 6 n/a 16.63 8/8 {}
cand-real-r3 1 1 0 0 6 100.00% 21.57 8/8 {}
Candidate − control Difference [95% family interval] Control spread Candidate spread
Precision among scored answers 50.00% [50.00%, 50.00%] 0.00% 0.00%
Comparison unverified 0.33 [0.33, 0.33] 3.00 1.00
Correct completed 0.33 [0.33, 0.33] 1.00 1.00
First-turn execution errors 0.33 [0.33, 0.33] 0.00 1.00
Fully correct conversations 0.00 [0.00, 0.00] 0.00 0.00
p95 application latency, shared host (s) -0.97 [-0.97, -0.97] 5.08 6.18
Refusals -1.33 [-1.33, -1.33] 1.00 0.00
Guard refusals -1.33 [-1.33, -1.33] 1.00 0.00
Verifier refusals 0.00 [0.00, 0.00] 0.00 0.00
Silent wrong -0.33 [-0.33, -0.33] 1.00 0.00

One family cluster: this interval is degenerate and does not quantify within-family or all-98 uncertainty. It cannot establish the inferential acceptance target.

The mean correct-completion difference is 0.33 per repeat, inside or equal to the control’s observed range of 1.00. The range is an observed noise floor, not a confidence bound.

Fresh excluding target (115 attempts)

Mean [minimum, maximum] from three repeats. Undefined values are excluded from means and remain n/a in the per-repeat table; the number of defined values is retained in results.json.

Metric Control Candidate
First-turn execution errors 27.00 [26.00, 29.00] 25.67 [24.00, 27.00]
Correct completed 2.00 [1.00, 3.00] 0.00 [0.00, 0.00]
Silent wrong 29.67 [26.00, 34.00] 26.00 [24.00, 29.00]
First-turn error rate 24.32% [23.42%, 26.13%] 23.12% [21.62%, 24.32%]
Precision among scored answers 6.41% [3.33%, 10.34%] 0.00% [0.00%, 0.00%]
Refusals 37.00 [35.00, 39.00] 41.33 [39.00, 44.00]
Guard refusals 37.00 [35.00, 39.00] 41.33 [39.00, 44.00]
Verifier refusals 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
Comparison unverified 19.00 [18.00, 20.00] 21.33 [17.00, 26.00]
Fully correct conversations 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
p95 application latency, shared host (s) 31.41 [27.67, 35.19] 29.38 [27.12, 33.00]
Task errors 27.33 [26.00, 29.00] 26.33 [25.00, 28.00]
Infrastructure failures 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
Blocked by parent 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
Arm / repeat First-turn errors Correct Silent wrong Refused Unverified comparison Precision p95 shared-host seconds Timed/total Refusal guards
base-fresh-r1 29 1 29 37 19 3.33% 31.37 115/115 {"CalendarSemanticsError": 32, "NativeDealUnitError": 4, "UNKNOWN-GOLD-COLUMN": 1}
base-fresh-r2 26 3 26 39 20 10.34% 35.19 115/115 {"AdditiveGrainError": 1, "CalendarSemanticsError": 37, "NativeDealUnitError": 1}
base-fresh-r3 26 2 34 35 18 5.56% 27.67 115/115 {"CalendarSemanticsError": 33, "NativeDealUnitError": 2}
cand-fresh-r1 26 0 24 39 26 0.00% 28.03 115/115 {"CalendarSemanticsError": 34, "LossReasonSourceError": 1, "NativeDealUnitError": 4}
cand-fresh-r2 24 0 25 44 21 0.00% 33.00 115/115 {"CalendarSemanticsError": 41, "NativeDealUnitError": 3}
cand-fresh-r3 27 0 29 41 17 0.00% 27.12 115/115 {"CalendarSemanticsError": 37, "NativeDealUnitError": 4}
Candidate − control Difference [95% family interval] Control spread Candidate spread
Precision among scored answers -6.32% [-13.48%, -1.28%] 7.01% 0.00%
Comparison unverified 2.33 [-2.40, 6.73] 2.00 9.00
Correct completed -2.00 [-3.87, -0.35] 2.00 0.00
First-turn execution errors -1.33 [-5.72, 3.43] 3.00 3.00
Fully correct conversations 0.00 [0.00, 0.00] 0.00 0.00
p95 application latency, shared host (s) -3.98 [-8.23, 0.90] 7.52 5.88
Refusals 4.33 [-2.53, 11.05] 4.00 5.00
Guard refusals 4.33 [-2.53, 11.05] 4.00 5.00
Verifier refusals 0.00 [0.00, 0.00] 0.00 0.00
Silent wrong -3.67 [-7.73, 0.33] 8.00 5.00

The mean correct-completion difference is -2.00 per repeat, inside or equal to the control’s observed range of 2.00. The range is an observed noise floor, not a confidence bound.

The registered regression check nevertheless fails because the paired interval is wholly negative: [-3.87, -0.35].

Real excluding target (93 attempts)

Mean [minimum, maximum] from three repeats. Undefined values are excluded from means and remain n/a in the per-repeat table; the number of defined values is retained in results.json.

Metric Control Candidate
First-turn execution errors 11.00 [10.00, 13.00] 11.33 [10.00, 14.00]
Correct completed 1.00 [1.00, 1.00] 1.00 [1.00, 1.00]
Silent wrong 22.67 [19.00, 27.00] 27.00 [25.00, 29.00]
First-turn error rate 13.75% [12.50%, 16.25%] 14.17% [12.50%, 17.50%]
Precision among scored answers 4.31% [3.57%, 5.00%] 3.58% [3.33%, 3.85%]
Refusals 30.33 [25.00, 35.00] 26.67 [25.00, 30.00]
Guard refusals 30.33 [25.00, 35.00] 26.67 [25.00, 30.00]
Verifier refusals 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
Comparison unverified 21.33 [18.00, 23.00] 21.33 [19.00, 24.00]
Fully correct conversations 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
p95 application latency, shared host (s) 29.21 [25.03, 35.63] 31.83 [29.22, 36.18]
Task errors 17.67 [17.00, 19.00] 17.00 [16.00, 19.00]
Infrastructure failures 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
Blocked by parent 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
Arm / repeat First-turn errors Correct Silent wrong Refused Unverified comparison Precision p95 shared-host seconds Timed/total Refusal guards
base-real-r1 10 1 22 35 18 4.35% 35.63 93/93 {"-": 1, "CalendarSemanticsError": 31, "NativeDealUnitError": 2, "WindowBindingError": 1}
base-real-r2 13 1 19 31 23 5.00% 25.03 93/93 {"-": 1, "AdditiveGrainError": 1, "CalendarSemanticsError": 28, "NativeDealUnitError": 1}
base-real-r3 10 1 27 25 23 3.57% 26.96 93/93 {"-": 1, "CalendarSemanticsError": 22, "NativeDealUnitError": 2}
cand-real-r1 14 1 29 25 19 3.33% 36.18 93/93 {"-": 1, "CalendarSemanticsError": 23, "NativeDealUnitError": 1}
cand-real-r2 10 1 27 25 24 3.57% 29.22 93/93 {"-": 1, "CalendarSemanticsError": 20, "FanOutSumError": 1, "NativeDealUnitError": 2, "WindowBindingError": 1}
cand-real-r3 10 1 25 30 21 3.85% 30.09 93/93 {"-": 1, "CalendarSemanticsError": 28, "NativeDealUnitError": 1}
Candidate − control Difference [95% family interval] Control spread Candidate spread
Precision among scored answers -0.65% [-2.75%, 0.00%] 1.43% 0.51%
Comparison unverified -0.00 [-3.44, 3.83] 5.00 5.00
Correct completed 0.00 [0.00, 0.00] 0.00 0.00
First-turn execution errors 0.33 [-2.63, 3.09] 3.00 4.00
Fully correct conversations 0.00 [0.00, 0.00] 0.00 0.00
p95 application latency, shared host (s) 1.85 [-0.72, 8.04] 10.59 6.96
Refusals -3.67 [-7.08, -0.64] 10.00 5.00
Guard refusals -3.67 [-7.08, -0.64] 10.00 5.00
Verifier refusals 0.00 [0.00, 0.00] 0.00 0.00
Silent wrong 4.33 [1.44, 7.38] 8.00 4.00

The mean correct-completion difference is 0.00 per repeat, inside or equal to the control’s observed range of 0.00. The range is an observed noise floor, not a confidence bound.

Full 98-question target

Population Source questions Direct corrected-oracle mappings Correct / wrong / refused per repeat Mean / min / max Paired interval
Full target 98 5 Unavailable for all 98 Unavailable Unavailable

The inventory contains 39 first questions and 59 follow-ups. User ancestor chains exist, but historical answer/result state is absent. Replaying them would create new populations and would not supply the 93 missing corrected references. target98-coverage.jsonl retains each ID, shape, envelope status and per-arm/repeat observed outcome or explicit missing-measurement status.

Target case outcomes

Population / case Shape Envelope Control r1 / r2 / r3 Candidate r1 / r2 / r3
fresh:F10:2-quotes-30d-no-trade count, lifetime_gate, quote_details, quote_threshold inside comparison_unverified / comparison_unverified / correct_completed comparison_unverified / comparison_unverified / comparison_unverified
fresh:F10:45d-no-booking-list client_list, lifetime_gate inside task_error / task_error / task_error comparison_unverified / comparison_unverified / comparison_unverified
fresh:F10:conv-may-drill:t1 count, lifetime_gate inside comparison_unverified / refusal_unverified / refusal_unverified comparison_unverified / comparison_unverified / comparison_unverified
fresh:F10:conv-may-drill:t2 client_list, lifetime_gate inside task_error / silent_wrong / task_error task_error / task_error / task_error
fresh:F10:conv-may-drill:t3 count, lifetime_gate, quote_details inside task_error / refusal_unverified / refusal_unverified refusal_unverified / refusal_unverified / silent_wrong
fresh:F10:july-never-booked count, lifetime_gate inside refusal_unverified / comparison_unverified / refusal_unverified comparison_unverified / comparison_unverified / comparison_unverified
fresh:F10:q2-same-window client_list, count, same_period_gate inside comparison_unverified / comparison_unverified / comparison_unverified comparison_unverified / comparison_unverified / comparison_unverified
real:F10:056d2c9babe2 count, lifetime_gate inside comparison_unverified / refusal_unverified / comparison_unverified comparison_unverified / comparison_unverified / task_error
real:F10:60e720a241cb count inside refusal_unverified / comparison_unverified / refusal_unverified comparison_unverified / comparison_unverified / comparison_unverified
real:F10:8690cb315864 client_list, count, prior_booking, quote_details inside comparison_unverified / comparison_unverified / correct_completed correct_completed / comparison_unverified / comparison_unverified
real:F10:878f1692df8e client_list, count, historical_trade_details, quote_threshold inside comparison_unverified / comparison_unverified / comparison_unverified comparison_unverified / comparison_unverified / comparison_unverified
real:F10:c5e4b0c91f6a client_list outside comparison_unverified / comparison_unverified / comparison_unverified comparison_unverified / comparison_unverified / comparison_unverified
real:F10:conv-f10-drill:t1 count, lifetime_gate inside comparison_unverified / comparison_unverified / comparison_unverified comparison_unverified / comparison_unverified / correct_completed
real:F10:conv-f10-drill:t2 client_list, lifetime_gate inside comparison_unverified / comparison_unverified / silent_wrong comparison_unverified / task_error / comparison_unverified
real:F10:conv-f10-drill:t3 count, lifetime_gate, quote_details inside comparison_unverified / comparison_unverified / refusal_unverified comparison_unverified / task_error / comparison_unverified

Acceptance

Target Verdict Scope
98 real-question family gain FAIL — NOT ESTABLISHED 93 direct references missing; all-98 coverage/interval unavailable.
Fresh non-target regression FAIL A registered non-target regression condition failed on measured cases.
Measured real target inferential gain FAIL — NOT ESTABLISHED Eight attempts, one family cluster; degenerate interval is not uncertainty evidence.
Real non-target regression PASS No regression detected under the registered conditions on measured cases; this is not equivalence.
Zero observed curated invariant violations PASS Executed present checks only; skipped/absent inherited checks remain unverified.

Remaining observed target classes

Counts pool three candidate repeats. Unverified comparisons and refusals are kept separate from confirmed wrong answers. Observed defect labels can overlap; row-set size, projection-width and recorded answer-contract failures follow D3 reporting conventions and do not isolate SQL root causes.

Population Class Attempts
fresh_target execution_boundary_unisolated 3
fresh_target oracle_comparison_unverified 15
fresh_target refused:CalendarSemanticsError 2
fresh_target row_set_size_mismatch 1
real_target execution_boundary_unisolated 2
real_target oracle_comparison_unverified 19
real_target planner_error_unresolved 1

Full target coverage gaps: 93 source questions lack direct corrected references; 59 source questions are follow-ups with unavailable historical result state. These counts overlap.

Transitions against control

Equal-weight Cartesian pairing of all nine repeat pairs, matched by case ID and scaled to one population repeat. These are descriptive transitions, not proof of a causal correction.

fresh

Control outcome Candidate outcome Mean count
comparison_unverified comparison_unverified 15.78
comparison_unverified refusal_unverified 2.89
comparison_unverified silent_wrong 0.89
comparison_unverified task_error 1.78
correct_completed comparison_unverified 0.89
correct_completed refusal_unverified 1.00
correct_completed silent_wrong 0.11
correct_completed task_error 0.33
refusal_unverified comparison_unverified 4.78
refusal_unverified refusal_unverified 25.78
refusal_unverified silent_wrong 4.78
refusal_unverified task_error 3.67
silent_wrong comparison_unverified 2.67
silent_wrong refusal_unverified 5.56
silent_wrong silent_wrong 19.33
silent_wrong task_error 2.44
task_error comparison_unverified 2.22
task_error refusal_unverified 6.78
task_error silent_wrong 1.22
task_error task_error 19.11

real

Control outcome Candidate outcome Mean count
comparison_unverified comparison_unverified 21.33
comparison_unverified correct_completed 0.56
comparison_unverified refusal_unverified 2.22
comparison_unverified silent_wrong 1.11
comparison_unverified task_error 2.11
correct_completed comparison_unverified 0.22
correct_completed correct_completed 1.11
refusal_unverified comparison_unverified 3.67
refusal_unverified refusal_unverified 21.33
refusal_unverified silent_wrong 5.22
refusal_unverified task_error 1.44
silent_wrong comparison_unverified 1.78
silent_wrong refusal_unverified 1.78
silent_wrong silent_wrong 18.33
silent_wrong task_error 1.11
task_error comparison_unverified 0.67
task_error refusal_unverified 1.33
task_error silent_wrong 2.33
task_error task_error 13.33

real_target

Control outcome Candidate outcome Mean count
comparison_unverified comparison_unverified 4.78
comparison_unverified correct_completed 0.56
comparison_unverified task_error 0.67
correct_completed comparison_unverified 0.22
correct_completed correct_completed 0.11
refusal_unverified comparison_unverified 1.11
refusal_unverified task_error 0.22
silent_wrong comparison_unverified 0.22
silent_wrong task_error 0.11

Validation and model evidence

The retained affected suite passed 1151 tests, with 0 failures, 0 errors and 88 skips. Its present curated subset passed 636, with 0 failures and 88 skips. The 19 absent inherited files and all skipped checks remain unverified. The full receipt is checks/summary.json.

Check Retained result
Affected suite 1151 passed, 88 skipped, 13 warnings in 193.34s (0:03:13)
tests/test_insights_f01_definition.py 71 passed in 13.72s
tests/test_insights_f01_harness.py 39 passed in 14.03s
tests/test_insights_f01_harness.py: unbounded idle admission and explicit empty-admission resume 63 passed in 15.85s
tests/test_insights_f01_harness.py: WAL-safe empty-admission resume 68 passed in 15.74s
tests/test_insights_f01_reporting.py 12 passed in 13.23s
tests/test_insights_f01_reporting.py: final arithmetic report corrections 12 passed in 13.77s
tests/test_insights_f01_reporting.py: interrupted attempt model provenance 13 passed in 13.96s
tests/test_insights_f01_reporting.py: final wording corrections 13 passed in 12.39s
Final prompt checks 103 passed, 8 warnings in 26.87s
Changed-file Ruff All checks passed
Project mypy Exit 0; no diagnostics
Project Ruff 11 inherited diagnostics, all in unchanged files

The earlier setup failure is retained separately in the check receipt. No failed or skipped check is converted to a pass by summing later checks.

Run Calls Failed Response model IDs Completed calls missing identity Attempts without mapped completed call
base-fresh-r1 248 0 {"glm-5.3": 248} 0 1
base-fresh-r2 257 0 {"glm-5.3": 257} 0 0
base-fresh-r3 251 0 {"glm-5.3": 251} 0 0
cand-fresh-r1 248 0 {"glm-5.3": 248} 0 1
cand-fresh-r2 255 1 {"glm-5.3": 254} 0 1
cand-fresh-r3 248 0 {"glm-5.3": 248} 0 1
base-real-r1 180 1 {"glm-5.3": 179} 0 2
base-real-r2 177 0 {"glm-5.3": 177} 0 2
base-real-r3 175 0 {"glm-5.3": 175} 0 2
cand-real-r1 172 0 {"glm-5.3": 172} 0 2
cand-real-r2 173 0 {"glm-5.3": 173} 0 3
cand-real-r3 170 0 {"glm-5.3": 170} 0 2

A response model ID is provider-reported identity, not an independent serving attestation. An attempt with no mapped completed generator call has unknown response identity; it is not assigned the requested model. Failed and unmapped calls remain visible in the traces.

New-run admission and integrity

Run Application Harness Lock wait (s) Terminal completion Listener clear Frozen inputs unchanged
cand-fresh-r1 da0010da 3212565d 0.94 True True True
cand-fresh-r2 da0010da 3212565d 0.86 True True True
cand-fresh-r3 da0010da 3212565d 0.91 True True True
base-real-r2 b24c1c3f 3212565d 187.89 True True True
base-real-r3 b24c1c3f 3212565d 1010.86 True True True
cand-real-r1 da0010da 3212565d 0.91 True True True
cand-real-r2 da0010da 3212565d 0.92 True True True
cand-real-r3 da0010da 3212565d 0.99 True True True

Review ledger

Pre-launch reviews: mechanics and method, separate Codex CLI calls requesting gpt-6-astra with high reasoning, strict no modifications. Both completed; their finding-to-action ledger is in the frozen plan.

Premeasurement code review: completed receipt, three independent lenses (correctness, security, adversarial), six parent-context lenses; fresh validator dispatch was unavailable at the CLI thread limit. Its initial verdict was not ready. Confirmed fixes below preceded every scored run. A separate read-only validation checked the reporting fixes. No claim of full independent code-review coverage is made.

Finding Action and verification
Contradictory non-target acceptance wording Use the stricter original correct-completion rule: loss beyond control range OR wholly negative interval. Silent-wrong spread-and-positive-interval condition retained. Synthetic boundary tests cover both.
Automatic report/inventory exports bypass private staging Generate under F01 private publish/planning directories; compact recomputation uses a separate private directory. Repository copies are explicit sanitized delivery preparation.
Standalone imports can emit bytecode Disable bytecode before local imports; initialize private cache/temp environment before report imports.
Cancellation can restart or advance Record cancellation, safely cancel/reap queued flock or await active cleanup, then exit the whole sequence. Synthetic queue/active cancellation tests.
Admission worker can outlive its wrapper Probe wrapper execs its worker; admission registers that process for centralized ownership and cleanup. Synthetic interruption/cleanup tests.
Follow-up booking override loses to old text Resolve current explicit booking override before inherited gate; activation/signup/historical-metric clauses do not change it.
Quote-period words can override lifetime booking gate Restrict override matching to booking clauses; synthetic tests cover both clause orders and explicit same-period wording.
Plural bookings missed Recognize plural booking nouns and exercise context, obligation and refusal paths.
Inherited snapshot writer precedes path checks Validate apps, concrete snapshot and manifest destinations before invoking the writer; symlink tests verify writer is never called.
Child diagnostic text may contain customer data Drain bounded stdout/stderr chunks to byte counts and SHA256 only; retain full attempts/traces in the private evidence directory.
Existing child symlink can redirect inherited export Validate existing export descendants before D3 export. Independent read-only follow-up confirmed closure; synthetic no-write escape test added.
FAIL table explained as no regression Make acceptance explanation conditional; independent read-only follow-up confirmed closure, synthetic wording assertion added.

The final source-contract inspection also clarified that the SQL-only generator must substitute every named scaffold parameter with a resolved SQL expression or literal; the executable builder still returns bound parameters for contract tests. This clarification preceded measurement and was checked by the focused generation/definition suite.

The application changes derive from ratified definitions, question shapes, source inspection and synthetic cases. No candidate answers were available during these corrections. Final result reviews and any report-only corrections follow below.

Empty-admission method parity correction

Read-only compact-export audit: no question/SQL/customer fields found in inspected exports; full98 and measured8 are separated. One misleading coverage field used true for first questions, although parent state is not applicable. Changed it to null; follow-up parent-state availability remains false. This changes a coverage description only.

Independent admission review caught a preservation issue before the resume guard was used: SQLite mode=ro may still create WAL sidecars, and its transaction context manager does not close the connection. Use immutable read-only access with explicit closure after rejecting unapplied WAL content; cover sidecar preservation and rejection in synthetic tests. No retained measurement database was opened by the unsafe check.

Before any scored question, inspection found an added 1,800-second idle timeout absent from D3. The empty admission was interrupted and retained; no application server or question was started. Restore unbounded admission while preserving the 95% gate and add explicit, fail-closed resume for this verified empty cancellation only. Keep the application revision frozen and record a separate harness revision. This correction follows a read-only method advisory that rejected relaxing the idle threshold without authorization. The admission receipt confirms cleanup, a clear listener and unchanged frozen inputs.

Resume evidence preparation

Resume verification: 12 passed in 13.23s; no failures or skips. Ruff passed for all 20 changed Python files. All eight new measurement receipts are complete and pass cleanup/input-integrity checks; the main checkout tracked diff, status and HEAD match the recorded resume baseline.

Export verification covers all 1,460 compact attempts, including 122 interrupted-run attempts. The repository ignores JSONL evidence by default; the approved compact records are explicitly included after scanning. Compact-only results match raw-derived results exactly. Review before/after hashes include ignored evidence and private raw JSON/JSONL records.

Final arithmetic review

Independent Codex CLI review requesting gpt-6-astra with high reasoning exited 0, changed no protected file (6,771 checked), and independently reproduced all 60 metric/interval entries, all transitions and all 1,338 scored raw/compact attempts. Arithmetic PASS; family acceptance FAIL. Correct stale requested-only model-provenance metadata to distinguish configuration from recorded response identity and unknown unmapped calls. Add direct-source correct counts (control 0/0/1; candidate 1/0/0) to the report and PR: the eight-attempt sample’s additional success includes the authored conversation. Neither correction changes any measured outcome or acceptance rule. The reviewer did not independently rescore SQL/reference answers or attest serving identity.

Post-correction reporting validation: 12 passed in 13.77s; changed-file Ruff passed for all 20 Python files.

Final leakage review

Independent Codex CLI review requesting gpt-6-astra with high reasoning exited 0 and changed no protected file (6,772 checked). It independently matched all 1,338 scored model joins, 12 correctness/silent-wrong intervals, 95 transition cells, 72 per-repeat precision/p95 values and 299 private-file hash entries. No leakage was detected within its stated inspected boundary; this is bounded evidence, not an anonymization certification. Family acceptance FAIL is supported.

Final wording review

A third independent Codex CLI review requesting gpt-6-astra with high reasoning exited 0 and changed no protected file (6,778 checked). It independently reproduced all 60 bootstrap entries, 95 transition cells, 72 per-repeat precision/p95 values, all 1,460 attempt identity annotations and 299 private-file hashes. Arithmetic PASS; family acceptance FAIL. No independent serving attestation is claimed for any review.

All final corrections affect report/evidence wording, publication preparation or provenance completeness. No application tuning, rescoring or additional measurement followed the scored runs.

Final report-source verification: 13 passed in 12.39s; changed-file Ruff passed for all 20 Python files. Exact review prompts and replies retain their original formatting and hashes.

Delivery launcher correction

The first publication launch exited before Wrangler could run: isolated XDG configuration made the mise shim select Node 20.20.2, below Wrangler’s required Node 22. A credential-free explicit-shim version command reproduced the exact retained stderr hash. Pin the existing Node 22.22.3 executable while keeping private caches and environment-only credentials; its version probe succeeds. Preserve the failed launch receipt. This launcher-only change followed the independent reviews; the reviewed ancestry/staging logic is unchanged, and all nine synthetic boundary checks passed again against the new source hash.

The cached Wrangler 4.131.0 also automatically delegated an agent-run Pages command to Workers, producing authentication error 10000 on a Workers service request. Local command-source inspection identified --force as the explicit Pages path for both creation and deployment. Use that flag to bypass delegation; no credential scope is expanded. This changes publication routing only, after review, while preserving the reviewed report-only staging boundary.

Direct Pages creation succeeded. Before upload, the regular-assets guard rejected Wrangler’s generated cache directory inside staging. Preserve that private attempt and give Wrangler a separate fresh private working directory, leaving only the report and headers in upload staging. No upload occurred before the guard stopped the attempt. This delivery-only correction preserves the reviewed boundary rather than relaxing it.

Final publisher validation adds a simulated CLI that creates a working-directory cache and verifies it stays outside both uploaded assets. All ten checks passed (the original nine boundaries plus this cache-flow check), using synthetic credentials and no remote calls or real credential reads.

Evidence and limits

Plan: docs/plans/2026-09-12-lore-goal-f01-plan.md. Compact records and recomputable results: docs/evidence/insights-goal-f01-2026-09-12/. Full records: /tmp/lore-goal3-eval/goal-f01/. scripts/insights_f01_report.py --compact recomputes metrics, paired intervals and transitions. Bootstrap: 10,000 whole-family draws, seed 2026090611, nearest-rank percentiles, pooled repeat counts normalized to each population; undefined draws are counted. Precision and p95 pool before differencing, so their points can differ from differences of repeat means.

No application was deployed or merged. Expected merge order: #104 safety release, sibling DCLASS, then F01 after resolving any overlap and rerunning affected checks. The sibling checkout and branch were not inspected or modified. These measurements do not verify the combined sibling/F01 result.