Sovereign DESIGN and DECOMPOSE on gpt-oss:20b: before and after the length-aware retry¶
- Object: how often
GptossloopDesignProvider.batch(thedesign.interviewstep) andGptossloopWorkPlanProducer.decompose(build.decompose) return a usable answer from gpt-oss:20b, before and after branchfix/sovereign-length-aware-retry - Host: the operator's macOS 26.6.2 workstation, ollama 0.34.4,
gpt-oss:20bdigest17052f91a42e; no other model request was in flight during either run - Ceilings: the defaults, no fit record (
VIBEY_OLLAMA_CONTEXT,VIBEY_OLLAMA_OUTPUT,VIBEY_OLLAMA_FITunset): context ceiling 8192, output 2048, temperature 0, JSON mode. Thenum_ctx/num_predict/thinkactually sent are recorded per exchange in the JSON records - Cutoff: before 2026-09-29T07:36:40Z to 08:04:42Z; after 2026-09-29T08:04:55Z to 08:37:39Z
- Instrument:
scripts/sovereign_retry_experiment.py --runs 10 --workloads design,decompose --replay-ledger <export>, which calls the production producers overOllamaChatClient.from_environmentwith only a recording transport added - Records:
sovereign-retry-2026-09-29-before.json,sovereign-retry-2026-09-29-after.json
Workloads¶
design.interview: 10 synthetic runs, cycling the seven DESIGN stages over three one-paragraph briefs (10 distinct prompts).design.interview (production replay): the DESIGN-phase ledger of the project whosedesign.interviewjob had failed 11 attempts withKeyError: 'question_id'(GitHub issue #133's project, 34 events), asked for the stage it was stuck on (walking_skeleton), 10 times. The ledger export is not committed.build.decompose: 10 runs cycling three synthetic specs of 3 to 5 criteria (3 distinct prompts).
Results¶
| workload | N | ok | empty content (done_reason) | schema violation | first call cut short (length) |
mean calls | mean latency |
|---|---|---|---|---|---|---|---|
| design.interview, before | 10 | 0 | 0 | 10 (KeyError: 'question_id') |
0 | 1.0 | 20.7 s |
| design.interview, after | 10 | 10 | 0 | 0 | 0 | 1.0 | 25.1 s |
| production replay, before | 10 | 0 | 0 | 10 (questions must be a list) |
0 | 1.0 | 24.4 s |
| production replay, after | 10 | 10 | 0 | 0 | 0 | 1.0 | 77.8 s |
| build.decompose, before | 10 | 0 | 7 (length x7) |
3 (verification must be an object) |
7 | 1.7 | 123.1 s |
| build.decompose, after | 10 | 10 | 0 | 0 | 10 (4 empty, 6 truncated JSON) | 2.0 | 93.5 s |
What the before run shows:
- Every DESIGN failure was a shape failure, not a transport one: in JSON mode the prompt carried no schema, so the model chose its own key names. The synthetic runs reproduced the production
KeyError: 'question_id'; the production replay put the questions under another key. - 7 of 10 decompose runs spent the whole 2048-token budget reasoning (
done_reason: length, empty content, a fullthinkingchannel). The existing empty-reply retry resent the same budget in the same mode, and every one of the 7 failed identically on the retry -- the productionOllama response carried empty message content.
What the after run shows:
- Stating the schema in the prompt removed every shape failure; no answer needed the re-ask (mean calls 1.0 for DESIGN), so the re-ask path is proven by unit tests only, not by this run.
- Every decompose first call was cut short at 2048 tokens (the schema lengthens the prompt and the answer), and every one was completed by the single widened retry:
num_predict4096 withinnum_ctx4931-5028,think: "low". The retries used 446-907 output tokens, so the lighter reasoning, not the larger budget, is what made room; the two levers were not separated.
Not measured¶
- Determinism: at temperature 0 an unchanged prompt repeats its answer (the replay's 10 runs and each spec's repeats returned identical token counts), so N counts runs, not independent samples: 10 distinct DESIGN prompts, 1 replay prompt, 3 decompose prompts.
- The latency cost of the stated schema on a large ledger (24 s to 78 s on the replay) was measured once on one host; whether it matters against the worker's 900 s timeout was not tested beyond this.
design.researchanddesign.synthesizego through the sameValidatedAskafter the change and were not replayed live.- The
OutputBudgetExhaustedpath (a second cut-short reply) did not occur live; it is covered by unit tests. - A fitted ceiling (
VIBEY_OLLAMA_FIT), other models (qwen3:14b acceptsthink: "low"but thinks regardless) and other hosts.