0058 — Concurrent runs of a local model are measured on each device, and unmeasured or stale means one¶
Status: accepted as mechanism; the amendment to sub-doctrine 8.c it proposes is not ratified by this record — only the operator's merge of the separate canon pull request ratifies it (Constitution Article II.3) · Date: 2026-09-24 · Cites: sub-doctrines 8.c, 8.j, 8.g, 10.f, 12.c, 12.e, and ADR-0045 · Related: ADR-0016, ADR-0018, ADR-0020, ADR-0045, ADR-0054 · Evidence: docs/architecture/evidence/slots-2026-09-24-mac17-2.json and its .md summary, produced by vibey-gh slots calibrate on this device
Owes: the conduct is the 8.c amendment drafted below, which is ratified separately
(ADR-0020: the record argues, the canon states). This record owes the advertised ADR count in
CLAUDE.md, AGENTS.md, GEMINI.md, README.md and docs/index.md
(tests/meta/test_adr_counts.py), a nav entry in properdocs.yml, the [local_models] and
vibey-gh slots rows in the vibey-gh configuration and CLI references, and pointers to them in
docs/reference/configuration.md and docs/reference/cli.md. It leaves two debts open, named
under Consequences: a per-node calibration Job in the Helm chart, and durable storm run
records.
Context¶
Sub-doctrine 8.c runs each loop as a single instance fed by a queue, and says of a model on the operator's own hardware: one run at a time. 8.j says every part of the family fits itself to the machine by measurement, never by a hunch, and never trades correctness for speed. Those two sentences meet at one number — how many runs of a resident model the machine serves at once — and on 2026-09-23 the operator ruled on it: "C — measure, then decide." Keep 8.c as written, benchmark one slot against two by replaying storm-depth turns, and bring numbers before any change to the canon. On 2026-09-24 the operator widened the ruling: "we should not stop at two, test and find out the ideal number, and same thing on all other devices that you run on; make that ADR / law."
The claim on the table was that two 32k-token slots reserve about the same KV cache as one 65k-token slot, and that few storm turns exceed 32k — so two half-size slots might double throughput at no memory cost. ADR-0045 had counted 71 of 838 recorded turns over 32,768 tokens.
A first run of this measurement was lost. It ran on the morning of 2026-09-24 against the
storm's own lane records under /private/tmp; at 09:09 the machine rebooted, /private/tmp
was wiped, and the records, the corpus, the evidence and the unpushed code went with it. The
Ollama app also applied a staged update on relaunch (0.34.2 → 0.34.4). Nothing measured that
morning is cited here as evidence. Two failure modes it met are the reason for two mechanisms
below: a runner that answers HTTP 200 with no done_reason when its decode fails underneath
it, and another client loading a model on the production runner mid-step.
Method¶
The corpus is storm-shaped, and says so. qwenloop records each run's tool results, answer
text, retries and per-turn input tokens (events.jsonl) under the storm's scratch directory,
not its payloads, and those records did not survive the reboot. So
docs/plans/qwenstorm-3.0.0/tools/storm_turn_pool.py specs builds the pool from durable,
committed material: each run is one of the storm's committed lane specs, planned with
qwenloop's own build_plan and system prompt, followed by read_file turns — with their real
arguments — whose results are the repository's real files at the corpus's commit, truncated as
qwenloop truncates them. That is the work lane time went to: reading code into context. When a
storm has fresh records, storm_turn_pool.py lanes rebuilds its real payloads instead, with the
same qwenloop functions. Depth in the pool is characters ÷ 3 (the fit calculus's conservative
estimate), used only to stratify and to keep every payload inside the window; the runner's own
prompt_eval_count on replay is the measured depth.
The pool is 1,096 turns in 40 runs (40 specs drawn with seed 0; sha256 508685035d8a645e…).
vibey-gh slots corpus drew 20 segments of 3 consecutive turns (60 turns per step), five
segments in each depth stratum by estimate — below 16k, 16–32k, 32–48k, and 48k and over — so
the deep tail is measured, not sampled around (corpus sha256 ae63b871ae3b62dc…).
The sweep. vibey-gh slots calibrate starts its own ollama serve — the Ollama app's own
bundled binary, the production runner's version — on port 11435 with OLLAMA_NUM_PARALLEL=N
and OLLAMA_NOPRUNE, beside the production runner, which it never restarts or reconfigures.
It replays the corpus with N closed-loop workers, each sending one segment's turns in order (so
a worker keeps its prefix warm exactly as a lane does), through /api/chat with
truncate: false, shift: false, temperature 0, seed 42 and 768 output tokens. It samples
every second: wired memory and the swap counters from vm_stat, the free share from
memory_pressure, and what both runners hold resident from /api/ps. It reads the runner's
own log for n_slots, n_ctx_slot, KV cache sizes, loads, truncations, context shifts and
Metal device failures. One slot is measured twice, so fidelity has a baseline. A step during
which the production runner held a model is discarded and measured again once production is
idle. It stops at a broken bound or after two steps without a 10% gain. Every completed step is
checkpointed under ~/.local/state/vibey-gh/slots/progress/, so this run was walked one step
at a time (--max-runs 1, 2, …), each step's evidence committed and pushed before the next
began — a push's own test suite never ran beside a measurement.
The bounds ([local_models], re-read at every decision): peak wired memory within 80% of
physical memory; swap-outs no more than twice the one-slot rate, never judged below 64 MB/min;
no turn that ran at one slot failing; no truncation, context shift, device failure, reload or
eviction; every answer ending stop or length; structural agreement with the one-slot answers
(the tool calls made and their argument names, or plain text) no more than 0.05 below one
slot's agreement with itself. At one slot the memory bounds are reported, never refusing: one is
8.c's floor.
Why beside production, not by restarting it. The brief's method was launchctl setenv
OLLAMA_NUM_PARALLEL and a restart of the Ollama app. A relaunch applies a staged update —
it did so on 2026-09-18 and again at the 09:09 reboot — so a restart per step can change the
runner mid-measurement and makes "restore the original state" impossible. So the production
runner was never touched: nothing was set with launchctl, and restoring was stopping the
calibration runner. The rule ADR-0045 learned — two servers bidding for one GPU breach 8.c — is
kept by refusing to start while production holds a model, and by discarding any step during
which it loads one.
Results¶
Measured on the operator's Mac17,2 (Apple M5, 24 GiB, Apple M5 GPU with 10 cores, macOS
26.6.2), Ollama 0.34.4, gpt-oss:20b (digest 17052f91a42e) at 65,536 tokens per slot —
fingerprint f99b07f610204948 — between 13:39 and 15:42 UTC on 2026-09-24, 60 turns per step.
The replayed prompts measured 5,243 to 43,215 tokens (p50 24,862); 19 of the 55 served at one
slot were longer than 32,768. They carried 0.75 of the pool's characters ÷ 3 estimate, so the
estimate was conservative as intended, and the deepest turns replayed were about 43k tokens,
not the 64k a lane can reach.
| step | slots × ctx | served | turns/h | out tok/s | p50 s | p95 s | peak wired | free min | swap-out | fidelity (structural / exact) | verdict |
|---|---|---|---|---|---|---|---|---|---|---|---|
| N=1 | 1 × 65,536 | 55/60 | 98.2 | 3.02 | 16.5 | 112.6 | 18.42 GB | 3% | 25.1 GB/min | — | floor; memory warned |
| N=1 again | 1 × 65,536 | 55/60 | 112.7 | 3.47 | 12.9 | 102.8 | 18.07 GB | 4% | 17.4 GB/min | 1.0 / 1.0 against the first | baseline |
| N=2 | 2 × 65,536 | 57/60 | 111.9 | 3.28 | 29.8 | 189.9 | 20.17 GB | 0% | 30.8 GB/min | 0.818 / 0.745 | breaks fidelity |
| claim | 2 × 32,768 | 36/60 | 172.7 | 5.74 | 20.6 | 129.3 | 18.16 GB | 2% | 22.8 GB/min | 0.889 / 0.778 | refuses 24 turns |
The ideal on this device is one. The sweep stopped at N=2 on fidelity: at temperature 0 with a fixed seed, one slot gave the same answer twice for every one of 55 turns, and two concurrent slots changed the tool calls or their argument names on 10 of 55 (18%), and the exact answer on 14. 8.j refuses a configuration whole when fidelity moves, whatever it saves — and it saved nothing here: 111.9 turns/h at two slots sits between the two one-slot runs (98.2 and 112.7), whose own difference, 14.8%, is this machine's noise; two slots nearly doubled median latency.
The claim about two 32k slots is half right. Their KV cache costs what one 65k slot's does — the same 65,536 non-sliding cells in the runner's log, and peak wired memory within 0.1 GB — but they refused, with an HTTP 400 and never silently, every turn longer than 32,768 tokens: 24 of 60 here, the shortest refused at 32,786. The 172.7 turns/h counts only the 36 turns it served, and its answers moved too (structural 0.889).
One slot already strains this machine. Loading the model at a 65,536-token window took wired memory from 3.4 GB at rest to 18.1 GB and free memory from 69% to 11%; during every step the free share fell to 0–4% and the machine swapped 17–31 GB a minute each way, the whole time (the per-second timeline is in the evidence). That is recorded as a floor warning, not a refusal — one is 8.c's floor — and it is the operator's workload beside the model as much as the model: the reading is the machine as it was used. Peak wired stayed under the 20.62 GB ceiling at every step; at two slots it came within 0.45 GB of it.
What else happened. The same five turns failed at both one-slot runs with the runner's
error parsing tool call (HTTP 500): the model's own malformed tool calls, which qwenloop
retries (#386); two of them succeeded at two slots, and no turn that one slot served failed at
two. No step saw a truncation, a context shift, a Metal device failure, a model reload, or a
model resident on the production runner, so no reading was discarded. Not everything was
idle beside the first N=1 step: four local commits (one installing pre-commit's environments),
the formatters, a type check and a five-second unit-test run at nice 19 took a few minutes of
CPU on the machine while the model served from the GPU. That step is also the slower of the two
one-slot runs, so the 14.8% noise figure may include it. No push, and no test suite, ran beside
any other step.
Confidence. Sixty turns per step, from 1,096 storm-shaped turns in 40 runs; one sweep; one device. Throughput differences under about 15% are noise on this machine at these swap rates. The fidelity result does not rest on throughput: it is a count of answers that changed, against a baseline that changed none. Not measured: other models, windows and devices; tool execution between turns (idle runner time favours more slots than this replay does); whole-lane prefix reuse (segments are three turns); thermals and power; stability over a night; turns deeper than about 43k tokens; and the storm's own payloads, which were not recorded durably.
Decision¶
- The mechanism is
vibey-gh slots, generic over any device that runs Ollama:corpus,calibrateandallowed, with the seams invibey_gh/interfaces/slots_interface.py(ADR-0016). It reads macOS throughsysctl,system_profiler,vm_statandmemory_pressure, and Linux through/proc,/sysandnvidia-smi. - Every platform assumption lives behind a seam (#1116: Ubuntu 26.04 LTS is
first-class, Windows follows, #1097).
DeviceProbeInterfacestates the machine (DarwinDeviceProbe:sysctl,system_profiler,sw_vers;LinuxDeviceProbe:/proc,/sys,nvidia-smi,/etc/os-release);HostMemorySamplerInterfacereads what it charges (vm_statandmemory_pressure;/proc/meminfo,/proc/vmstatand GPU memory, with cgroup limits to follow);RunnerParallelismInterfacesets the production runner's parallelism and restarts it (MacOSAppParallelism:launchctl setenvand an app restart;SystemdParallelism: a clearly marked stub that renders theEnvironment=drop-in and raises until a Linux host measures it).PlatformProbespicks them; nothing else names an operating system. The calibration's own runner is a plainollama serve, the same on both. - Evidence is keyed to a device fingerprint: hardware model, processor, memory, accelerator, operating system, runner version, model digest and context window. Evidence for another fingerprint is stale; so is evidence older than 30 days.
[local_models] concurrent_runsis the declaration (12.c), default1: 8.c as written, which probes nothing and needs no evidence."measured"takes the ideal N the device's evidence supports, re-judged against today's bounds. A number above one runs only where the device measured it inside every bound and faster than one; otherwise one runs, and the refusal names what is missing. This repository declares1.- Unmeasured or stale means one, out loud, and the gap closes itself (12.e).
slots allowedwrites a calibration request beside the evidence;storm-queue.shasksslots allowedon every pass instead of the host-widepgrepit hard-coded, logs the answer and its reason, and when its queue empties it builds a corpus (its own lane records, else its committed specs) and runsslots calibrate --if-requestedunder the shared model lock. - A calibration that is not clean is not evidence. A step taken while the production
runner held a model is discarded and measured again (twice at most); a run that never gets a
clean step, or that ran on a runner version other than production's, is written to
--outfor the record and not recorded for the device. A 200 without adone_reasonis a failed turn. - A sweep survives a reboot. Each completed step is checkpointed, keyed by fingerprint, corpus and method, and a rerun takes what it has.
Proposed amendment to 8.c (not ratified here)¶
The sentence "The instance takes on as much work at once as its capacity allows — for a model running on the operator's own hardware, one run at a time — and no more." would read:
The instance takes on as much work at once as its capacity allows, and no more. For a model running on the operator's own hardware, that capacity is the number of concurrent runs measured on that device and recorded as evidence (8.j) — measured against the loop's own work, keyed to the device, the runner, the model and its context window, and held to the bounds 8.j names: wired memory within its ceiling, swap not rising, no prompt refused or cut that one run would have served, and every answer as faithful as one run's. Unmeasured or stale means one. A number the evidence does not support is never run, however it is declared, and the operator may always declare fewer (12.c).
It rules the method, not a number: it stays true whatever any device measures, including this
one. Ratification is the operator's merge of the separate canon pull request (Article II.3);
nothing in this record or its pull request changes the canon. With the amendment ratified,
setting concurrent_runs = "measured" is the operator's choice; this record does not make it.
Consequences¶
- The storm reads one lane at a time on this device today, as before, but now because the
declaration and the evidence say so rather than because a
pgrepwas written into a shell script. A device whose evidence supports more can run more once the operator declares"measured", and one whose runner or model changes falls back to one and asks to be measured again — as this device's own evidence would have gone stale at the 0.34.2 → 0.34.4 upgrade. - The production runner must match the calibrated one. The lanes call
/v1withoutnum_ctx, so the app'sOLLAMA_CONTEXT_LENGTHsizes every slot. Running N lanes needs the app started withOLLAMA_NUM_PARALLEL=Nand the calibrated context per slot; a runner that sizes slots differently is not the runner that was measured. - Owed: the cluster. Each node that serves a model is its own device. The chart runs one
Ollama pod with
contextLength: 32768— a window that refuses or truncates every storm turn deeper than that — and does not template a calibration Job. The documented procedure is a Job pinned to the node, inside the runner's pod network, withVIBEY_GH_SLOTS_DIRon a volume the workers share; until it is templated, a node without evidence runs one, which is safe. No cluster node was measured, and no number is claimed for one. - Owed: durable storm records. The storm keeps its lanes, and so every run's record, under
/private/tmp; a reboot erases the only record of what the model was asked. They belong somewhere that survives one. - Other devices — the self-hosted runner's host, cluster workers, adopters' machines — were not reachable from this run and carry no evidence; each runs one until it is calibrated.
Alternatives considered¶
- Keep the host-wide
pgrep. It is 8.c as written, but it is a literal, not a measurement, and it cannot say why or when to change (12.c, 8.j). - A single number for every device in the canon. Rejected: the right number depends on the device, the runner, the model and its window; a canon that named one would be false on the next machine. The canon rules the method.
- Restart the production runner per step. Rejected: a relaunch applies a staged update the operator had not asked for, and changes the runner under the measurement.
- Reconstruct the lost run from memory. Rejected (10.f): a number nobody can re-read is not evidence.