Host optimization¶
How the machine vibey runs on is kept healthy and fitted to its workload, with every change declared in the repository, applied under a gate, measured week after week, and reversible. It answers the operator's request of 2026-10-01: "anything you can do to improve system health over time and keep the system still fully optimized".
Everything here is declared in scripts/host_tuning.toml
(ADR-0045, fitted to the iron; ADR-0051, configuration in TOML). A setting is never changed by
a one-off command. It is declared, then applied by the tool, which journals the value it
replaced.
# Run these with a plain python3 (3.11 or later), never under `uv run`: uv holds its own
# cache while its child runs, so the uv-cache reclaim would wait on itself.
python3 scripts/host_health.py tune check # declared against actual; exit 1 on drift
python3 scripts/host_health.py tune plan # what apply would do, touching nothing
python3 scripts/host_health.py tune apply # apply what each item's gate allows
python3 scripts/host_health.py tune apply --only uv_cache
python3 scripts/host_health.py tune undo ollama_kv_cache_type
Classes and gates¶
Each item has a class, and the class decides what lets tune apply touch it
([host_tuning.gates]; sub-doctrine 12.d: a gate, never judgement).
| Class | What it is | Its gate |
|---|---|---|
| A | Cannot reach the model or a running experiment, and is reversible or regenerable | Always |
| B | Changes the model's behaviour, speed or memory | adopted = true, set in a merged pull request that cites a review-canary run which held against the baseline |
| C | The operator's call: a desktop settings pane, the queue's database, deleting a model, hardware | operator_approved names the operator's approval. Even then, a step that needs root or a settings pane is printed for a person, never performed |
The tool never restarts a service and never runs sudo. A change that needs a restart is
reported as pending-restart by tune check until the service's own start time is later
than the change.
tune check state |
Meaning |
|---|---|
in-force |
The host has what is declared |
proposed |
It differs, and the item's gate is closed: nothing is wrong yet |
drift |
It differs, and the gate is open: apply it, or the declaration is wrong |
pending-restart |
Applied, but the service started before the change |
not-received |
Set, but the model runner's argv shows the server never passed it on (ADR-0045 found this for the KV cache type on Ollama 0.34.2) |
over-limit |
A log past its declared size |
reclaimable, listed |
A cache's size, or the models on disk, for information |
absent, not-applicable, unknown |
Not on this host, not on this platform, or unreadable (with the reason) |
check exits 1 when an item whose gate is open is drift, not-received,
pending-restart or over-limit.
How the Ollama environment is set, per platform:
- macOS. The Ollama app hands its server the launchd session environment when it launches
(
launchctl setenv, as Ollama's FAQ documents). Alaunchctl setenvdoes not survive a log-out, soapplyalso writes a login agent,~/Library/LaunchAgents/dev.vibey.host-tuning.ollama-env.plist, that re-asserts exactly what the journal says is applied.undorestores the prior value, or unsets the variable if it had none, and rewrites the agent (or removes it when nothing applied remains). The app must then be quit from its menu-bar icon and opened again. - Linux.
applystages a systemd drop-in withEnvironment=lines in~/.local/state/vibey/host-tuning/staged/and prints thesudo installandsystemctl restartcommands that install it.
The journal is ~/.local/state/vibey/host-tuning/journal.jsonl. It is append-only: an undo is
a new entry that supersedes the apply.
The measured budget¶
Measured on the operator's MacBook Pro (Mac17,2, Apple M5, 10 cores, 24 GiB, APPLE SSD
AP1024Z) on 2026-10-01 and 2026-10-02, while the large-diff review experiment
(research/large-diff-review, PR #1328) ran on its Ollama. Every figure was read without
sudo: proc_pid_rusage per process (footprint, and disk bytes written since each process
started), vm_stat, sysctl vm.swapusage, top, and smartctl on disk0, sampled every
30 s for 29.5 minutes (2026-10-02 00:39–01:09 UTC). gpt-oss:20b stayed loaded at
num_ctx 65536 the whole time, serving requests on and off.
Memory¶
Demand is about twice the machine. The process footprints add up to roughly 49 GiB on a 24 GiB host, so macOS holds the rest compressed and in swap.
| Group | Footprint, mean (GiB) | Of which compressed or swapped (top CMPRS, one sample) |
|---|---|---|
Ollama (llama-server, gpt-oss:20b) |
18.5 (16.3–20.2) | 3.7 GiB |
| Docker Desktop's VM (with a 10-node Kubernetes cluster and the runner) | 8.9 | 8.1 GiB, almost all of it |
| Other processes (about 230) | 4.5 | |
| Browser | 3.7 | |
| VS Code | 3.7 | |
node (53 processes, mostly agent MCP servers) |
2.7 | |
| Agent sessions (Claude and others) | 2.0 | |
Python (including uvx MCP servers) |
1.1 | |
| Postgres | 0.7 |
- Wired: 16.8–16.9 GiB, almost all of it the model's weights and KV cache held through Metal. Wired memory cannot be paged out.
- Swap in use: 15.7 to 22.7 GiB during the window (16.9 GiB at its start, 20.9 GiB at its end). macOS grows swap files on demand; it is not a ceiling.
- The compressor held 31.9 GiB of pages in 3.5 GiB.
- The model's own footprint was 3.7 GiB compressed or swapped in one sample. Its weights are paged out between requests and paged back in for the next one, which is a likely reason generation speed varies with host state (12–34 tok/s in the evidence track). That reading is inferred, not measured here.
- Docker Desktop's VM may use 24576 MiB, the whole machine (
MemoryMiBin its settings), and runs Docker Desktop's built-in Kubernetes with 10 nodes.docker statsmeasured the kind nodes at 16 MiB to 740 MiB each, about 6.4 GiB together. The Kubernetes API server did not answerkubectl(TLS handshake timeouts), itself a sign of the thrashing. The self-hosted runner's own container used 48 MiB idle.
What drives the SSD's writes¶
Over the 29.5-minute window:
| Source | GB/h | Share |
|---|---|---|
SSD host writes (smartctl data units written) |
152.5 | 100% |
Swap-outs (vm_stat Swapouts × 16 KiB page) |
130.1 | 85% |
Every readable process's own file writes (proc_pid_rusage) |
1.8 | 1.2% |
| — of which Postgres | 1.2 | |
| Pageouts of file-backed pages | 2.7 | |
| Unattributed (kernel, file-system metadata, root's processes) | about 20 |
Swap-ins ran at 119.6 GB/h at the same time: the machine was moving about 120–130 GB/h each
way between memory and the SSD. Over the whole uptime the counters agree. top reported
2095 GiB written to disk in about 29 hours since boot, and vm_stat 96.8 million swap-outs
(1.59 TB at 16 KiB), about three quarters of it. The intervals with an active request ran
higher (median 135 GB/h) than those without (103 GB/h), but swap churned in both, because the
model stayed resident and wired in both.
Measured, and inferred. Measured: every counter above. Inferred: that a vm_stat
swap-out is 16 KiB written to the swap file. The compressor writes compressed segments, so the
swap-out bytes are an upper bound on what reached the swap file (not verified). Per-write
attribution would need fs_usage, which needs sudo, so the evidence is the counters. Even so,
nothing else readable writes more than 2 GB/h, so the earlier hypothesis holds by
elimination: the SSD's 36.5 TB in 337 power-on hours is mostly the price of memory
overcommit, and the lever is memory, not the disk.
Model reloads¶
Each starting llama-server line in the Ollama server log is one model load. There were
166 loads on 2026-10-01, at 168 distinct num_ctx values since 2026-09-27. Of 358
loads since 2026-09-28, 298 changed the context size or the model: clients ask for a
different num_ctx per request, and Ollama reloads the model for each one. The evidence track
measured 72% of CI review requests reloading, at a median of 4.8 s each. Of 631 gaps between
requests, 129 were longer than Ollama's default 5-minute keep-alive and 38 longer than 15
minutes.
Disk¶
The data volume had 458 GiB free (49% used). The regenerable caches were large. On
2026-10-01, ~/.cache/uv held 82 GB, ~/.npm/_cacache 11.9 GB, and Docker's build cache
6.8 GB (none of its 68 records active). The Ollama models take 48 GB, and the Docker VM's
disk image 20 GB. The Postgres server log was 29 MiB. Free disk is not this machine's
problem, so reclaiming caches is hygiene, not a health lever.
The plan¶
| Item | Class | Expected effect | Risk |
|---|---|---|---|
ollama_fixed_num_ctx: every lane sends one fixed num_ctx |
B | Most reloads stop (298 of 358 were context changes); seconds per request, and the model is no longer read again from the SSD | A code change in the callers (the review lane, OllamaChatClient); a fixed size large enough for the biggest diff costs KV memory on every request |
ollama_max_loaded_models = 1 |
B | A second model can never be co-resident with gpt-oss:20b |
Another model's request evicts and reloads it |
ollama_num_parallel = 1 |
B | None today (the runner already has -np 1); pins it |
Queueing only |
ollama_keep_alive = 15m |
B | About 91 fewer expiry reloads over those four days | ~16 GiB stays wired up to 10 minutes longer in each idle gap; the weekly swap figures must not rise |
ollama_flash_attention = 1 |
B | Little on its own (--flash-attn auto already); the precondition for a quantised KV cache |
Can change numerics slightly |
ollama_kv_cache_type = q8_0 |
B | About 0.65 GiB of wired memory back at 65536 tokens | Can change outputs; check must first show --cache-type-k reached the runner |
docker_kubernetes off |
C | About 6.4 GiB of demand gone, the largest single lever | Anything using the docker-desktop context stops |
docker_memory_cap = 8 GiB |
C | Bounds the VM's worst case (now the whole machine) | A CI job needing more fails; lower it only after Kubernetes is off |
postgres_shared_buffers stays 128 MB |
C (guard) | None: Postgres is 0.7 GiB and 1.2 GB/h, not a lever. Raising it would cost the model memory | None while it stays |
unused_models listed |
C | Disk only | Deleting a model is irreversible without a re-download |
uv_cache, npm_cache, docker_build_cache reclaimed |
A | Disk only | A lane's next sync or build may download or rebuild |
postgres_log, host_health_log rotated past a size |
A | Bounded disk | None: archives are kept |
Noted but not declared, because nothing measured them here:
- The operator's own applications. VS Code, the browser, agent sessions and their MCP
servers (
node,uvx) together hold about 13 GiB of footprint. Closing what is idle helps, but it is the operator's workflow, not a setting. - Docker Desktop's Model Runner and AI features (
EnableInference,EnableDockerAI) are on. Whether they hold memory was not measured. - Spotlight indexing the storm home.
mdworkerprocesses were busy, but they wrote 0.03 GB/h, so it is not a write lever, and its CPU cost was not measured.
Class A: applied on 2026-10-02¶
python3 scripts/host_health.py tune apply --only uv_cache --only npm_cache --only
docker_build_cache ran from 01:19 to 01:29 UTC. Before it ran, each target was listed with
its size and the reason it is regenerable. The journal holds each before and after.
| Item | Before | After | Outcome |
|---|---|---|---|
npm_cache (npm cache verify) |
11.90 GB | 5.73 GB | 6.16 GB freed |
docker_build_cache (docker builder prune --filter until=168h) |
6.81 GB | 6.81 GB | Nothing freed: every record was newer than the declared week, so the filter kept them all |
uv_cache (uv cache prune) |
82.02 GB | 82.02 GB | Deferred. uv waited 300 s for its cache lock and gave up (exit 2). Other uv processes were holding the cache: uvx MCP servers that live as long as their agent sessions. --force would route around that in-use check (12.d), so the tool does not use it. Run it again when no uv process is running |
The logs were under their limits (Postgres 29.3 of 64 MiB), so nothing was rotated. No class-A item reaches memory, so none of this changes swap or the SSD's writes. That needs classes B and C.
Class B: the adoption procedure¶
Class B waits until the experiments on this host finish: research/large-diff-review and
its draft PR #1328, and any live review the canary would contend with. Then, one item at a
time, in this order (the first two cannot change an output, and the output-changing ones
come last):
ollama_max_loaded_modelsandollama_num_parallelollama_fixed_num_ctx: a code change in the callers, on its own pull requestollama_keep_aliveollama_flash_attention, thenollama_kv_cache_type
For each item:
- On a branch, set the item's
adopted = trueinscripts/host_tuning.toml. - On the host, from that branch, run
python3 scripts/host_health.py tune apply --only ITEM, then restart Ollama (quit the app and open it again). - Run
tune check. The item must bein-force, and for a runner flag the detail must show the runner received it. Anot-receiveditem is undone at once: the canary would be measuring the old setting. - Run the canary:
vibey-gh review-canary run, thenvibey-gh review-canary status. - Compare with the baseline, the 2026-10-01 measurement in
docs/architecture/evidence/review-canary/ledger.jsonl(PR #1331). That run caught 18 of 25 planted defects (95% Wilson interval 52.4–85.7%) and blocked 0 of 14 controls (0.0–21.5%). The item is kept only ifstatusexits 0 (the floor holds), recall is no lower than the baseline's interval allows, and the false-positive rate is no higher. - Kept: open the pull request with
adopted = true, citing the canary's ledger line. Not kept: runtune undo ITEM, restart Ollama, and record the result in the item'sevidencewithadopted = false, so the refusal is on the record (12.d).
After adoption, the weekly host-health record judges the item against the weeks before it,
through the figures in its judged_by (next section). An item whose weeks get worse is undone
the same way.
Class C: the operator's asks¶
Each of these is the operator's decision. To approve one, set its operator_approved to the
approving pull request or date. tune apply then prints the step (it performs nothing that
needs a settings pane, root or a deletion).
- Turn off Docker Desktop's built-in Kubernetes, if nothing uses the
docker-desktopcontext: Docker Desktop > Settings > Kubernetes, then Apply & restart. This is the largest single memory lever measured (about 6.4 GiB). Do it after the experiments, because it changes the host state that tok/s depends on. - Then cap Docker Desktop's VM at 8 GiB (Settings > Resources > Advanced), and watch the runner's CI jobs for a week.
- Decide about the models not loaded in 30 days:
gemma4:26b(16.9 GB),qwen2.5-coder:14b(9.0 GB) andqwen2.5-coder:1.5b(1.0 GB). This frees disk only.qwen3:14bis qwenloop's model andgpt-oss:20bthe sovereign default. release-binaries-scratchin the storm home is 8.8 GB, is not a registered worktree, and is not a cache this tool can prove regenerable. Its owner decides.- Memory. 24 GiB is vibey's minimum and 32 GiB its recommendation. The weekly
ram_capacityandswapdrivers already track this machine against both. Size the next machine for at least 32 GiB, after the levers above have been measured. - Postgres needs nothing.
postgres_shared_buffersis a guard that keeps it from being raised.
How the weekly record judges each change¶
The weekly host-health record (scripts/host_health.py weekly) carries what was in force
(tuning.state) beside the figures that judge it, so the weeks after a change can be read
against the weeks before:
| Figure | What it shows |
|---|---|
memory.swap_used_gib, memory.swap_used_ratio |
Swap in use, and against memory |
memory.swapout_gb_per_hour |
Swap-out volume per hour since boot |
memory.swap_share_of_writes |
Swap-outs as a share of all bytes written to disk since boot |
storage.host_writes_gb_per_hour_since_boot |
Disk writes per hour since boot |
storage.bytes_written_per_week |
The SSD's own counter, week over week. It does not reset at boot, so it is the one to trend |
ollama.loads_per_day, ollama.distinct_num_ctx |
Model reloads, and the context sizes that cause them |
throughput.gen_tok_s_week_median |
Generation rate at comparable context (prompts of 1,024 to 32,768 tokens) |
memory.budget |
Memory by process group |
The since-boot rates restart at every boot, so compare them only between records made a
similar number of hours after a boot. The SSD's weekly bytes written do not reset. As wear
slows, the ssd_endurance and ssd_wear drivers move their forecast later on their own.