Host health¶
How healthy is the machine vibey and krypton run on, and how long until it needs replacing?
The tables on this page are generated. scripts/host_health.py writes them from the
committed, append-only weekly record,
docs/architecture/evidence/host-health.jsonl.
The text between the markers is never edited by hand. The prose around them is.
Where the weekly run happens, and why there¶
The probes run on the host itself, once a week. A launchd agent (macOS) or a systemd
user timer (Linux) runs them, and python scripts/host_health.py install renders that unit
from [host_health.schedule] in scripts/host_health.toml. They do not run in GitHub
Actions. The project's self-hosted runner is a Linux container on the host (vibey-gh
runner starts it with docker run), and a container cannot see the host's SSD health,
battery, thermal state or real memory pressure.
install clones the repository into a directory of its own ([host_health.publish]
clone_dir). It then writes the unit into the service manager's directory and prints the one
command that loads it. vibey never loads a unit into your session itself.
# macOS, from your vibey checkout
uv run --no-project --python 3.12 python scripts/host_health.py install
launchctl bootstrap "gui/$(id -u)" ~/Library/LaunchAgents/dev.vibey.host-health.plist
# Linux
uv run --no-project --python 3.12 python scripts/host_health.py install
systemctl --user daemon-reload && systemctl --user enable --now dev.vibey.host-health.timer
Each week the unit runs scripts/host_health.py weekly, which does three things:
- It fetches the clone and starts a fresh
automation/host-healthbranch fromdevelop. - It probes the host and appends one record to the host's own durable copy
(
[host_health] local_record). - It merges in every local record that the committed file does not hold yet, matched by run id. A week whose pull request never merged is carried by the next one, never dropped. Then it re-renders this page, checks it, commits, pushes the branch, and opens or updates one pull request.
Nothing reaches develop except through that pull request and the merge train. If
publishing fails, the week stays in the local record and the run says so.
A hosted workflow (.github/workflows/host-health.yml) reads the committed record every
week. It keeps one tracking issue open while any of these conditions holds:
- a host has not recorded for 14 days;
- an expected probe is missing from a host's last two records;
- a driver is past its threshold;
- a replacement is predicted within the warning horizon.
It closes the issue when none of them holds. vibey doctor prints the newest record's
summary, its age and the forecast.
What is measured, and how¶
Nothing needs sudo. Some probes cannot run: one would need privileges, a tool is not
installed, or the machine lacks the hardware (a desktop has no battery). Each of those is
recorded as skipped, with the reason, and no number is invented for it.
| Probe | macOS | Linux |
|---|---|---|
| Storage | smartctl --json -a disk0 when smartmontools is installed (wear, spare, data written, media errors, power-on hours). diskutil info for the SMART verdict. df for free space. |
smartctl --json -a, which usually needs root and is skipped without it. df. |
| Battery | ioreg -rn AppleSmartBattery (cycles, full-charge against design capacity, design cycle count). system_profiler SPPowerDataType (condition). |
/sys/class/power_supply/BAT* |
| Thermal | pmset -g therm (CPU speed limit, warning level) |
Thermal zones, cpufreq caps against the hardware maximum, throttle counters |
| Memory | sysctl hw.memsize vm.swapusage, memory_pressure -Q, vm_stat rates since boot |
/proc/meminfo, /proc/pressure/memory, /proc/vmstat |
| Reliability | Panic, reset, shutdown-stall and jetsam reports, by file name and date only. The reports are never opened. Time since boot. | pstore crash records, /proc/uptime |
| Throughput | The sovereign model's generation rate per week, mined passively from the Ollama server log and its rotations. No request is made. | Same |
| Microbenchmark | SHA-256 and a flushed sequential write. They run only when the minimum-specs idle gate says the host is quiet. | Same |
| Platform | The OS and its vendor support end, and the hardware's support status. Declared in the TOML with a source and a last-verified date. | Same |
| Capacity | vibey's own memory and disk requirements, copied in from the minimum-specs record with their dates | Same |
| Model loads | How often the sovereign model was loaded, per day, and at how many distinct context sizes, over the trailing week, mined passively from the Ollama server log | Same |
| Budget | Memory by process group from top (each process's footprint, and how much of it is compressed or swapped out). All bytes written to disk since boot, and the share of them that swap-outs account for (vm_stat) |
ps resident sizes, /proc/diskstats, /proc/vmstat |
| Tuning | Which items of the declared host tuning (scripts/host_tuning.toml) were in force, so the weeks after a change can be compared with the weeks before it (Host optimization) |
Same |
Privacy (SD-01 §1): the host is named by a fingerprint, a truncated SHA-256 over its hardware facts and platform identifier. The identifier, serial numbers and the host name never reach the record.
How the forecast is made¶
Each driver is one reason the machine might need replacing. [host_health.drivers]
declares each one with its threshold and where that threshold comes from.
- Trend drivers are SSD wear, battery capacity and cycles, generation rate, and free
disk. Each is fitted over the weekly history with the Theil-Sen estimator, which an odd
week cannot drag, and with Sen's rank-based interval on the slope
(
scripts/host_health_forecast.py). The projected date is where the line reaches the threshold. The interval comes from the slope's bounds. When the data cannot rule out "never", the interval says not bounded. With fewer than four weekly points, a trend driver says insufficient history and projects nothing. - Level drivers are memory against vibey's minimum, sustained swap, kernel panics and thermal limits. One is past its threshold when the last confirm records all are.
- Dated drivers are the OS leaving vendor support and the hardware becoming obsolete. They take the date the vendor publishes. Where the vendor publishes none, they say so.
The machine's replacement date is the earliest driver's date, and the page names the
driver that binds. The interval runs from the earliest bound across the drivers to the
smallest finite latest bound. vibey's own requirements are read from the minimum-specs
machinery, never restated here. The minimum and recommended memory and free disk come from
its record, and the minimum generation rate from scripts/minimum_specs.toml.
The hosts¶
Host sha256:ec5e81b92e3a09da (Mac17,2 · Apple M5 · 24 GiB · macOS 26.6.2 (25G83)): 2 weekly record(s), 2026-10-01 to 2026-10-02.
The forecast¶
Host sha256:ec5e81b92e3a09da (Mac17,2 · Apple M5 · 24 GiB · macOS 26.6.2 (25G83)): 2 weekly record(s), 2026-10-01 to 2026-10-02.
No driver projects a replacement date yet. Trend drivers need 4 weekly points; dated drivers need a vendor date. Each driver's reason is in the table below. As of 2026-10-02.
Hypothesis, not a conclusion: 16.38 GiB of swap in use on 24 GiB of memory, with swap-outs of about 62.3 GB/h since boot, may be a large part of the SSD's 108.5 GB per power-on hour of writes. The record cannot attribute writes to swap; per-process disk I/O would test it.
| Driver | Latest | Threshold | State | Date (interval) | Points |
|---|---|---|---|---|---|
| SSD data written vs rated endurance (TBW) | 36.55 TB | — | unknown | threshold unavailable: storage.rated_endurance_tb: Apple publishes no endurance (TBW) rating for its internal SSDs; a search on 2026-10-01 found only third-party estimates and forum figures, which are not a rating (verified 2026-10-01); the forecast rests on SSD wear (NVMe percentage used) instead | 1 |
| SSD wear (NVMe percentage used) | 1 % | 100 (declared in scripts/host_health.toml) | insufficient-history | 1 of the 4 weekly points a trend needs | 1 |
| SSD available spare at its threshold | 1 percentage points | 0 (declared in scripts/host_health.toml) | within | — | 1 |
| Battery full-charge capacity vs design | 1.0162 ratio | 0.8 (declared in scripts/host_health.toml) | insufficient-history | 2 of the 4 weekly points a trend needs | 2 |
| Battery cycles vs the design cycle count | 18 count | 1000 (this host's battery.design_cycle_count (2026-10-02)) | insufficient-history | 2 of the 4 weekly points a trend needs | 2 |
| Sovereign-model generation rate | 25.51 tokens/s | 10 (scripts/minimum_specs.toml minimum_specs.assumptions.minimum_gen_tok_s) | insufficient-history | 2 of the 4 weekly points a trend needs | 2 |
| Free disk vs vibey's minimum | 492.3 GB | 20 (minimum-specs record disk.minimum_gb (stale, 2026-10-02)) | insufficient-history | 2 of the 4 weekly points a trend needs | 2 |
| Memory vs vibey's minimum | 24 GB | 24 (minimum-specs record ram.minimum_gb (stale, 2026-09-30)) | within | — | 2 |
| Swap in use vs memory (sustained) | 0.682 ratio | 0.5 (declared in scripts/host_health.toml) | unconfirmed | past the threshold in 2 of the 3 records that confirm it | 2 |
| Kernel panics in the trailing window | 1 count | 3 (declared in scripts/host_health.toml) | within | — | 2 |
| Thermal CPU speed limit (sustained) | 100 % | 99 (declared in scripts/host_health.toml) | within | — | 2 |
| Operating system out of vendor support | — | — | unknown | Apple publishes no end-of-support date for a macOS release (its security-releases page lists updates, not end dates) (source https://support.apple.com/en-us/100100, verified 2026-10-01) | 0 |
| Hardware obsolete (vendor policy) | — | — | unknown | Apple has not stopped selling it as far as this table records; set last_sold when it does (policy: obsolete 7 years after the last sale; source https://support.apple.com/en-us/102772, verified 2026-10-01) | 0 |
The newest record¶
Host sha256:ec5e81b92e3a09da (Mac17,2 · Apple M5 · 24 GiB · macOS 26.6.2 (25G83)): 2 weekly record(s), 2026-10-01 to 2026-10-02.
| Figure | Value | Status | How |
|---|---|---|---|
| Battery full-charge capacity / design | 1.0162 ratio | measured | ioreg -rn AppleSmartBattery: AppleRawMaxCapacity / DesignCapacity |
| Battery condition | Normal | measured | system_profiler SPPowerDataType |
| Battery cycle count | 18 count | measured | ioreg -rn AppleSmartBattery: CycleCount |
| Battery design cycle count | 1000 count | measured | ioreg -rn AppleSmartBattery: DesignCycleCount9C |
| vibey's minimum free disk | 20 GB | declared | read docs/architecture/evidence/minimum-specs.json: disk.minimum_gb (source: docs/architecture/evidence/minimum-specs.json) |
| vibey's recommended free disk | 50 GB | declared | read docs/architecture/evidence/minimum-specs.json: disk.recommended_gb (source: docs/architecture/evidence/minimum-specs.json) |
| vibey's minimum memory | 24 GB | declared | read docs/architecture/evidence/minimum-specs.json: ram.minimum_gb (source: docs/architecture/evidence/minimum-specs.json) |
| vibey's recommended memory | 32 GB | declared | read docs/architecture/evidence/minimum-specs.json: ram.recommended_gb (source: docs/architecture/evidence/minimum-specs.json) |
| Memory compressions per hour | 64,502,520 pages/h | measured | vm_stat counter / hours since kern.boottime |
| System-wide memory free | 11 % | measured | memory_pressure -Q |
| Installed memory, GB as sold | 24 GB | measured | sysctl -n hw.memsize |
| Installed memory | 24 GiB | measured | sysctl -n hw.memsize |
| Swap size | 17 GiB | measured | sysctl vm.swapusage |
| Swap in use | 16.38 GiB | measured | sysctl vm.swapusage |
| Swap in use / installed memory | 0.682 ratio | measured | sysctl vm.swapusage |
| Swap-out volume per hour | 62.3 GB/h | measured | vm_stat Swapouts x page size / hours since kern.boottime |
| Swap-outs per hour | 3,802,632 pages/h | measured | vm_stat counter / hours since kern.boottime |
| CPU: SHA-256 throughput, one core | 2965.2 MiB/s | measured | best of 3: SHA-256 over 256 MiB in memory |
| Disk: fsynced sequential write | 1763.6 MiB/s | measured | best of 3: 128 MiB written and flushed to the drive (F_FULLFSYNC on macOS, fsync elsewhere), scratch file removed |
| Hardware obsolete under the vendor's policy | — | skipped | skipped: Apple has not stopped selling it as far as this table records; set last_sold when it does (policy: obsolete 7 years after the last sale; source https://support.apple.com/en-us/102772, verified 2026-10-01) |
| Operating system | macOS 26.6.2 (25G83) | measured | sw_vers |
| OS vendor support ends | — | skipped | skipped: Apple publishes no end-of-support date for a macOS release (its security-releases page lists updates, not end dates) (source https://support.apple.com/en-us/100100, verified 2026-10-01) |
| Crash, panic and shutdown events found | 9 event(s) | measured | list report names in /Library/Logs/DiagnosticReports, /Library/Logs/DiagnosticReports/Retired, ~/Library/Logs/DiagnosticReports |
| jetsam events in the last 30 days | 1 count | measured | list report names in /Library/Logs/DiagnosticReports, /Library/Logs/DiagnosticReports/Retired, ~/Library/Logs/DiagnosticReports |
| kernel panic events in the last 30 days | 1 count | measured | list report names in /Library/Logs/DiagnosticReports, /Library/Logs/DiagnosticReports/Retired, ~/Library/Logs/DiagnosticReports |
| reset events in the last 30 days | 1 count | measured | list report names in /Library/Logs/DiagnosticReports, /Library/Logs/DiagnosticReports/Retired, ~/Library/Logs/DiagnosticReports |
| shutdown stall events in the last 30 days | 6 count | measured | list report names in /Library/Logs/DiagnosticReports, /Library/Logs/DiagnosticReports/Retired, ~/Library/Logs/DiagnosticReports |
| Days since boot | 1.01 d | measured | sysctl -n kern.boottime |
| SSD available spare | 100 % | measured | smartctl --json=c -a disk0 |
| SSD data written, in bytes | 36,554,771,968,000 bytes | measured | smartctl --json=c -a disk0 |
| SSD data written per week | — | skipped | skipped: needs an earlier record of this host with storage.bytes_written |
| SSD critical warning bits | 0 bits | measured | smartctl --json=c -a disk0 |
| SSD error-log entries | — | skipped | skipped: smartctl could not read the error-information log page: Read 1 entries from Error Information Log failed: GetLogPage failed: system=0x38, sub=0x0, code=745 |
| Free disk on the data volume | 492.3 GB | measured | df -Pk /System/Volumes/Data |
| SSD media and data integrity errors | 0 count | measured | smartctl --json=c -a disk0 |
| SSD model | APPLE SSD AP1024Z | measured | smartctl --json=c -a disk0 |
| SSD wear: NVMe percentage used | 1 % | measured | smartctl --json=c -a disk0 |
| SSD power-on hours | 337 h | measured | smartctl --json=c -a disk0 |
| SSD rated endurance (TBW) | — | skipped | skipped: Apple publishes no endurance (TBW) rating for its internal SSDs; a search on 2026-10-01 found only third-party estimates and forum figures, which are not a rating (verified 2026-10-01) |
| Data volume size | 994.6 GB | measured | df -Pk /System/Volumes/Data |
| SSD SMART overall status | passed | measured | smartctl --json=c -a disk0 |
| SSD available spare above its threshold | 1 percentage points | measured | smartctl --json=c -a disk0 |
| SSD data written | 36.55 TB | measured | smartctl --json=c -a disk0 |
| SSD unsafe shutdowns | 10 count | measured | smartctl --json=c -a disk0 |
| SSD writes per power-on hour (lifetime average) | 108.5 GB/h | measured | smartctl --json=c -a disk0 |
| CPU speed limit | 100 % | measured | pmset -g therm |
| Thermal warning level | 0 level | measured | pmset -g therm |
| Generation rate, newest week's median | 25.51 tokens/s | measured | mine ~/.ollama/logs/server*.log (passive; no request is made) |
| Generation rate by week | 2 week(s) of medians | measured | mine ~/.ollama/logs/server*.log (passive; no request is made) |