dotbabel

Experiment: CPU lane size — 2026-09-27

The dotbabel fleet CPU lanes give each heavy test run about 5 CPUs: k = (usable + 2) / 5 lanes, and 1 CPU is reserved when a machine has 6 or more (plugins/dotbabel/scripts/fleet-lane.sh:99-111). Nobody measured the 5. This experiment measures which lane size finishes the most test runs per hour on one real machine. The limits are no test failures from load and a responsive host.

Decision rule (locked before running)

  1. Drop a layout that has a failure from load, meaning a failure that its solo run did not have. Also drop a layout whose node -e 0 probe p95 is more than 3× the idle p95.
  2. From the remaining layouts, pick the highest phase B jobs/hour. If two are within 5%, pick the one with fewer, larger lanes.
  3. Convert the winner to CPUs per lane W, as k = (usable + ⌊W/2⌋) / W. W = 5 means no change.

Environment

Id Repo @ SHA Command
DV dotbabel @ 2bc2de8 npx vitest run (141 files)
DB dotbabel bash plugins/dotbabel/scripts/run-bats.sh (75 files, -j 8)
DBJ dotbabel the same, with BATS_JOBS = the lane’s CPUs
SV squadranks @ a4eb90a3 npx vitest run (222 files)
SG squadranks api/ go test ./... -race -count=1
MP moneyballer @ cbec9bc pytest elt/tests/ -m 'not real_api' --timeout=120 -n auto -p no:cacheprovider

Phase A: one suite alone (median wall time in seconds, 2 runs)

CPUs DV DB DBJ SV SG MP
2 78.5 137.9 345.7 232.0 47.6 653.0
3 43.9 127.6 184.0 140.9 31.9 501.6
5 32.5 102.3 122.9 87.3 30.4 311.8
6 28.4 96.1 107.7 89.7 29.6 311.9
8 25.7 100.4 94.1 87.4 28.3 228.4
15 23.7 97.4 79.6 82.2 24.6 187.0

SG is flat after 3 CPUs, and SV and DB are flat after 5. DV and MP keep gaining up to 15 CPUs.

Phase B: 10 jobs (2 of each suite) through each layout, 2 runs

Layout Makespan s Jobs/h Mean turnaround s Failures (from load) node -e 0 p95 ms PSI median / p95 Max load
unlaned 897 40.1 473 1 (1) 450 60.9 / 87.8 58.0
1×15 940 38.3 410 1 (0) 124 3.6 / 24.6 21.6
2×7 949 38.0 457 0 121 9.2 / 21.4 19.5
4×4 962 37.4 467 0 106 8.4 / 45.4 28.8
3×5 (today) 1000 36.0 539 0 105 10.5 / 35.2 23.4
5×3 1121 32.1 546 0 101 6.5 / 40.8 29.6

The idle node -e 0 p95 was 22 ms (59 samples).

The two failures:

Findings

  1. Lanes are worth their cost. Without lanes, 5 suites at once finish 5% more jobs per hour. But the host is 4× slower to start a process (450 ms p95), CPU pressure has a median of 61%, the load goes to 58, and a timing test fails. With any lane layout, the p95 stays at 101–124 ms, and no test fails from load.
  2. Fewer, larger lanes do slightly better than today’s 3×5. 1×15, 2×7, and 4×4 are within 2.5% of each other. Today’s 3×5 is 6% slower than the best, and 5×3 is 16% slower. Responsiveness is the same for every lane layout.
  3. Lanes do not isolate runs completely. In the pilot, two DV runs on separate 5-CPU lanes took 33 s each, against 24 s for one run alone. They share memory bandwidth and the Windows scheduler on the hybrid CPU.
  4. Keep bats at -j 8. The default (-j 8 from getconf) is faster than BATS_JOBS = the lane’s CPUs at every width up to 6: 102 s against 123 s at 5 CPUs. run-bats.sh needs no change.
  5. The host was slower overnight for multi-core work. DV at 15 CPUs took 23–24 s, against 10.8 s in the daytime pilot. The single-thread drift probe changed only 8% (median 1101 ms), and 11 of 192 trials were more than 15% off (max 2.8 s). The comparisons stay fair because the layouts were interleaved, but absolute times are high.

Decision

Under the locked rule, no layout qualifies. The responsiveness limit (3 × 22 ms = 66 ms) is below what any CPU-bound test load gives on this machine. Every lane layout had a p95 of 101–124 ms.

Recommendation (post-hoc, for the user to decide). Two lanes of 7–8 CPUs (W = 7, k = (usable + 3) / 7, which gives lanes 0-6 and 7-14 on this machine) would do this:

The throughput gain is close to the noise of 2 repetitions on a host that drifts. A 1-hour daytime run of only 3×5 against 2×7, with 3 repetitions, would confirm it before the formula changes.

Reproduce

cd docs/experiments/cpu-lane-size
bash bench.sh setup
BENCH_NS="2 3 5 6 8 15" BENCH_REPS_A=2 BENCH_REPS_B=2 setsid nohup bash bench.sh full > /tmp/lane-bench-logs/full.out 2>&1 &
node analyze.mjs results
bash bench.sh teardown

results/trials.jsonl has one row per trial, and results/probe.jsonl.gz has the host probe. The pilot data is in results/pilot/.

Test 1: jobs that arrive over time (2026-09-28)

Phases A and B give all jobs at once. Sessions do not work that way: they start test runs at random times, so a short run can wait behind a long one. Test 1 measures that wait.

Method.

Load control (added after an invalid first attempt). The first attempt, on 2026-09-27 05:49, was stopped after 2 of 9 runs. Other sessions ran Docker, compiles, and eslint outside the lanes, the load reached 45, and the same jobs ran 4–5 times slower in one layout. Its data is in results/arrivals-invalid/, and nothing here uses it. The second attempt adds these controls:

Run. It ran on 2026-09-28 from 21:53 to 00:13. The gate waited 11 minutes, until 2 busy squadranks sessions went quiet.

Layout Jobs Mean turnaround s p95 turnaround s p50 wait s p95 wait s Short-job (DV, SG) p95 s Failures node -e 0 p95 ms
2×7 33 374.3 634.9 206.2 546.9 614.5 1 140
3×5 (today) 33 399.8 717.6 88.8 530.9 629.7 1 137
4×4 33 410.4 726.8 4.8 457.6 582.9 0 127

2×7 has the lowest mean turnaround in every repetition: 484, 140, and 382 s, against 511, 158, and 410 s for 3×5. With 2 lanes, jobs wait longer for a lane (p50 206 s against 89 s), but each job finishes sooner, so the total time is lower.

The failures:

Test 1 decision (rule locked before the run)

A layout replaces 3×5 only if all of these are true:

2×7 passes all four: 6.4% lower mean turnaround, a better short-job p95, 1 failure against 1, and 140 ms against 137 ms. 4×4 does not pass. On this machine, the data supports 7 CPUs per lane: k = (usable + 3) / 7, which gives lanes 0-6 and 7-14.

The margin (6.4%) is just above the 5% limit, and only a 16-CPU machine was measured. With W = 7, a 10-CPU machine would get 1 lane of 9, and that is not measured. So a formula change needs its own decision, and it may need a floor of 2 lanes.

Reproduce: BENCH_OUT=$PWD/results/arrivals BENCH_REPS_C=3 systemd-run --user --scope --unit=lane-bench-launch-$(date +%s) -- setsid nohup bash bench.sh launch-c 23:00 &, then node analyze.mjs results/arrivals.