The dotbabel fleet CPU lanes give each heavy test run about 5 CPUs: k = (usable + 2) / 5 lanes, and 1 CPU is reserved when a machine has 6 or more (plugins/dotbabel/scripts/fleet-lane.sh:99-111). Nobody measured the 5. This experiment measures which lane size finishes the most test runs per hour on one real machine. The limits are no test failures from load and a responsive host.
node -e 0 probe p95 is more than 3× the idle p95.k = (usable + ⌊W/2⌋) / W. W = 5 means no change.taskset pin fixes vCPUs, not physical cores.| Id | Repo @ SHA | Command |
|---|---|---|
| DV | dotbabel @ 2bc2de8 | npx vitest run (141 files) |
| DB | dotbabel | bash plugins/dotbabel/scripts/run-bats.sh (75 files, -j 8) |
| DBJ | dotbabel | the same, with BATS_JOBS = the lane’s CPUs |
| SV | squadranks @ a4eb90a3 | npx vitest run (222 files) |
| SG | squadranks api/ |
go test ./... -race -count=1 |
| MP | moneyballer @ cbec9bc | pytest elt/tests/ -m 'not real_api' --timeout=120 -n auto -p no:cacheprovider |
cpu-lane-size/bench.sh and cpu-lane-size/analyze.mjs.
node_modules hardlinked from the main checkout.taskset, DOTBABEL_LANE, and PYTEST_XDIST_AUTO_NUM_WORKERS.node -e 0 and bash -c :, load1, and CPU PSI.| CPUs | DV | DB | DBJ | SV | SG | MP |
|---|---|---|---|---|---|---|
| 2 | 78.5 | 137.9 | 345.7 | 232.0 | 47.6 | 653.0 |
| 3 | 43.9 | 127.6 | 184.0 | 140.9 | 31.9 | 501.6 |
| 5 | 32.5 | 102.3 | 122.9 | 87.3 | 30.4 | 311.8 |
| 6 | 28.4 | 96.1 | 107.7 | 89.7 | 29.6 | 311.9 |
| 8 | 25.7 | 100.4 | 94.1 | 87.4 | 28.3 | 228.4 |
| 15 | 23.7 | 97.4 | 79.6 | 82.2 | 24.6 | 187.0 |
SG is flat after 3 CPUs, and SV and DB are flat after 5. DV and MP keep gaining up to 15 CPUs.
| Layout | Makespan s | Jobs/h | Mean turnaround s | Failures (from load) | node -e 0 p95 ms |
PSI median / p95 | Max load |
|---|---|---|---|---|---|---|---|
| unlaned | 897 | 40.1 | 473 | 1 (1) | 450 | 60.9 / 87.8 | 58.0 |
| 1×15 | 940 | 38.3 | 410 | 1 (0) | 124 | 3.6 / 24.6 | 21.6 |
| 2×7 | 949 | 38.0 | 457 | 0 | 121 | 9.2 / 21.4 | 19.5 |
| 4×4 | 962 | 37.4 | 467 | 0 | 106 | 8.4 / 45.4 | 28.8 |
| 3×5 (today) | 1000 | 36.0 | 539 | 0 | 105 | 10.5 / 35.2 | 23.4 |
| 5×3 | 1121 | 32.1 | 546 | 0 | 101 | 6.5 / 40.8 | 29.6 |
The idle node -e 0 p95 was 22 ms (59 samples).
The two failures:
TestResultSync_RunLoop_TransitionsBetweenIntervals failed. It is a timing test, and it failed at load 58. This is a failure from load.e2e-seed-residue.test.mjs failed. It is a live check against the repository, and in this layout no other job ran at the same time. This is not a failure from load.-j 8. The default (-j 8 from getconf) is faster than BATS_JOBS = the lane’s CPUs at every width up to 6: 102 s against 123 s at 5 CPUs. run-bats.sh needs no change.Under the locked rule, no layout qualifies. The responsiveness limit (3 × 22 ms = 66 ms) is below what any CPU-bound test load gives on this machine. Every lane layout had a p95 of 101–124 ms.
Recommendation (post-hoc, for the user to decide). Two lanes of 7–8 CPUs (W = 7, k = (usable + 3) / 7, which gives lanes 0-6 and 7-14 on this machine) would do this:
The throughput gain is close to the noise of 2 repetitions on a host that drifts. A 1-hour daytime run of only 3×5 against 2×7, with 3 repetitions, would confirm it before the formula changes.
cd docs/experiments/cpu-lane-size
bash bench.sh setup
BENCH_NS="2 3 5 6 8 15" BENCH_REPS_A=2 BENCH_REPS_B=2 setsid nohup bash bench.sh full > /tmp/lane-bench-logs/full.out 2>&1 &
node analyze.mjs results
bash bench.sh teardown
results/trials.jsonl has one row per trial, and results/probe.jsonl.gz has the host probe. The pilot data is in results/pilot/.
Phases A and B give all jobs at once. Sessions do not work that way: they start test runs at random times, so a short run can wait behind a long one. Test 1 measures that wait.
Method.
fleet-lane.sh, with its FIFO queue, taskset, and environment. The jobs use a private state directory, and the benchmark holds the real lanes.model-intelligence-adapter-claude.test.mjs. That test fails about half the time under load (fixed later on main in #435), and a random failure would count against one layout.Load control (added after an invalid first attempt). The first attempt, on 2026-09-27 05:49, was stopped after 2 of 9 runs. Other sessions ran Docker, compiles, and eslint outside the lanes, the load reached 45, and the same jobs ran 4–5 times slower in one layout. Its data is in results/arrivals-invalid/, and nothing here uses it. The second attempt adds these controls:
/proc/stat minus the scope’s cpu.stat usage.Run. It ran on 2026-09-28 from 21:53 to 00:13. The gate waited 11 minutes, until 2 busy squadranks sessions went quiet.
| Layout | Jobs | Mean turnaround s | p95 turnaround s | p50 wait s | p95 wait s | Short-job (DV, SG) p95 s | Failures | node -e 0 p95 ms |
|---|---|---|---|---|---|---|---|---|
| 2×7 | 33 | 374.3 | 634.9 | 206.2 | 546.9 | 614.5 | 1 | 140 |
| 3×5 (today) | 33 | 399.8 | 717.6 | 88.8 | 530.9 | 629.7 | 1 | 137 |
| 4×4 | 33 | 410.4 | 726.8 | 4.8 | 457.6 | 582.9 | 0 | 127 |
2×7 has the lowest mean turnaround in every repetition: 484, 140, and 382 s, against 511, 158, and 410 s for 3×5. With 2 lanes, jobs wait longer for a lane (p50 206 s against 89 s), but each job finishes sooner, so the total time is lower.
The failures:
e2e-seed-residue.test.mjs (live), a live check against the repository. It also failed with no other job running in phase B.forwards TERM to the command and exits with its status, a timing-sensitive signal test.A layout replaces 3×5 only if all of these are true:
node -e 0 p95 is no more than 1.5×2×7 passes all four: 6.4% lower mean turnaround, a better short-job p95, 1 failure against 1, and 140 ms against 137 ms. 4×4 does not pass. On this machine, the data supports 7 CPUs per lane: k = (usable + 3) / 7, which gives lanes 0-6 and 7-14.
The margin (6.4%) is just above the 5% limit, and only a 16-CPU machine was measured. With W = 7, a 10-CPU machine would get 1 lane of 9, and that is not measured. So a formula change needs its own decision, and it may need a floor of 2 lanes.
Reproduce: BENCH_OUT=$PWD/results/arrivals BENCH_REPS_C=3 systemd-run --user --scope --unit=lane-bench-launch-$(date +%s) -- setsid nohup bash bench.sh launch-c 23:00 &, then node analyze.mjs results/arrivals.