Benchmarking contract¶
Benchmark results are build artefacts, not permanent source-code claims. Raw
outputs belong under /tmp or another external artefact directory and must not
be committed.
Lean-native comparison harness¶
lean_native_compare.py compares exact commits in fresh interpreters using
complete ABBA cycles and same-code A/A controls. Its self-subscribed native
client exercises both MQTT directions, QoS 0/1/2, memory/SQLite, iterator/callback,
and bursts of 1/2/8 or a long progressive lot. A phase requires ordered,
payload-verified delivery, publication receipts and final inbound handshakes.
CPU covers the combined client process; the broker runs separately.
Per-message latency starts when the producer constructs the element immediately before admission, so it includes prefetch queue residence. A long batch's overall elapsed time also reports work waiting before an element is read. Timing runs exclude tracemalloc; a separate phase measures Python peak allocations. RSS is diagnostic. Topic, payload, flow and queue limits are identical across arms.
This explicitly authorized comparison campaign has no performance acceptance threshold. Functional correctness and resource bounds remain mandatory. Record all ratios and A/A noise, and label results diagnostic if the runner is ineligible. These measurements are not release qualification or public cross-client evidence.
Record the broker's TCP settings. Small QoS 0 bursts can alternate between fast
delivery and roughly 40-ms TCP stalls, including with identical client source;
a short pilot may then choose an unsuitable message count. Inspect raw A/A
durations as well as aggregate ratios. A separate broker with Mosquitto's
set_tcp_nodelay true is a useful controlled condition, but it changes the
stimulus. Preserve the original attempt and report the two conditions separately
instead of presenting the new result as an unchanged-workload rerun.
lean_native_paced.py supplements those saturated lots with fixed offered load.
It freezes rates at 50/75/90% of the reference long-lot A/A median before comparing
the two source revisions. This calibration is a fraction of that measured
workload, not a separately established sustainable open-loop capacity. A separate
process emits scheduled tokens; the client loop does not sleep to pace itself.
The harness retains planned arrival, the pacer's pre-send timestamp, publication
call, admission return, delivery, and observed receipt-completion clocks in hashed
companion files. Include those files with the result JSON. Admission return is an API
observation, not the internal commit instant. Report deliveries before return
explicitly; residual latency clamps those observations to zero. Scheduled-to-call
and scheduled-to-delivery latency retain producer backlog and pacer lateness.
The pacer's token socket is blocking: a full buffer may delay subsequent
emissions. The offered rate defines the fixed schedule, not guaranteed actual
emission under overload. Inspect pre-send lateness alongside planned-to-call
delay rather than assuming that moving the clock to another process removes
all backpressure from the load generator.
lean_native_diagnostics.py separates publication, ingress with iterator or
callback delivery, callback invocation alone, and routed dispatch. Its existing
packet-aware test transport isolates local orchestration; publication includes
the transport's broker emulation and is not a network capacity result. Fresh
processes run A/A and ABBA trials. Optional profiles run in separate phases and
count Python calls and resumptions, not operating-system context switches or
system calls. Evaluate routing candidates separately against their immediate
predecessor and the measured same-code noise.
What each benchmark answers¶
hotpath_profile.pycounts calls, primitive calls, and allocations. These exact measurements are the first place to look for redundant work.paired_regression.pycompares isolated implementation paths in fresh processes and alternating order.paired_network.pyrecords advisory closed-loop QoS 1 capacity and PUBACK latency. It is not a release gate because its raw-arm A/A stability is not reliable enough to use its legacy CV rule as a general release decision.network_release_gate.pywraps a calibratedpaired_network.pysubset in mandatory baseline/candidate A/A controls and confidence-bound ABBA-cycle evaluation. It is a deep manual release/audit gate; seenetwork-release-gate.md.paired_open_loop.pymeasures completion and loop lag at calibrated or fixed absolute load. It can sweep outbound windows while callingAsyncClientdirectly; no cross-client adapter participates in the measurement.open_loop_release_gate.pyanchors fractional load to baseline capacity and uses a two-stage loop-lag decision. The initial ABBA screen must exceed the relative threshold with relative and additive 95% lower bounds above the no-effect boundary before it consumes one of the bounded same-code A/A confirmation slots. Final rejection still requires the additive increase to exceed the measured same-code noise envelope.
The open-loop loop-lag rule is a targeted regression detector, not an
equivalence proof. Its initial screen uses necessary conditions from the final
pre-existing verdict: a point estimate above 1.05 whose relative or additive
95% interval still crosses the no-effect boundary remains diagnostic. It cannot
fail the release or consume scarce same-code controls by itself. Throughput,
runner eligibility, exact completion, and the deeper controlled network gate
remain separate evidence.
- paired_writer_capacity.py protects the native publish_nowait closed-loop
writer regime for QoS 0/1. It yields once per application outstanding window
and yields/retries on synchronous backpressure, matching the scheduling shape
used by the external native capacity harness. Its primary metric is the
candidate/base completed-rate ratio, not an absolute cross-machine rate.
- paired_writer_waiter_contention.py isolates WritePump.enqueue() waiters
against a tight writer message window. It is the contention harness for the
targeted-wake experiment; default concurrency is 1/4/16 (64/256 are opt-in).
It does not replace paired_writer_capacity.py.
- paired_qos1_rtt.py measures a synchronous message handler's QoS 1 reply
from immediately before publish_nowait() to transport exposure. The
publish_nowait_call_to_transport interval includes admission and writer
scheduling. Historical callback-return-to-transport results use a different
starting point and require separate controls.
- application_stress.py exercises callbacks, iterators, backpressure, memory,
and SQLite persistence.
- memory_profile.py enforces versioned tracemalloc and logical-counter limits.
Comparisons must exercise equivalent public completion semantics. If a library
cannot express a contract, report N/A; do not manufacture equivalence with an
extra barrier that changes only one side.
Valid local evidence¶
runner_probe.py records CPU affinity, model, governor, load, temperature,
Python, and broker metadata. A strict performance run requires --enforce. An
ineligible host produces no release evidence, even if its ratio looks good.
Paired measurements use fresh interpreters and alternate base/candidate order in
complete ABBA cycles. Read the pair distribution as well as the aggregate.
Benchmark-specific validity rules still apply: the legacy strict harnesses that
explicitly gate raw arms retain their CV limits and neutral-control budgets.
network_release_gate.py is deliberately different: raw arm CV is diagnostic,
and validity comes from bounded same-code bias plus 95% equivalence intervals on
complete ABBA-cycle ratios. A benchmark that cannot satisfy its own declared A/A
control is diagnostic, regardless of whether it exposes a strict exit mode.
Hosted GitHub runners are useful for functional coverage and advisory numbers. They are not authoritative for latency or small throughput changes.
Closed-loop writer-capacity regression gate¶
The eager-write optimisation has two intentionally different regimes: paced
traffic should keep the zero-hop first write, while a synchronous producer burst
must seed once and then let the writer task batch the rest. Open-loop latency
cells cannot prove the latter. paired_writer_capacity.py therefore runs the
native producer on its own event loop with the same application discipline as
the external capacity harness: MQTT 3.1.1, 256-byte payloads, protocol inflight
20, application outstanding 64, and one cooperative yield per 64 successful
submissions. A FlowControlError yields once and retries the same unit of work.
QoS 0 counts successful native admission, then drains the writer outside the
timed interval; QoS 1 counts actual publish completion.
For a writer-regime change, first validate the harness as A/A on one source tree, then compare the approved baseline and candidate on the same eligible host. Record the exact baseline commit in the release manifest. The strict sequence is:
python benchmarks/runner_probe.py \
--output /tmp/mqttium-runner.json --enforce
# Harness control: BASELINE is a checkout/worktree of the recorded baseline.
python benchmarks/paired_writer_capacity.py \
--base-root "$BASELINE" --candidate-root "$BASELINE" \
--protocol 311 --qos-values 0,1 --payload-bytes 256 \
--inflight 20 --outstanding 64 --max-queued 200 --repeat 8 \
--policy strict --preflight-report /tmp/mqttium-runner.json \
--output /tmp/mqttium-writer-capacity-aa.json
python benchmarks/paired_writer_capacity.py \
--base-root "$BASELINE" --candidate-root . \
--protocol 311 --qos-values 0,1 --payload-bytes 256 \
--inflight 20 --outstanding 64 --max-queued 200 --repeat 8 \
--policy strict --preflight-report /tmp/mqttium-runner.json \
--output /tmp/mqttium-writer-capacity-ab.json
The A/A median completed-rate ratio must stay within 2% and baseline CV at or below 5%. The A/B candidate must retain at least 95% of the recorded baseline completed rate for both QoS 0 and QoS 1. This is a regression floor, not a cross-client performance claim. The paced open-loop acceptance cells at 2,500 and 7,500 messages/s remain separate evidence and must still be retained; recovering capacity by simply disabling eager writes would fail that side of the contract.
The GitHub Paired Regression workflow runs a shorter version of this cell as
advisory functional/diagnostic coverage. Its hosted-runner numbers are not a
substitute for the strict eligible-host A/A and A/B sequence above.
Keeping the harness out of the result¶
A benchmark can become its own bottleneck. MQTTium's network harness follows five rules to prevent that:
- The process running
AsyncClientcontains no subscriber reader thread.mosquitto_suboutput is timestamped in a separate observer process, so payload parsing cannot take the publisher's GIL. - The observer subscribes at QoS 0. It verifies delivery and ordering without adding a second PUBACK stream to a benchmark whose target is publisher PUBACK capacity.
- Every closed-loop cell calibrates its message count toward a configured target duration and records the actual duration. The target is not presented as a guarantee because fresh-process and broker rates can change after calibration.
- Open-loop calibration runs the same subscriber, completion tracking, and telemetry path as the paced sample. The only difference is pacing itself.
- Each publication's start timestamp stays paired with its own receipt. Receipt identity distinguishes publications even when a packet identifier is reused before the observer runs. Tracking uses neither publication callbacks nor a single timestamp slot per MID.
When a CPU is selected, the publisher worker is pinned only after the subscriber and observer have started. The observer therefore does not inherit the publisher's single-CPU affinity.
Harness changes require an A/A control on one source tree before they can support an A/B claim. Record throughput, CPU time, delivery latency, CV, and the A/A ratio. If instrumentation changes the scenario materially, retain both the old and new controls and explain why the new result is more representative.
Latency semantics¶
Network latency starts immediately before the application calls publish().
It includes local admission and queue residence as well as broker and transport
time. PUBACK proves broker acceptance; independent subscriber completion proves
delivery to the observer.
Receipt completion is observed by an awaiting task, so its timestamp includes that task's scheduling delay. Its own A/A cell must pass before it supports a latency comparison. Exact call/allocation profiles complement these timings when isolating changes to publication orchestration.
Larger inflight windows can improve throughput through batching while increasing latency. Sweep the window before calling a high-window latency change a protocol regression. Open-loop measurements are preferable when the question is latency at a known fraction of capacity.
Use fixed absolute rates when the question names a concrete workload rather than a fraction of the implementation's capacity. For example, an independent MQTT 3.1.1 diagnostic at 5,000 and 10,000 messages/s can use:
python benchmarks/runner_probe.py \
--output /tmp/mqttium-runner.json --enforce
python benchmarks/paired_open_loop.py \
--base-root . --candidate-root . \
--protocols 311 --payloads 64,4096 \
--completions receipt --windows 8,32,64,128 \
--target-rates 5000,10000 --repeat 12 \
--policy strict --preflight-report /tmp/mqttium-runner.json \
--output /tmp/mqttium-open-loop-aa.json
When --target-rates is present, fixed-rate points run by themselves unless
--fractions is also supplied explicitly. With neither option, the retained
0.50,0.75,0.90,1.00 capacity-fraction sweep remains the default. Passing the
same source root on both sides is an A/A control and enforces the neutral
completed-rate ratio within 2% in addition to the normal CV checks.
A practical optimisation order¶
- Count exact calls and allocations per operation.
- Isolate the suspicious operation with a microbenchmark.
- Confirm the change in paired A/B runs on an eligible, idle host.
- Include a neutral control and verify non-targeted paths.
- Use CI last, for reproducibility and portability rather than precise timing.
This order distinguishes demonstrated redundant work from attractive but unmeasured allocation theories. Several retained MQTTium optimisations removed a duplicate call or conversion. Several allocation-motivated rewrites were reverted because the replacement operation cost more.
Acceptance thresholds¶
A micro optimisation is retained when it removes demonstrated work, improves by at least 2%, and favours the candidate in at least 8 of 11 pairs. A network optimisation requires a reproducible gain of at least 5% at two load points.
Unless a benchmark-specific contract explicitly replaces them, the generic acceptance checks are:
- baseline CV at most 5%;
- neutral control within 2%;
- non-targeted throughput not down by more than 3%;
- loop lag not up by more than 5%;
- memory limits and public semantics unchanged;
- added complexity justified by an explicit, measured trade-off.
network_release_gate.py is a no-regression release gate rather than the generic
optimisation-acceptance test above. Its calibrated same-code equivalence and A/B
confidence-bound contract, including diagnostic-only raw CV, is defined in
network-release-gate.md.
The loop-lag ratio is only readable when both arms sit in the same pacing
regime, and it silently penalises the faster one when they do not.
loop_lag_p95 measures how late the paced publisher wakes relative to its own
deadline, so it is bimodal: while the publisher still has slack it genuinely
sleeps between messages and the value sits on a plateau set by timer wake-up
granularity (~1 ms on the reference host); once its per-iteration work exceeds
the pacing interval it stops sleeping and the value collapses by roughly 5× to
something that measures loop congestion. A candidate that changes publisher
per-iteration cost moves the rate at which that transition happens, so at rates
inside the transition band the ratio compares a plateau value against a
collapsed one and reports a large regression for what is in fact an
improvement. Baseline CV inflates in the same band, because the slower arm
flips between modes from sample to sample.
Before trusting a lag verdict, compare the two arms' absolute
loop_lag_p95: values near the plateau mean the publisher is still sleeping and
the number is a timer artifact, not congestion. Choose load points where both
arms are on the same side of the transition. A worked example, including the
sweep that identified the band, is in
the historical native writer-hop report.
Memory thresholds¶
check_memory_thresholds.py validates memory_profile.py immediately after the
profile. memory_thresholds.json contains reviewed limits, not generated
measurements. It gates tracemalloc peaks and exact logical counters; RSS remains
diagnostic because allocators and kernels retain pages differently.
Reference measurements live in the historical memory-results report. Raising a threshold is a reviewable source change and requires a documented reason.