Serving load testing¶
The approved specifications are implemented by a disposable Locust harness. Each invocation prepares synthetic sources, publishes four Recommendation Snapshots, verifies the public Serving API, measures traffic, writes aggregate reports, and removes its Docker resources. The harness measures ordinary serving; preparation is outside the measured interval.
Run a profile¶
Start from make setup with sibling ../telemetry and ../serving-limits checkouts. Docker Compose
must provide a Linux VM with at least 8 CPUs, 12 GiB memory, and 20 GiB free disk; the host also needs
20 GiB free disk. Preflight fails when these prerequisites are missing. Runs are serialized across
local worktrees and in CI.
make load-test
make load-test-capacity LOAD_PROFILE=stepped
make load-test-capacity LOAD_PROFILE=sustained
make load-test-capacity LOAD_PROFILE=stepped LOAD_VARIANT=hot-set
make load-test-capacity LOAD_PROFILE=stepped LOAD_VARIANT=limit-100
make load-test-ui LOAD_PROFILE=stepped
| Profile | Users | Pacing | Warm-up after full population | Measured duration |
|---|---|---|---|---|
| Reference | 100 | 2.5 requests/s per user, nominal 250 total | 60 s | 300 s |
| Stepped | 20, 50, 100, 200 | No intentional think time | 30 s at each step | 120 s at each step |
| Sustained | 100 | Same pacing as reference | 60 s | 1,800 s, six 300 s windows |
Pacing waits for each response and cannot guarantee an independent arrival rate. Startup phases spread the 100 users across the 400 ms cycle; a slow response delays that user's next request. Measured valid throughput remains a separate gate. Reference requires uniform anchors and limit 20. Capacity variants select either an 80/20 anchor hot set or limit 100; they cannot be combined.
All profiles allocate 25% of requests to each scope and use all 11 ordinary strategies: 70% anchored
and 30% global, equally weighted within each family. Personalized serving and training traffic are
excluded. The hot set is the first ceil(20%) of each sorted eligible inventory.
The UI command prints an allocated loopback dashboard URL after preparation. Start with the selected profile's first population, 20 for stepped or 100 for sustained. The dashboard starts that complete timed profile; its host, populations, phase lengths, statistics reset, and restart controls are fixed. Stop or cancellation makes the run fail and triggers cleanup. Initial start waiting is at most 600 s.
Fixture and resource envelope¶
Preparation selects the new serving-load Generation Profile v1 and fixture serving-load-v1, preserving
the smoke recipe's structural controls. Each of four scopes has 10,000 Items, 250,000 views,
100,000 purchases, 90 days of history, 20 categories, and 9,500 eligible anchors. Total input is
40,000 Items, 1,000,000 views, and 400,000 purchases. Released smoke, development, and qualification
profiles keep their existing sizes and versions.
The scopes vary Data Source, Tracking ID, and Catalog ID independently. Item identifiers repeat intentionally. Source materialization uses owned PostgreSQL schemas; canonical adapter queries retain the complete Commerce Scope. The normal worker runs once per scope and publishes atomically. Source PostgreSQL and preparation stop before readiness and measured serving.
| Container | CPU limit | Memory limit |
|---|---|---|
| API, one process | 2 | 2 GiB |
| Control PostgreSQL | 2 | 2 GiB |
| Source PostgreSQL | 1 | 2 GiB |
| Preparation | 2 | 4 GiB |
| Locust, one process | 1 | 1 GiB |
Preparation uses two pipeline threads, a 1 GB engine budget, and 4 GB scratch budget. Python 3.14, PostgreSQL 18, and locked optional Locust 2.46.6/FastHttpUser run in separate images/processes. Locust is absent from the ordinary service environment. Pytest explicitly disables its automatic plugin, including when the optional extra happens to be installed locally.
No developer database URL is needed. Passwords are generated privately for each invocation and passed through the owned container environment. Only API and dashboard HTTP ports are published, on loopback with Docker-allocated port numbers. Database ports are internal.
Default root seed is 20260805; cutoff is the invocation's aware UTC time. Reproduce both explicitly:
make load-test LOAD_SEED=20260805 LOAD_CUTOFF=2026-10-02T12:00:00Z
Fixture manifests, materialization receipts, actual Snapshot UUIDs, and package/build identities
are evidence. Snapshot UUIDs need not match between runs. Request streams use MT19937 with
root seed + 1000000 + spawn ordinal; selection is reproducible, while completion order depends
on scheduling and service time.
Interpret the gate¶
Reference and every sustained window require all of the following:
- Complete measured execution and all 44 scope/strategy buckets with at least 1,000 observations.
- At least 237.5 valid served requests/s over the full 300 s denominator, including idle time.
- At least 95% of original request-event durations at or below 90 ms and 99% at or below 150 ms, overall and in every bucket.
- Zero HTTP, transport, task, semantic, or admission errors; required reports and observations; no detected generator saturation or lost users; and successful owned cleanup.
Warm-up is excluded by request start time. Completions belong to that start cohort through the two-second request deadline and ten-second drain; errors remain latency observations. Rounded nearest-rank p50/p95/p99 are labelled estimates and never determine acceptance. A fast HTTP 200 counts as valid only when its expected head, strategy, anchor, timestamps, staleness, ranks, limit, eligibility, category relation, uniqueness, and exact published lane digest agree. Digests include scores, confidence, reasons, and provenance. Empty, fallback, and stale lanes are valid when expected.
Stepped stress reports achieved rate and latency without applying reference latency/rate thresholds.
Only declared service_limit_exceeded 429 and service_limit_unavailable 503 are allowed overload
categories. Unexpected HTTP responses, semantic/transport/task errors, unobserved buckets, missing
evidence, generator saturation, or cleanup failure still fail. Sustained uses the reference policy.
Each command exits zero only for its applicable passing gate and complete cleanup. Failures have finite diagnostic codes. No automatic retry changes an outcome; an explicit rerun has a new run ID. The 180-minute total ceiling includes setup (at most 60 min), fixed traffic phases, and cleanup (at most 120 s). A timeout cannot lengthen a measured phase into a passing result.
Reports and recovery¶
Reports live under ignored artifacts/load-testing/<32-character-run-id>/:
| File | Contents |
|---|---|
summary.json |
Final pass/fail, precise gates, resources, versions, and cleanup result |
aggregate.csv |
Window/overall/bucket counts, original threshold counters, rates, and approximate percentiles |
generator.json |
Client runtime, phase boundaries, warm-up outcome counts, saturation evidence |
fixture.json |
Safe generation manifests, receipts, seeds/cutoff, and four actual heads |
diagnostics.json, preparation.json |
Finite lifecycle, preparation, and generator failure categories |
owner.json, completed.json |
Owned identity and whether cleanup is complete |
Retained reports have no per-run byte quota. Producers, cardinality, sampling, filenames, in-memory buffers, and runtime remain bounded. No raw interactions, per-request URLs/bodies, credentials, bearer material, Shopper identifiers, raw Order IDs, or raw Browsing Session IDs are retained. The private oracle and configuration disappear with the work volume. Keep the last ten completed owned runs, including failures, preserving active and foreign directories. CI retains artifacts for 14 days. Storage/write failures fail the command.
The final summary is written after required evidence markers and retention succeed. A late storage failure leaves no passing summary, even when the measured response counters themselves passed.
Command failures also print their phase, finite operation name, cause, and exit code to the console.
The additive diagnostics.command_failures list in summary.json and diagnostics.json retains
at most the primary and cleanup command failures. Preflight operations distinguish docker_info,
docker_load_inventory, docker_compose_version, postgres_image_pull, and docker_disk_probe;
other subprocess failures use command together with their phase. Exit code is null when no exit
status was observed, such as a missing executable or timeout. Arguments and raw output remain
excluded. A preflight failure produces no serving-performance evidence.
SIGINT/SIGTERM attempt owned cleanup. After a forced process or machine interruption, recover the run recorded in its owner file:
uv run --frozen python scripts/load_testing/run.py --cleanup-run RUN_ID
Recovery verifies run and Compose labels before mutation and preserves a failed outcome. A foreign
resource with inconsistent labels is refused. Cleanup removes only verified containers, networks,
volumes, and run-tagged images; shared Docker build cache remains. CI uses --cleanup-active in an
always-run step before uploading available evidence.
CI and verification limits¶
The serving-load workflow is manual initially, read-only,
serialized, and separate from established verification. It checks out the sibling dependencies'
main branches in an isolated workspace, runs the selected Make command, recovers owned resources,
and uploads reports even after failure. It needs the existing sibling read tokens described in
testing, plus the Docker prerequisites above. Missing prerequisites fail.
Before recovery or execution, CI prepares a job-private DOCKER_CONFIG under RUNNER_TEMP for
anonymous pulls of the harness's public images. It preserves the selected Docker context and
Compose/Buildx plugin discovery, without copying registry credentials or changing the user's
Docker configuration. Context metadata is referenced in place. The runner owns temporary-directory
cleanup. A credential-free Docker Hub entry prevents Docker from auto-selecting a native credential
helper; an empty configuration alone does not. This avoids Docker Desktop credential lookup failures
in the macOS Actions service's separate security session. Public-registry rate limits and network
failures still fail the run; CI does not retry them into a passing result.
The recommendation checkout preserves ignored ownership reports so a later job can recover an interrupted environment before starting its new run. Recovery also removes the owned private scratch directory; all active-run recovery shares a single 120-second cleanup budget.
Promote to automatic pull-request gating only after a valid runner reference baseline proves the declared execution and prerequisites. Local evidence alone does not qualify the CI runner or a deployment. Hosted required-check settings remain outside this implementation.
make load-test-runtime PYTHON_VERSION=3.12 (also 3.13 or 3.14) uses an isolated optional environment
to check real HTTP events, timeout and sanitization, bounded UI, profile timing, and pacing. The
same smoke check runs inside the actual generator image before fixture preparation. Deterministic
gate and contract tests run with make test; actual latency belongs to the disposable load run.
See the design record,
requirement evidence, and
implementation handoff.