Skip to content

Serving load testing

The approved specifications are implemented by a disposable Locust harness. Each invocation prepares synthetic sources, publishes four Recommendation Snapshots, verifies the public Serving API, measures traffic, writes aggregate reports, and removes its Docker resources. The harness measures ordinary serving; preparation is outside the measured interval.

Run a profile

Start from make setup with sibling ../telemetry and ../serving-limits checkouts. Docker Compose must provide a Linux VM with at least 8 CPUs, 12 GiB memory, and 20 GiB free disk; the host also needs 20 GiB free disk. Preflight fails when these prerequisites are missing. Runs are serialized across local worktrees and in CI.

make load-test
make load-test-capacity LOAD_PROFILE=stepped
make load-test-capacity LOAD_PROFILE=sustained
make load-test-capacity LOAD_PROFILE=stepped LOAD_VARIANT=hot-set
make load-test-capacity LOAD_PROFILE=stepped LOAD_VARIANT=limit-100
make load-test-ui LOAD_PROFILE=stepped
Profile Users Pacing Warm-up after full population Measured duration
Reference 100 2.5 requests/s per user, nominal 250 total 60 s 300 s
Stepped 20, 50, 100, 200 No intentional think time 30 s at each step 120 s at each step
Sustained 100 Same pacing as reference 60 s 1,800 s, six 300 s windows

Pacing waits for each response and cannot guarantee an independent arrival rate. Startup phases spread the 100 users across the 400 ms cycle; a slow response delays that user's next request. Measured valid throughput remains a separate gate. Reference requires uniform anchors and limit 20. Capacity variants select either an 80/20 anchor hot set or limit 100; they cannot be combined.

All profiles allocate 25% of requests to each scope and use all 11 ordinary strategies: 70% anchored and 30% global, equally weighted within each family. Personalized serving and training traffic are excluded. The hot set is the first ceil(20%) of each sorted eligible inventory.

The UI command prints an allocated loopback dashboard URL after preparation. Start with the selected profile's first population, 20 for stepped or 100 for sustained. The dashboard starts that complete timed profile; its host, populations, phase lengths, statistics reset, and restart controls are fixed. Stop or cancellation makes the run fail and triggers cleanup. Initial start waiting is at most 600 s.

Fixture and resource envelope

Preparation selects the new serving-load Generation Profile v1 and fixture serving-load-v1, preserving the smoke recipe's structural controls. Each of four scopes has 10,000 Items, 250,000 views, 100,000 purchases, 90 days of history, 20 categories, and 9,500 eligible anchors. Total input is 40,000 Items, 1,000,000 views, and 400,000 purchases. Released smoke, development, and qualification profiles keep their existing sizes and versions.

The scopes vary Data Source, Tracking ID, and Catalog ID independently. Item identifiers repeat intentionally. Source materialization uses owned PostgreSQL schemas; canonical adapter queries retain the complete Commerce Scope. The normal worker runs once per scope and publishes atomically. Source PostgreSQL and preparation stop before readiness and measured serving.

Container CPU limit Memory limit
API, one process 2 2 GiB
Control PostgreSQL 2 2 GiB
Source PostgreSQL 1 2 GiB
Preparation 2 4 GiB
Locust, one process 1 1 GiB

Preparation uses two pipeline threads, a 1 GB engine budget, and 4 GB scratch budget. Python 3.14, PostgreSQL 18, and locked optional Locust 2.46.6/FastHttpUser run in separate images/processes. Locust is absent from the ordinary service environment. Pytest explicitly disables its automatic plugin, including when the optional extra happens to be installed locally.

No developer database URL is needed. Passwords are generated privately for each invocation and passed through the owned container environment. Only API and dashboard HTTP ports are published, on loopback with Docker-allocated port numbers. Database ports are internal.

Default root seed is 20260805; cutoff is the invocation's aware UTC time. Reproduce both explicitly:

make load-test LOAD_SEED=20260805 LOAD_CUTOFF=2026-10-02T12:00:00Z

Fixture manifests, materialization receipts, actual Snapshot UUIDs, and package/build identities are evidence. Snapshot UUIDs need not match between runs. Request streams use MT19937 with root seed + 1000000 + spawn ordinal; selection is reproducible, while completion order depends on scheduling and service time.

Interpret the gate

Reference and every sustained window require all of the following:

  • Complete measured execution and all 44 scope/strategy buckets with at least 1,000 observations.
  • At least 237.5 valid served requests/s over the full 300 s denominator, including idle time.
  • At least 95% of original request-event durations at or below 90 ms and 99% at or below 150 ms, overall and in every bucket.
  • Zero HTTP, transport, task, semantic, or admission errors; required reports and observations; no detected generator saturation or lost users; and successful owned cleanup.

Warm-up is excluded by request start time. Completions belong to that start cohort through the two-second request deadline and ten-second drain; errors remain latency observations. Rounded nearest-rank p50/p95/p99 are labelled estimates and never determine acceptance. A fast HTTP 200 counts as valid only when its expected head, strategy, anchor, timestamps, staleness, ranks, limit, eligibility, category relation, uniqueness, and exact published lane digest agree. Digests include scores, confidence, reasons, and provenance. Empty, fallback, and stale lanes are valid when expected.

Stepped stress reports achieved rate and latency without applying reference latency/rate thresholds. Only declared service_limit_exceeded 429 and service_limit_unavailable 503 are allowed overload categories. Unexpected HTTP responses, semantic/transport/task errors, unobserved buckets, missing evidence, generator saturation, or cleanup failure still fail. Sustained uses the reference policy.

Each command exits zero only for its applicable passing gate and complete cleanup. Failures have finite diagnostic codes. No automatic retry changes an outcome; an explicit rerun has a new run ID. The 180-minute total ceiling includes setup (at most 60 min), fixed traffic phases, and cleanup (at most 120 s). A timeout cannot lengthen a measured phase into a passing result.

Reports and recovery

Reports live under ignored artifacts/load-testing/<32-character-run-id>/:

File Contents
summary.json Final pass/fail, precise gates, resources, versions, and cleanup result
aggregate.csv Window/overall/bucket counts, original threshold counters, rates, and approximate percentiles
generator.json Client runtime, phase boundaries, warm-up outcome counts, saturation evidence
fixture.json Safe generation manifests, receipts, seeds/cutoff, and four actual heads
diagnostics.json, preparation.json Finite lifecycle, preparation, and generator failure categories
owner.json, completed.json Owned identity and whether cleanup is complete

Retained reports have no per-run byte quota. Producers, cardinality, sampling, filenames, in-memory buffers, and runtime remain bounded. No raw interactions, per-request URLs/bodies, credentials, bearer material, Shopper identifiers, raw Order IDs, or raw Browsing Session IDs are retained. The private oracle and configuration disappear with the work volume. Keep the last ten completed owned runs, including failures, preserving active and foreign directories. CI retains artifacts for 14 days. Storage/write failures fail the command.

The final summary is written after required evidence markers and retention succeed. A late storage failure leaves no passing summary, even when the measured response counters themselves passed.

Command failures also print their phase, finite operation name, cause, and exit code to the console. The additive diagnostics.command_failures list in summary.json and diagnostics.json retains at most the primary and cleanup command failures. Preflight operations distinguish docker_info, docker_load_inventory, docker_compose_version, postgres_image_pull, and docker_disk_probe; other subprocess failures use command together with their phase. Exit code is null when no exit status was observed, such as a missing executable or timeout. Arguments and raw output remain excluded. A preflight failure produces no serving-performance evidence.

SIGINT/SIGTERM attempt owned cleanup. After a forced process or machine interruption, recover the run recorded in its owner file:

uv run --frozen python scripts/load_testing/run.py --cleanup-run RUN_ID

Recovery verifies run and Compose labels before mutation and preserves a failed outcome. A foreign resource with inconsistent labels is refused. Cleanup removes only verified containers, networks, volumes, and run-tagged images; shared Docker build cache remains. CI uses --cleanup-active in an always-run step before uploading available evidence.

CI and verification limits

The serving-load workflow is manual initially, read-only, serialized, and separate from established verification. It checks out the sibling dependencies' main branches in an isolated workspace, runs the selected Make command, recovers owned resources, and uploads reports even after failure. It needs the existing sibling read tokens described in testing, plus the Docker prerequisites above. Missing prerequisites fail.

Before recovery or execution, CI prepares a job-private DOCKER_CONFIG under RUNNER_TEMP for anonymous pulls of the harness's public images. It preserves the selected Docker context and Compose/Buildx plugin discovery, without copying registry credentials or changing the user's Docker configuration. Context metadata is referenced in place. The runner owns temporary-directory cleanup. A credential-free Docker Hub entry prevents Docker from auto-selecting a native credential helper; an empty configuration alone does not. This avoids Docker Desktop credential lookup failures in the macOS Actions service's separate security session. Public-registry rate limits and network failures still fail the run; CI does not retry them into a passing result.

The recommendation checkout preserves ignored ownership reports so a later job can recover an interrupted environment before starting its new run. Recovery also removes the owned private scratch directory; all active-run recovery shares a single 120-second cleanup budget.

Promote to automatic pull-request gating only after a valid runner reference baseline proves the declared execution and prerequisites. Local evidence alone does not qualify the CI runner or a deployment. Hosted required-check settings remain outside this implementation.

make load-test-runtime PYTHON_VERSION=3.12 (also 3.13 or 3.14) uses an isolated optional environment to check real HTTP events, timeout and sanitization, bounded UI, profile timing, and pacing. The same smoke check runs inside the actual generator image before fixture preparation. Deterministic gate and contract tests run with make test; actual latency belongs to the disposable load run. See the design record, requirement evidence, and implementation handoff.