Prepare and own a disposable local serving load environment¶
Status: Agreed requirements implemented by the disposable serving harness. Verification and qualification limits are recorded in the handoff.
Tracker: Locust serving load testing: environment, traffic, and CI gates.
Problem Statement¶
Developers need a reproducible environment for measuring Serving API behavior without sharing or damaging an existing local stack. The current local stack has persistent resources, and starting its services does not publish a Recommendation Snapshot. A load test against an unprepared service could measure missing-snapshot errors instead of useful recommendation serving.
Solution¶
Provide a bounded local workflow that owns its environment, prepares synthetic Catalogs and published Recommendation Snapshots through existing service boundaries, verifies readiness, and exposes serving traffic to the load generator. The workflow records its resource and fixture identity and cleans up only the resources it created, including after failures and cancellation.
User Stories¶
- As a developer, I want a supported local setup command, so that I can prepare a load environment without assembling undocumented steps.
- As a developer, I want each execution to own its services, storage and ports, so that concurrent or existing local environments remain independent.
- As a developer, I want synthetic inputs with explicit seeds and logical cutoffs, so that workload preparation is reproducible.
- As a tester, I want every prepared source and Snapshot bound to its complete Commerce Scope, so that similarly named Items cannot combine evidence across Catalogs or Commerce Properties.
- As a developer, I want the selected data scale visible before preparation, so that a large Generation Profile is never started accidentally.
- As a developer, I want setup to publish a Recommendation Snapshot before measurement, so that the Serving API has valid evidence to serve.
- As a tester, I want readiness to verify representative serving responses, so that an open port is insufficient to start a measured run.
- As a tester, I want each execution to identify its actual published serving head, so that different Snapshot UUIDs do not undermine fixture validation.
- As a developer, I want the load generator isolated from the API and setup processes, so that Locust's runtime behavior cannot modify the service runtime.
- As an operator, I want API, database and generator resource settings recorded, so that performance differences can be interpreted.
- As a developer, I want the local environment exposed only through its intended local interfaces, so that a test does not direct traffic at another deployment.
- As a developer, I want finite setup, training, readiness and shutdown deadlines, so that a failed prerequisite cannot hang the workflow.
- As a developer, I want cleanup after success, failure, timeout and cancellation, so that repeated runs do not accumulate services or storage.
- As a developer, I want cleanup to verify ownership, so that it cannot remove an unrelated service, database, volume or artifact.
- As a maintainer, I want dependencies installed through the repository's supported locked workflow, so that local execution matches the selected runtime versions.
- As a CI maintainer, I want clear prerequisite failures and bounded diagnostics, so that missing local infrastructure is distinguished from a service performance failure.
Implementation Decisions¶
- User-confirmed scope is Serving API load in a disposable local environment. Locust is the selected generator; capacity exploration and CI gates are both required.
- Use existing synthetic-source generation, Data Source Adapter, Training API/worker and snapshot-publication boundaries for preparation. A setup Training Run occurs before measurement; it is not a Training API load test or a training-capacity measurement.
- Every source, prepared run, Snapshot and serving request carries the complete Commerce Scope: Data Source, Tracking ID and Catalog ID. No production commerce records are used.
- Readiness must establish a complete published Snapshot and a valid HTTP serving response. Record the expected head and scope-local eligible Item inventory; deterministic preparation does not require deterministic Snapshot UUIDs.
- The environment controller owns a unique execution identity, resource inventory and bounded cleanup. Sharing a developer's existing control database or relying on persistent stack defaults does not meet the disposal contract.
- Locust runs in a separate process from the API, worker, preparation and ordinary tests. Load dependencies belong to an optional development/testing dependency surface and must not enter the service runtime dependency boundary.
- The supported command surface is the repository's committed Make workflow. Setup, execution, report retention and teardown must have defined outcomes and finite deadlines.
- Record the declared API process/thread configuration, database topology, generator topology and resource envelope. A pipeline aggregation-memory setting alone is not a process-memory limit.
- Keep credentials, bearer material, Shopper identifiers, raw Order IDs, raw Browsing Session IDs and raw interaction rows out of diagnostics and retained evidence. Any generated source belongs to the owned synthetic environment.
- The user approved a 10x increase over the proposed smoke-scale fixture: 10,000 Items, 250,000 views and 100,000 purchases per Commerce Scope. Four scopes total 40,000 Items, 1,000,000 views and 400,000 purchases. Preserve the smoke recipe's 90-day history and structural/distribution controls.
- Prepare a reproducible versioned local load fixture with those counts. The dedicated
serving-loadGeneration Profile v1 uses fixture identityserving-load-v1. Released presets retain their parameters; the scaled fixture is not labelled as an unmodified smoke/development Generation Profile. - The four scopes exercise isolation differences in Data Source, Tracking ID and Catalog ID using intentionally repeated identifiers and independently published expected serving heads.
- The approved reference allocates 25% of traffic to each scope and targets 250 requests/second across 100 virtual users, with 60 seconds of warm-up after the population is established and 5 minutes measured. The sustained capacity profile uses that reference load for 30 minutes after 60 seconds of warm-up. A separate stepped profile measures 20, 50, 100 and 200 concurrent users, settling for 30 seconds and measuring for 2 minutes at each level.
- The user approved a dedicated Docker Compose stack with API, control PostgreSQL, source PostgreSQL, preparation worker and a separate Locust service. Each execution owns a unique project, network, volumes and identity. Expose only loopback interfaces with automatically allocated ports; clean up only owned resources on success, failure, timeout and cancellation.
- Prepare and publish all four expected scope heads before warm-up. Stop the preparation worker and source PostgreSQL service before serving measurement. Synthetic source data disappears when owned source volumes are removed. Metadata Similar Items preparation remains outside serving measurements; its cost has not been measured.
- Apply the approved initial process/container limits below. One API process and one generator process define the reference topology. Record actual resource/runtime/host identities; these allocations do not establish achieved serving capacity.
| Component | CPU limit | Memory limit |
|---|---|---|
| API, one process | 2 vCPUs | 2 GiB |
| Control PostgreSQL | 2 vCPUs | 2 GiB |
| Source PostgreSQL | 1 vCPU | 2 GiB |
| Preparation worker | 2 vCPUs | 4 GiB |
| Locust, one process | 1 vCPU | 1 GiB |
- Require at least 8 Docker vCPUs, 12 GiB VM memory and 20 GiB available disk initially. Serialize load runs and require the declared resource observations and generator saturation warnings for a valid gate. Detected saturation or unmet prerequisites produce an invalid/failing outcome. Preparation initially uses 2 threads, a 1 GB engine budget and a 4 GB scratch budget; engine memory is not the process-memory limit.
- Use the existing Python 3.14/container and PostgreSQL 18 baseline with both sibling build contexts. The user approved Locust 2.46.6 and FastHttpUser in an optional load dependency extra and a separate generator image/process. Ordinary service and test processes do not import Locust or receive its monkey patches.
- Provide
make load-testfor the headless reference,make load-test-capacitywith an explicit stepped or sustained profile, andmake load-test-uifor the loopback UI. Each command owns preparation, publication/readiness, execution, reporting and teardown. UI controls remain within the approved 200-user and finite profile bounds. - The user increased the maximum lifecycle to 180 minutes per invocation. Retain the recommended sublimits: setup including build/generation/publication at most 60 minutes, individual request timeout 2 seconds, measured request drain 10 seconds, cleanup at most 2 minutes and initial UI-start wait at most 10 minutes. These fit inside the lifecycle ceiling. The existing warm-up and measured phase lengths do not increase; the ceiling is not an expected duration.
- Retain a versioned JSON summary, aggregate CSV and sanitized diagnostics under ignored
artifacts/load-testing/<run-id>/. The user explicitly requested unlimited storage per run: impose no harness byte quota on retained reports and no derived 500 MiB total cap. Keep the last 10 completed local runs and retain CI artifacts for 14 days; retention affects only owned completed reports and preserves active runs. - Unlimited per-run storage changes the storage budget only. Bound report structure, filenames, label cardinality, sampling, producer/runtime duration and in-memory buffers; retain the privacy and source/serving boundaries. Do not substitute per-request dumps or raw interaction rows for aggregate evidence. Storage/write failures cannot become a passing result or silently truncate required evidence.
- Start with a manually triggered CI load workflow that checks Docker/resource/disk prerequisites, serializes execution, performs owned cleanup and uploads available safe evidence on every outcome. Missing prerequisites fail rather than skip the gate. Enable automatic pull-request gating after a valid baseline run proves runner prerequisites and the declared reference execution; a failed target requires diagnosis rather than silently lowering thresholds. Hosted runner, secret and required-check settings are outside the approved scope.
Testing Decisions¶
- Primary acceptance boundaries are the public HTTP Serving API and the local workflow's commands, reports and exit outcomes. Validate externally observable preparation and cleanup without introducing service-only load-test hooks.
- Reuse the existing deterministic synthetic-source, end-to-end Training Run, snapshot-publication and API contract testing conventions.
- Given a new owned environment, when preparation completes, then the expected same-scope head is published and representative serving requests satisfy its fixture contract.
- Given an existing unrelated local stack, when the load workflow starts and later cleans up, then the unrelated resources and published serving head remain available.
- Given a setup, training or readiness failure, when the workflow exits, then it reports the failed phase, performs bounded owned-resource cleanup and starts no measurement.
- Given timeout, interruption or incomplete cleanup, when the workflow exits, then the report cannot claim a complete successful load execution.
- Given all expected heads are published, when the source service and preparation worker stop, then same-scope Serving API requests remain correct throughout the measured profile.
- Given insufficient Docker resources, unavailable required observations or a breached phase/lifecycle deadline, when the command runs, then it identifies the failed prerequisite/phase, cannot pass and performs owned cleanup.
- Exercise scope mismatches, missing or corrupt Snapshots, invalid configuration, port conflicts and ownership mismatches through the command/API outcomes.
- Verify report structure/cardinality, sampling and memory bounds, retention ownership and secret exclusions. A report larger than the previously proposed 50 MiB quota must not fail solely for its size; actual storage failure remains an execution failure. Avoid elapsed-time assertions in ordinary unit tests; actual timing measurements belong to the load runs.
- During later implementation, run focused harness checks and repository static/contract/integration gates. Changes to container artifacts require the Compose gate; migration checks apply only if migrations are introduced.
Out of Scope¶
Training API load, Training Run scale qualification, real merchant data, staging or production traffic, personalization, new ANN enablement, hosted runner/secret/required-check changes, package publication and deployment.
Further Notes¶
Status: agreed requirements from the user-confirmed conversation. The environment is not implemented, no measured run is claimed and technical design is not self-approved.
This specification is related to the Locust serving load-test map. Define the serving workload and reference profile, Define repeatable CI performance gates and capacity evidence and Choose the disposable load environment and command lifecycle are resolved. No open human decision blocks these requirements.
The tracker label records the original implementation routing. The local design record and story evidence matrix describe implemented boundaries and checks. Runner and measured profile qualification remain separate evidence.