Exercise reproducible Serving API traffic with Locust¶
Status: Agreed requirements implemented by the disposable serving harness. Verification and qualification limits are recorded in the handoff.
Tracker: Locust serving load testing: environment, traffic, and CI gates.
Problem Statement¶
The current qualification client repeatedly requests one anchor and one Recommendation Strategy. Developers need representative serving traffic to discover throughput limits, tail-latency changes and bottlenecks across the API. Counting HTTP 200 responses alone would miss invalid recommendations, incorrect Snapshot use and Commerce Scope isolation failures.
Solution¶
Provide a Locust suite that sends bounded, reproducible ordinary anchored and global Serving API requests against prepared synthetic Recommendation Snapshots. Named workloads support a fixed reference run for CI and controlled increasing or sustained traffic for capacity exploration. Each request is validated against the seeded serving contract, and its result contributes to safe, bounded statistics.
User Stories¶
- As a developer, I want load scenarios written with Locust, so that traffic behavior is maintainable in the project's Python ecosystem.
- As a developer, I want ordinary anchored and global Serving API traffic, so that capacity evidence covers the selected serving workload.
- As a tester, I want multiple anchor Items and Recommendation Strategies, so that one hot request cannot stand in for the whole API.
- As a developer, I want the request mix and Item selection distribution recorded, so that runs can be compared against the same workload.
- As a tester, I want every request scoped by Data Source, Tracking ID and Catalog ID, so that traffic preserves Commerce Scope isolation.
- As a tester, I want expected Snapshot and eligible Item checks, so that HTTP success cannot conceal an incorrect Recommendation Set.
- As a tester, I want recommendation limits, ordering, uniqueness and anchor exclusions checked, so that load does not hide public response-contract violations.
- As a tester, I want declared empty and fallback results interpreted correctly, so that valid sparse-evidence behavior is not mislabeled as failure.
- As a developer, I want reproducible random choices and a recorded workload definition, so that changing traffic distributions is an explicit experimental change.
- As a developer, I want a fixed bounded reference profile, so that CI performance comparisons describe a stable population.
- As a developer, I want stepped load increases, so that I can locate the point where throughput or latency stops meeting the objectives.
- As an operator, I want a sustained traffic profile, so that memory, connection and latency changes become visible over a defined interval.
- As a developer, I want concurrent-user count distinguished from achieved request rate, so that service slowdown does not create a false capacity pass.
- As an operator, I want admission rejections distinguished from unexpected service failures, so that overload behavior can be interpreted.
- As a developer, I want finite request timeouts and a finite total execution deadline, so that a stalled service cannot leave unbounded pending work.
- As a developer, I want Locust's dashboard available for bounded capacity exploration, so that I can inspect changes while a run is active.
- As a CI maintainer, I want a headless workload mode, so that the same traffic contract can run without manual interaction.
- As a maintainer, I want route and error statistics to have bounded safe names, so that varying Items and exception details cannot leak data or inflate reports.
- As a developer, I want generator saturation reported, so that a generator bottleneck is not claimed as service capacity.
- As a maintainer, I want the selected load dependencies compatible with the project's supported Python versions, so that the suite fits existing development and CI environments.
Implementation Decisions¶
- User-confirmed: Locust; Serving API only; disposable local execution; both capacity exploration and CI performance gates.
- User-confirmed fixture scale is 10,000 Items, 250,000 views and 100,000 purchases per scope across four independently prepared Commerce Scopes. The dedicated load recipe preserves the smoke fixture's other structural controls and records its own identity.
- Cover all 11 nonpersonalized strategies: eight anchored and three global. Send 70% of requests to anchored strategies and 30% to global strategies, with equal weights within each family.
- Use four scopes to exercise Data Source, Tracking ID and Catalog ID isolation with deliberately repeated identifiers. Allocate traffic equally: 25% to each scope.
- The CI reference selects eligible anchor Items uniformly and uses limit 20. A separate capacity variant sends 80% of anchored requests to a declared 20% eligible-Item hot set; another capacity variant exercises limit 100. Record each variant's changed controls explicitly.
- Initial measured traffic uses the existing nonpersonalized anchored and global HTTP serving contracts. No recommendation computation or source data read moves into the generator.
- The generator consumes the prepared workload and Snapshot identities from the owned environment. Data Source and Tracking ID remain request inputs, and Catalog ID remains part of scope selection.
- Request validation uses the expected same-scope published Snapshot, eligible Item inventory and existing strategy/anchor/limit/provenance semantics. It must not assume the response itself echoes the complete Commerce Scope.
- Declare each workload's strategy mix, required request buckets, selection distribution, response limit, seed, scope distribution, warm-up and measured phases, concurrency/pacing and maximum duration. Configurations must reject values outside the execution's bounds.
- The approved reference targets a nominal paced 250 requests/second across 100 virtual users total, with 60 seconds of warm-up after the full population is established and then 5 minutes measured. This is a workload target, not measured serving capacity or an independent arrival-rate guarantee.
- The approved stepped capacity profile uses 20, 50, 100 and 200 concurrent users, with 30 seconds to settle and 2 minutes measured at each level. It has no intentional think time and no further automatic concurrency steps. The approved sustained profile uses 100 users at the nominal reference pacing of 250 requests/second, with 60 seconds warm-up and 30 minutes measured.
- Offer the reference and capacity profiles as separately selectable bounded runs. Hot-set and limit-100 variants are individually selectable and explicitly declared. Interactive adjustments stay within the approved concurrency and finite phase bounds; they do not silently change a reference run into capacity exploration.
- Report requested settings and achieved valid served throughput separately. Locust's normal user pacing waits for responses, and a virtual-user count or pacing target does not guarantee an independent arrival rate.
- Ordinary serving uses already published Snapshots throughout measurement. Preparation training is complete before the measured interval; synchronous training and merchant-source access remain prohibited.
- HTTP status alone is insufficient. Semantic failures, malformed JSON, transport failures, task errors and unexpected statuses have finite safe outcome categories.
- Declared overload/admission responses remain visible and distinct from valid served throughput. Rate/concurrency 429 and admission-state/policy 503 are different outcomes. The approved reference allows neither; a complete stepped stress run reports explicitly expected admission overload without failing solely for reaching capacity. Unexpected service errors and invalid content still fail execution.
- The approved measurement excludes warm-up and assigns requests to their start phase through a bounded completion/drain deadline. Reference evidence covers all 44 Commerce Scope × strategy combinations, with at least 1,000 observations each. Reference and sustained-window gate acceptance follows the performance-gate specification.
- Use fixed route-template/strategy buckets for Locust statistics. Sanitization must occur before failures, exceptions, URLs or response bodies reach retained statistics or logs.
- Keep connections, in-flight requests, generator processes, in-memory histories, report structure/cardinality and sampling bounded. The approved reference uses one generator process. Retained report storage has no per-run byte quota; privacy, finite runtime and bounded in-memory state still apply.
- The user approved Locust 2.46.6, FastHttpUser and a separately invoked optional load dependency. Runtime HTTP/event, timeout/sanitization, phase/pacing, and UI checks pass on macOS ARM64 CPython 3.12–3.14 and the Linux ARM64 Python 3.14 generator image; exact versions and measured qualification limits are recorded in the handoff.
- Apply the approved 2-second request timeout, 10-second measured request drain and 180-minute total lifecycle ceiling. Existing measured phase lengths stay fixed; initial UI-start waiting is limited to 10 minutes. Use the selected container/process resources and lifecycle in the environment specification.
- The workload scale, strategy mix, scope allocation, Item-selection variants, response limits, nominal rate, concurrency, phase lengths, performance-gate policy, client/version and initial environment envelope are user-approved. Technical design and measured/runtime validation remain work; no achieved-capacity claim follows from these decisions.
Testing Decisions¶
- Exercise the existing HTTP Serving API as the primary service seam; observe workload behavior through the load command's safe reports and exit outcomes.
- Reuse existing API contract tests for missing/empty/stale Snapshots, limits, ordering, eligibility, provenance, invalid scopes and failure semantics.
- Given a seeded same-scope Snapshot, when an anchored or global workload runs, then each valid result uses the expected head and respects the selected request's public contract.
- Given identical Product IDs in different Commerce Scopes, when both scopes are exercised, then each result matches its expected head and eligible inventory without cross-scope evidence.
- Given a fast HTTP 200 response with invalid recommendation content, when validation completes, then the request is classified as invalid and cannot count as valid served throughput.
- Given a legitimate empty response or Fallback Strategy declared by the fixture, when it is received, then its documented semantics determine correctness.
- Given declared admission traffic, when rate/concurrency rejection or admission-state unavailability occurs, then each outcome is retained in its distinct bounded category.
- Given a timeout, task exception, worker loss or saturated generator, when the run concludes, then the report records the measurement limitation and cannot present an unqualified success.
- Verify explicit seeds, request distributions, phase attribution, bounds, response classifications and sanitization with deterministic harness tests. Validate capacity and actual timings in the disposable load environment.
- Perform dependency resolution, import and HTTP smoke checks against the selected runtime combinations during implementation; package metadata is not runtime evidence.
Out of Scope¶
Personalized serving and For You, Shopper interaction ingestion, authenticated Swimlane load, ANN enablement or index comparisons, browser rendering and storefront journeys, Training API load, concurrent training/publication experiments, training-scale or monthly-availability qualification, staging/production traffic and hosted settings changes.
Further Notes¶
Status: the bounded Locust suite and profiles are implemented. Runtime and measured reference outcomes are recorded in the handoff. Stepped/sustained measured capacity and CI-runner qualification require their own complete runs; the design record does not assert separate design approval.
Primary-source findings are recorded in Verify Locust compatibility and load generation semantics, research branch research/locust-serving-load-testing at commit 2766803bd121c3f8642016f97ee42ad2efcf5a48.
This specification is related to the Locust serving load-test map. Define the serving workload and reference profile, Define repeatable CI performance gates and capacity evidence and Choose the disposable load environment and command lifecycle are resolved. The ready-for-agent label routes the agreed requirements to later design and implementation work.