Locust serving load-testing specifications¶
These agreed specifications synthesize the conversation's confirmed scope: Serving API only, a disposable local environment, and both capacity exploration and CI performance gates. The harness is implemented. The run guide, design record, and verification handoff separate implementation evidence from measured performance and CI-runner qualification.
All three complete specs are published in the Linear specification bundle, labelled ready-for-agent. The workspace issue limit required a single issue for the bundle.
- Prepare and own a disposable local serving load environment.
- Exercise reproducible Serving API traffic with Locust.
- Gate and report local serving performance in CI.
Confirmed scope and acceptance boundary¶
Locust exercises ordinary anchored and global Serving API requests against synthetic, published, Commerce Scope-specific Recommendation Snapshots. A setup Training Run may prepare those snapshots before measurement. Measured Training API traffic and personalization are outside the initial suite.
Acceptance is observed through the existing HTTP Serving API and the load workflow's command reports, exit outcomes and resource cleanup. No service-only test hooks are required.
The existing Serving API objectives remain p95 at most 90 ms and p99 at most 150 ms. Local synthetic performance evidence does not establish deployment-scale qualification or monthly availability.
The user approved a 10x increase over the initially proposed smoke-scale fixture: 10,000 Items, 250,000 views and 100,000 purchases per Commerce Scope, preserving the smoke recipe's other structural controls. Four scopes exercise Data Source, Tracking ID and Catalog ID isolation. This requires a dedicated reproducible local load fixture and does not change released profiles.
The reference covers all 11 nonpersonalized strategies: 70% anchored requests and 30% global, with equal weights within each family, uniform eligible-Item anchors and limit 20. Separate capacity variants exercise an 80/20 anchor hot set and limit 100. Allocate traffic equally: 25% to each of the four scopes.
The approved reference targets a nominal paced 250 requests/second across 100 virtual users, with 60 seconds of warm-up after the full population is established and 5 minutes measured. A separately selected sustained profile uses that reference load for 30 measured minutes after 60 seconds warm-up. The stepped capacity profile uses 20, 50, 100 and 200 concurrent users, with 30 seconds settling and 2 minutes measured per level. Nominal pacing is a requested workload; achieved valid served throughput is separate evidence.
The approved reference gate applies the 90/150 ms latency objectives overall and to all 44 Commerce Scope × strategy combinations, with at least 1,000 observations each and unrounded threshold counters. It requires at least 237.5 valid served requests/second and zero unexpected errors, invalid responses or admission rejections. Warm-up is excluded; requests are attributed by start phase through a bounded completion period using the complete measured interval as the rate denominator. Sustained runs satisfy those gates in every five-minute window. Stepped stress reports expected overload under its own policy. Invalid/noisy runs fail CI with a recorded cause; explicit reruns retain prior results and automatic retries cannot turn a failure into a pass.
The approved environment is a dedicated owned Docker Compose stack with fixed process/container resources, loopback ports and automatic cleanup. Source PostgreSQL and the preparation worker stop before measurement. Locust 2.46.6 with FastHttpUser runs separately through an optional load extra. Make commands expose headless reference, selected capacity and local UI modes. The maximum total lifecycle is 180 minutes; the recommended request/setup/drain/cleanup sublimits remain fixed.
The user explicitly selected unlimited report storage per run: no harness byte quota or derived 500 MiB total cap. JSON summaries, aggregate CSV and sanitized diagnostics retain bounded structure, cardinality, sampling, in-memory state and runtime. Retain the last 10 completed local runs and CI artifacts for 14 days. Storage failures cannot create a passing result. CI starts manually and promotes to automatic pull-request gating after a valid baseline verifies runner prerequisites and the declared reference execution.
Decision gates¶
The Locust serving load-test map indexes the completed decisions. No pending human questions remain for these requirements.
Define the serving workload and reference profile is complete and records the user-approved fixture, distributions and bounded profiles above.
Define repeatable CI performance gates and capacity evidence is complete and records the approved measurement, acceptance and outcome policies above.
Choose the disposable load environment and command lifecycle is complete and records the approved stack, resources, commands, deadlines, report retention and CI rollout.
Verify Locust compatibility and load generation semantics is complete. Runtime HTTP/event, timeout, phase, pacing, and UI checks now pass on macOS ARM64 Python 3.12–3.14 and the Linux ARM64 Python 3.14 generator image. These checks do not measure capacity.
The specification bundle carries the ready-for-agent label and native blocking relationships. It is related to the map rather than becoming a decision-ticket child. Publishing the specs records the resolved human decisions without asserting technical design approval or measured capacity.
Governing documentation¶
- Commerce language and isolation boundaries.
- Approved service requirements.
- Repository testing and verification.
- Agent readiness and CI context.
- Operational observations.
Repository Make targets and documentation checks verify these requirements. The 58-story evidence matrix records implementation and verification coverage. The manual CI workflow awaits a valid runner reference baseline before pull-request promotion; measured load outcomes are recorded in the handoff.