Gate and report local serving performance in CI¶
Status: Agreed requirements implemented by the disposable serving harness. Verification and qualification limits are recorded in the handoff.
Tracker: Locust serving load testing: environment, traffic, and CI gates.
Problem Statement¶
Developers want both capacity measurements and an automated regression gate. A load run can appear successful despite insufficient traffic, missing routes, empty measurement windows, invalid responses or an overloaded generator. Approximate displayed percentiles can also hide requests exceeding an objective near a histogram bucket boundary.
Solution¶
Evaluate a declared reference workload over a complete measured interval and publish bounded evidence of correctness, achieved throughput, tail latency, failures and resource conditions. A CI command succeeds only when the applicable gate conditions and execution completeness are satisfied. Capacity exploration reports overload and bottlenecks explicitly under its own declared mode.
User Stories¶
- As a CI maintainer, I want a bounded headless performance command, so that serving regressions can be checked automatically.
- As a developer, I want the existing Serving API latency objectives applied, so that the load suite evaluates the service's intended behavior.
- As a developer, I want warm-up and measurement distinguished, so that startup traffic does not silently change the measurement population.
- As a tester, I want requests crossing phase boundaries attributed explicitly, so that warm-up or drain responses cannot contaminate a gate.
- As a developer, I want the full measured interval used for achieved throughput, so that trailing inactivity cannot be omitted.
- As a developer, I want a minimum valid request population and achieved throughput, so that a lightly loaded service cannot receive a capacity pass.
- As a tester, I want required route/strategy buckets checked, so that a fast dominant route cannot conceal an untested or slow route.
- As a developer, I want latency acceptance protected against histogram rounding, so that a request above a threshold cannot create a false pass.
- As an operator, I want failed and rejected responses retained in the evidence, so that reported objectives are not satisfied by silently removing failures.
- As a tester, I want semantic response failures to affect the outcome, so that successful HTTP status is not mistaken for correct serving.
- As an operator, I want rate/concurrency rejection distinguished from state/policy unavailability and service errors, so that the failure boundary is observable.
- As a developer, I want generator saturation and local contention recorded, so that capacity claims name their measurement limits.
- As a CI maintainer, I want zero requests, missing required buckets and truncated runs rejected, so that incomplete executions cannot pass.
- As a CI maintainer, I want timeout, task errors, worker loss and missing final evidence to produce a failing outcome, so that automation cannot treat missing work as success.
- As a developer, I want capacity exploration results distinguished from reference-gate results, so that deliberate overload is interpreted under the correct workload contract.
- As a developer, I want versions, seeds, data/profile identity, Snapshot identity and environment resources recorded, so that another run can reproduce the experiment.
- As an operator, I want throughput, p50/p95/p99, classified outcomes and resource observations reported, so that bottlenecks can be investigated.
- As a maintainer, I want bounded aggregate evidence without credentials or sensitive identities, so that reports can be retained safely.
- As a CI maintainer, I want prerequisite and noisy-run conditions explicit, so that runner problems are not confused with validated service regressions.
- As a developer, I want the gate evaluator tested with deterministic observations, so that gate correctness does not depend on fast unit-test timing.
- As a maintainer, I want repository CI integration to preserve existing verification, so that adding a load check does not weaken established gates.
- As a developer, I want local performance evidence labeled accurately, so that local synthetic measurements are not promoted into deployment-scale qualification.
Implementation Decisions¶
- User-confirmed: both capacity/bottleneck exploration and CI performance gates, for serving-only traffic in a disposable local environment.
- The approved fixture has 10,000 Items, 250,000 views and 100,000 purchases per scope across four isolated scopes. The reference covers all 11 nonpersonalized strategies with a 70% anchored/30% global mix, equal weights within each family, uniform eligible-Item anchors and limit 20. Capacity variants add a declared 80/20 anchor hot set and limit 100; they do not redefine the reference gate.
- The approved reference allocates 25% of traffic to each scope and targets a nominal paced 250 requests/second across 100 virtual users. Warm-up lasts 60 seconds after the population is established; measurement lasts 5 minutes. The separately selected sustained profile uses this reference load for 30 measured minutes after 60 seconds warm-up. The stepped profile uses 20, 50, 100 and 200 concurrent users, with 30 seconds settling and 2 minutes measured per level. No achieved-capacity result is claimed.
- Reuse the existing approved Serving API objectives: p95 at most 90 ms and p99 at most 150 ms. The user approved these local regression checks overall and for every Commerce Scope × strategy combination. Preserve the governing population of successful and product-defined error responses; exclude only caller cancellation as the service requirement defines. Training API latency and monthly availability are outside this gate.
- The reference gate checks response correctness, complete execution, the agreed latency objectives, sufficient samples, achieved valid served throughput and the selected error/admission policy. A latency-only success does not meet the contract.
- Exclude warm-up and attribute requests by their start phase. Include every attempt started in the measured interval through the approved 2-second request timeout and 10-second completion/drain deadline; retain timeouts as failures. Valid completions from this measured-start cohort use the full declared interval, including idle time, as the throughput denominator. The full lifecycle is capped at 180 minutes without extending measured phases.
- The approved minimum achieved valid served throughput is 237.5 requests/second, 95% of the nominal 250 target, over the complete 300-second reference measurement. Invalid responses and admission rejections do not count toward this minimum. Requested pacing and observed throughput remain separate evidence.
- Retain all relevant response outcomes and their counts. Failures and product-defined errors must not be silently removed from latency evidence to satisfy the objectives. Valid served throughput excludes invalid responses and admission rejections.
- The reference allows zero unexpected HTTP statuses, request/transport/task failures, invalid response content or admission 429/503 responses. Expected empty/fallback/stale HTTP 200 results are successful only when they match the same-scope published-head contract and eligibility semantics. A canceled or incomplete test fails execution completeness even where caller cancellation is excluded from the latency population.
- Report cumulative measured-interval p50/p95/p99 with their estimator and precision. Locust's ordinary histogram rounds observations, including to 10 ms buckets near 150 ms, so the gate must use an acceptance method that cannot falsely pass through this rounding.
- The user approved bounded counters over original unrounded request-event durations. Require at least 95% of observations at or below 90 ms and at least 99% at or below 150 ms, overall and in every required combination. Report Locust histogram percentiles with their approximation labelled; histogram rounding cannot determine gate acceptance.
- Require at least 1,000 measured observations in each of the 44 Commerce Scope × strategy combinations. Retain required-bucket sample evidence; an undersampled combination cannot pass. Apply the approved process/container resources in the environment specification.
- A CI success cannot override incomplete measurement, invalid responses, request/task failures outside the approved error policy, worker loss, missing reports, timeouts, unresolved generator saturation or failed owned-resource cleanup.
- A complete stepped stress run reports reached concurrency, achieved served throughput, latency and explicitly classified admission/overload conditions. It does not fail solely for exceeding reference performance or expected admission capacity. Unexpected statuses such as snapshot corruption, invalid content, transport/task/worker failures, missing evidence or failed cleanup still fail execution.
- The sustained run applies the selected reference gates in every consecutive five-minute measured window. Each window meets the same latency, 44-combination sample, 237.5 valid-RPS and zero-error requirements; late degradation cannot be concealed by a whole-run aggregate. It does not receive the stress profile's overload allowance.
- Mark generator saturation, lost workers, shortened or missing measurement, missing required evidence and unmet declared resource prerequisites as invalid/failing CI outcomes with a recorded cause. No automatic retry converts a failure to a passing result; an explicit rerun retains both outcomes. Require declared resource observations and generator saturation warnings; record actual host/runtime/resource identities and unavailable optional measurements.
- Record bounded CPU/memory, database-connection and service/generator observations where supported, together with unavailable measurements. Native pipeline limits and package availability do not prove process headroom.
- Evidence identifies the build/runtime/dependency versions, operating system/architecture, declared resources/topology, seeds, data/workload identity, expected published heads, phase durations, request buckets, sample counts, rates, approximate percentiles, precise threshold evidence and classified outcomes.
- Retain versioned JSON summaries, aggregate CSV and sanitized diagnostics under the repository's ignored artifact convention. Per the user's explicit choice, retained report storage has no per-run byte quota and no derived total byte cap. Retain the last 10 completed local runs and CI artifacts for 14 days. Bound report structure/cardinality, sampling, filenames, runtime and in-memory buffers; storage failure cannot silently truncate required evidence into a pass. No raw interaction rows, per-request identity-bearing paths, credentials, bearer material, Shopper identifiers, raw Order IDs or raw Browsing Session IDs are retained.
- The user approved a manually triggered CI workflow first, with Docker/resource/disk preflight, serialized execution, owned cleanup and safe artifact upload on every outcome. Missing prerequisites fail rather than skip the gate. Enable automatic pull-request gating after a valid baseline run proves runner prerequisites and the declared reference execution. The current CI environment is self-hosted macOS ARM64; hosted runner, secret and required-check setting changes remain outside scope.
- No additional production telemetry imports or service-only inspection endpoints are required by this specification. Reuse the existing observability adapter boundary where service measurements are needed.
Testing Decisions¶
- Observe acceptance through the public HTTP Serving API and the load command's report and exit status. Test the gate logic with controlled observation summaries rather than machine-speed assertions.
- Reuse the repository's contract, privacy, disposable-integration and delivery-workflow testing conventions.
- Given correct responses, sufficient samples/throughput and a complete measurement satisfying the approved policy, when the gate evaluates, then its report explains each criterion and the command succeeds.
- Given a request above a latency objective that rounds into an apparently passing Locust histogram bucket, when the precision-sensitive gate evaluates, then it applies the declared precise convention and cannot falsely pass.
- Given zero samples, a missing required bucket, under-rate traffic or a shortened interval, when the run ends, then it cannot pass even if the displayed latency percentiles appear acceptable.
- Given an aggregate that passes while one scope/strategy combination misses a threshold or has fewer than 1,000 samples, when the gate evaluates, then it fails with that bounded combination identified.
- Given invalid JSON or recommendation content, an unexpected failure or a declared admission response, when results are evaluated, then each outcome is retained and accepted or rejected only under its explicit policy.
- Given warm-up overlap, a request finishing during drain or trailing inactivity, when measurements are summarized, then phase attribution and the rate denominator follow the declared population.
- Given task failure, worker loss, generator saturation, process timeout, missing final evidence or failed cleanup, when automation receives the result, then it cannot record a successful reference gate.
- Given a sustained run whose last five-minute window fails despite a passing whole-run aggregate, when it is evaluated, then the run fails. Given expected admission overload during a complete stress run, then it remains classified capacity evidence under the approved stress policy.
- Given an invalid or failed run followed by an explicit rerun, when evidence is retained, then both outcomes remain available and an automatic retry cannot erase the first outcome.
- Verify evidence structure/cardinality, sampling, retention ownership, safe names, serialization/version identity and secret exclusion with bounded deterministic tests. Confirm that no per-run byte quota rejects otherwise valid evidence, while actual write/storage failure prevents a pass. Compare compatible reference runs only.
- Actual latency and capacity assertions run in the disposable environment with the selected workload/resource envelope. Ordinary unit tests verify classifications and math.
- During implementation, run focused harness tests and the applicable static, contract, integration and delivery gates. Workflow integration must retain existing required verification; no hosted setting change is implied.
Out of Scope¶
Formal Training Run scale qualification, recommendation quality assessment, monthly availability measurement, personalization or ANN qualification, production/staging traffic, a telemetry-platform redesign, hosted infrastructure or required-check changes, deployment and package publication.
Further Notes¶
Status: the user-approved gate policy, evidence lifecycle, and manual CI job are implemented. Actual reference outcomes and verification limits are recorded in the handoff. The CI runner needs a valid baseline before automatic pull-request gating; no deployment qualification or separate design approval is asserted.
The Locust serving load-test map indexes the completed decisions. Define the serving workload and reference profile, Define repeatable CI performance gates and capacity evidence and Choose the disposable load environment and command lifecycle are resolved. No open human decision blocks these requirements.
Measurement and compatibility findings come from Verify Locust compatibility and load generation semantics. The ready-for-agent label routes agreed requirements to design and implementation; runtime verification and technical design review remain engineering work.