Serving load requirement evidence
This matrix covers all 58 user stories in the approved specifications. Story numbers refer to
the numbered lists in environment,
traffic, and
gates.
Deterministic checks prove controls and failure behavior; disposable runs prove actual preparation,
HTTP traffic, measured performance, and teardown. See the
handoff for commands, outcomes, and qualification limits.
Environment: 16 stories
| Stories |
Implementation |
Observable evidence |
| 1–2, 11 |
Controller, owned Compose, Make commands |
Full invocation, allocated loopback endpoints, unique run resources, ownership refusal test |
| 3–5 |
Preparation, dedicated profile |
Seed/cutoff manifests, verified row receipts, four independently scoped heads, announced fixture scale |
| 6–8 |
Normal worker publication and all 44 HTTP readiness checks in controller |
Actual four-scope PostgreSQL setup and public serving; wrong-head contract tests |
| 9, 15 |
Separate Docker stages, optional load extra, pytest plugin exclusion |
Delivery contracts, ordinary socket isolation test, Python runtime smoke checks |
| 10, 16 |
Resource preflight, inspected caps, finite diagnostics and aggregate observations |
Full run resource/version evidence; setup failure, five preflight command failures with safe operation/exit metadata, and required-report lifecycle tests |
| 12–14 |
Setup/request/drain/total/cleanup budgets, signal handling, verified ownership and recovery |
Completed run teardown; foreign-volume refusal, setup failure cleanup, failed recovery tests |
Serving traffic: 20 stories
| Stories |
Implementation |
Observable evidence |
| 1–5, 9 |
Workload policy, FastHttpUser driver |
Seeded selection, 44 finite labels, scope/family distribution unit tests, full reference traffic |
| 6–8 |
Source eligibility/category metadata plus exact published-head lane digests |
Real API contract tests: wrong scope, rank, score, provenance, corrupt eligibility/category, declared empty/stale |
| 10–13 |
Fixed reference/stepped/sustained populations, closed-loop phase offsets, achieved valid RPS |
Profile tests, actual user pacing regression probe, virtual-time phase checks, measured reference report |
| 14–15 |
Finite outcomes, no retry/redirect, absolute two-second timeout, bounded drain |
Actual HTTP admission/timeout smoke, controlled start-cohort and overload-policy tests |
| 16–17 |
Fixed bounded dashboard and headless runner |
Real Flask mutation hooks reject outside host/population/reset/restart; stop invalidates run; runtime smoke |
| 18–19 |
Sanitization before event listeners, fixed labels, finite counters, CPU saturation detection |
Actual request-event/error sanitization and timeout smoke; report resource observations and original-counter tests |
| 20 |
Locked Locust 2.46.6 in isolated optional runtime |
Real client smoke on macOS ARM64 Python 3.12, 3.13, 3.14; Linux ARM64 generator-image smoke |
| Stories |
Implementation |
Observable evidence |
| 1–2, 21 |
Make reference/capacity commands, manual serialized CI workflow |
Full headless invocation and existing repository gates; delivery tests preserve manual rollout, anonymous job-private Docker configuration, and recovery/upload; client configuration tests preserve context/plugins and personal settings |
| 3–5 |
Half-open request-start windows, post-population settling, full fixed denominator |
Warm-up/drain attribution, idle-tail and virtual-time phase checks |
| 6–8, 13 |
Measurement policy: exact counters, required buckets/samples, rate and completeness |
154 ms rounding regression, exact 90/150 boundaries, sparse/missing samples, under-rate and incomplete cases |
| 9–11 |
Every measured attempt counted; independent semantic validation and admission categories |
Failure-latency and zero-error tests, contract corruption tests, real sanitized HTTP events |
| 12, 14, 19 |
Lost-user checks, generator CPU warnings, required observations/reports, deadline and cleanup failures |
Controlled incomplete evidence and setup/recovery tests; measured generator/container resources |
| 15 |
Distinct stress policy; six sustained windows with reference gates |
Allowed admission stress test; bad sustained window cannot hide in overall aggregate; real phase adapter smoke |
| 16–18 |
Atomic evidence and retention, manifests, versions, aggregate CSV, sanitized diagnostics |
51 MiB report accepted without quota, fsync failure, late marker/retention failure cannot leave a passing summary, completed/failed/active/foreign retention checks; full run reports |
| 20, 22 |
Deterministic gate tests separated from measured runs; explicit qualification labels |
All gate arithmetic tests and the handoff's measured result/CI-runner limitations |
Qualification boundary
Reference is the complete measured local experiment. Stepped and sustained controls are checked
with actual client adapters and virtual time; their capacity and 30-minute performance claims
require separately selected complete runs. CI dispatch and runner qualification require a valid
runner baseline. This matrix makes no production capacity, monthly availability, ANN, personalization,
or training-scale claim.