Serving load harness design record¶
Status: implemented engineering design for the user-approved Locust specification bundle. This record describes the choices made during implementation; it does not assert separate technical-design approval or performance qualification. The run guide describes the operational contract.
Boundaries and ownership¶
flowchart LR
M[Make command] --> C[Host lifecycle controller]
C --> P[Owned preparation process]
P --> S[(Synthetic source PostgreSQL)]
P --> W[Normal TrainingWorker]
W --> D[(Control PostgreSQL published heads)]
P --> O[Private oracle and safe fixture receipt]
C --> A[Ordinary API process]
A --> D
C --> L[Isolated Locust process]
O --> L
L -->|Public serving HTTP| A
L --> E[Aggregate evidence]
C --> E
C --> T[Verified owned cleanup]
Source and preparation stop before the HTTP readiness checks. During measurement the API reads published snapshots and Locust makes ordinary public requests. Preparation uses normal training and publication contracts; no test endpoint or synchronous serving computation is introduced.
| Responsibility | Implementation |
|---|---|
| Approved profiles, scoped selection, semantic oracle | src/recommendations/load_testing/workload.py |
| Original-duration counters and interval gates | src/recommendations/load_testing/measurement.py |
| Atomic reports and owned completed-run retention | src/recommendations/load_testing/evidence.py |
| Docker preflight, deadlines, observations, ownership, recovery | scripts/load_testing/run.py |
| Anonymous CI image pulls with preserved Docker context/plugins | scripts/load_testing/docker_client.py |
| Source generation and normal worker publication | scripts/load_testing/prepare.py |
| FastHttpUser, request events, profiles, dashboard, saturation | load_testing/driver.py |
| Actual client runtime compatibility probe | scripts/load_testing/runtime_smoke.py |
Policy imports no Locust and no API/worker composition roots. Outer scripts own framework and
composition dependencies. Docker ordinary/preparation images contain ordinary dependencies; only
the generator installs the optional load extra. The normal pytest process blocks Locust's plugin
to prevent accidental gevent patches when that extra is installed.
The sibling telemetry build context remains available to ordinary service packaging. Aggregate Docker/psutil observations provide the required resource evidence; load reporting adds no production telemetry dependency or direct telemetry import outside the existing observability adapter.
Preparation and independent validation¶
Generate four scope-bound PostgreSQL schemas with deterministic seeds and an explicit aware cutoff. Two Data Source configurations expose canonical scoped UNION views for their allowed schemas. Each scope is trained once and publishes through the existing complete, atomic snapshot boundary.
Read eligible inventory/category metadata from preparation's source Catalog, then stream the published global-geography Recommendation Sets from control storage. Retain two canonical digests per lane, for limits 20 and 100, alongside the actual head and source inventory. The private oracle is bounded by four scopes, 9,500 anchors, eight anchored strategies, three global strategies, and two fixed limits. It contains derived validation data, never interaction rows.
The generator checks both independent source eligibility/category facts and exact published lane content. Thus a corrupted published lane cannot validate its own ineligible candidates. A repeated Item identifier across scopes cannot conceal the wrong head. Publication UUIDs are recorded rather than forced to deterministic values. Ordinary contract tests inject wrong heads, ranks, scores, provenance, and corrupt published eligibility/category lanes.
Timing and acceptance¶
Requests use one sequential FastHttpUser connection at a time, no retries or redirects, and a two-second absolute cooperative deadline. Timeout attempts produce safe finite events. Request metadata is sanitized before Locust listeners retain statistics: names are exactly the 44 finite scope/strategy buckets; exception/body/URL material is removed.
Declare measurement boundaries after the full population is established. Attribute outcomes by original monotonic request-start time through drain. Preserve failed attempts in latency counters, and divide valid completions by the entire declared interval. Sustained uses six independent gates. Stepped overload allows only explicitly classified admission responses.
Use exact counters over original durations for 90/150 ms acceptance. Finite Locust-style histogram bins provide labelled approximate percentiles, capped by a 3,000 ms overflow bin. No raw timing history is retained. Acceptance is integer arithmetic over threshold fractions, with required sample and throughput checks. This prevents a 154 ms observation rounded to 150 ms from passing the p99 gate.
A deterministic probe reproduced synchronized starts from a shared pacing phase: all 100 requests
started in one 40 ms slice. Initial per-user offsets now cover the 400 ms reference cycle. Subsequent
wait is max(0, 0.4 - elapsed_since_actual_request_start), preserving closed-loop pacing without
catch-up arrivals. This changes no approved users, rate target, phase lengths, or thresholds.
Failure and privacy contracts¶
A unique run/project identity and two independent ownership labels protect cleanup. Verify the complete resource inventory and image labels before any mutation. Local file locking prevents concurrent harness invocations across worktrees; CI also serializes its workflow. Loopback HTTP ports are allocated by Docker; source/control database ports are internal.
Finite deadline budgets cover setup, UI start, measured profiles, request completion, and teardown within 180 minutes. Repeated cancellation cannot bypass the bounded cleanup attempt. Recovery checks report ownership and leaves interrupted evidence failed. Missing reports, storage failure, lost users, generator saturation, unmet resource prerequisites, or incomplete cleanup cannot pass.
JSON/CSV writes use temporary files, fsync, and replacement. Report storage has no byte quota; finite producers and fixed cardinality bound memory and work. Subprocess output is suppressed except small structured queries, whose in-memory buffer is bounded separately from report storage. Credentials exist only in the private owned environment. Source/work volumes disappear on cleanup. Safe manifests omit planted interaction context identifiers.
The manual macOS CI job uses an isolated Docker client configuration for its public image pulls, builds, execution, and recovery. It references the existing context store and plugin directories, but excludes personal registry authentication and native credential helpers. This addresses a reproduced image-pull failure in the runner's separate security session without changing the runner's security settings or personal Docker configuration. The job's temporary directory owns the client configuration. Finite operation names and observed exit codes are retained for at most the primary and cleanup command failures; raw subprocess output remains suppressed.
Verification and rollout¶
Deterministic tests verify gate math, start-window attribution, error policy, retention, write failure, ownership refusal, setup failure, and recovery. Contract tests use the real public API. The isolated client smoke uses real HTTP, Locust request events, Flask dashboard hooks, and virtual time for profile/pacing checks, on Python 3.12–3.14 and the generator's Linux image.
The full reference is the actual four-scope PostgreSQL publication/serving/cleanup experiment. Passing code checks cannot establish latency. Stepped and sustained runtime phase checks establish their controls; measured capacity or sustained qualification requires their own complete runs. The initial CI workflow stays manual until the runner's valid reference baseline is established. Evidence and qualification limits are maintained in the implementation handoff.