Skip to content

Design: fixes for repository review findings

Control and scope

Field Value
Status Implementation ready for review; reduced execution measured with failed Oracle observation
Revision R3, 2026-09-30; maintainer accepted H1 and reduced Q1 execution
Code baseline 1dfb34a
Author Codex
Decision owner and reviewer Repository maintainer; implementation requested by user
Authority Initial planning request, followed by $implement implement the plan
Evidence Review record and its reproducible probes
Deliverable Implement and verify the remediation; no deployment

This design covers all 28 confirmed findings: C1–C16, S1–S7 and E1–E5. It also covers M1, S8, S9, H1, Q1 and Q2 as qualification or investigation plans. Those six concerns are not promoted to confirmed production defects. The OpenMP probe abort was disproved as a production-builder defect and requires no fix.

Current implementation evidence, completed slices, and remaining qualification are tracked in the remediation handoff. The plans below remain the full scope; completing one slice does not complete the implementation goal.

The governing contracts are Commerce Scope and domain language, service requirements, personalization, ANN lifecycle, admission design, and repository verification. Their checked-in revisions at the code baseline are the test basis. Review finding IDs are local risk identifiers, not newly approved product requirements.

Planning is complete when every finding has a change surface, an invariant owner, an implementation sequence, discriminating verification, and applicable compatibility and operations treatment. Implementing these plans and qualifying production are subsequent work. Draft recommendations below do not manufacture approval of new API semantics, retention policy, deployments, or capacity claims.

Shared decisions and verification rules

  1. Keep all reads and writes scoped by Data Source, Tracking ID and Catalog ID. Shopper state adds its existing Shopper key. Serving remains source-independent.
  2. Preserve exact ranking and evidence definitions. Sampling, score perturbations, approximate replacement of exact metadata ranking, and silent cohort reweighting are consequential contract changes, not shortcuts to a performance fix.
  3. Put invariants inside existing repositories, selection routines and value objects. Introduce a port or state object only where it hides a real consistency, lifecycle or backend boundary. The reusable limiter never imports recommendations.
  4. For each verification ID below, promote the synthetic scratch counterexample to the named test, run it before the fix, record the specified red signal, implement that behavior, then rerun the same case and its existing file. Test names below are proposed names, not claims that tests already exist or pass.
  5. Use fixed aware clocks, explicit seeds and small literal or independent reference oracles. Coordinate races with barriers or controlled iterators, never sleeps. Use actual disposable databases for transaction claims. Assertions about SQL count, body consumption or object residency are appropriate when that bound is the contract.
  6. Use make test-focused TEST=path::test_name for one case, or TEST=path for its existing file. The target accepts one quoted path; do not pass multiple filenames. Before Python toolchain work, obtain the module's PyCharm environment as required.
  7. Finish each implementation slice with make test. Add make migration-check for migrations, make smoke for installed-package changes, and make compose-check for Compose changes. PostgreSQL evidence requires an explicitly disposable TEST_CONTROL_DATABASE_URL and make test-postgres; ordinary SQLite tests do not prove PostgreSQL locking. Final integrated release candidate uses make verify.
  8. Benchmarks run in fresh processes with fixed workload, backend/thread budgets and correctness signatures. Retain repetitions, median and tail times, peak RSS, traced allocations where useful, temporary disk, environment and dependency versions. Do not assert arbitrary timing limits in ordinary unit tests. Assert structural bounds there, then evaluate deployment SLOs with controlled measurements.

Dependency order and useful patterns

Slice Findings Main responsibility and pattern Depends on
A C6 Fenced publication unit of work None
B C7, C1 Consistent source adapter and snapshot-bound repository read None; C1 benefits from A
C C10, C8, C9, S3, H1 Monotonic state header, retention policy and resumable maintenance C10 before maintenance changes
D C2, C3, C5, S5, S8 Feature-block composition and bounded deterministic selection C2 before C5; C3 before optimizing candidates
E C11, C13, S7, C12, M1 Bounded load scheduler, artifact validator and retrieval selection Correctness before retrieval tuning
F C4, S1, S2 Typed policy resolution and indexed permit lifecycle C4 before state changes; S1/S2 together
G S6, C14, C15, C16 Transport guard and canonical configuration policy None
H E1–E5, S9 Evaluation-unit validation and exposure identity/state adapter E3/E2 together; E5 before report changes
I Q1, Q2 Capacity evidence and dependency rules Qualification after relevant fixes

These are cohesive implementation slices, not a requirement to land one large PR. Split migrations from consumers where additive rollout permits it. Each finding below has its own observable acceptance signal even when it shares a slice with others.

C1 — One snapshot for For You candidates and fallback

Owner and surface: storage.py:SnapshotRepository, serving.py:SnapshotReader/ForYouCandidateLoader; ordinary lane loading remains separate. REQ-012, REQ-018 and the personalization one-head contract govern consistency.

Plan: Add one bounded repository operation returning a snapshot, the deduplicated global union and its Best Sellers fallback. Resolve head identity and required stored lanes through one SQL statement, or a genuinely repeatable read view if multiple statements are necessary. Build the returned immutable bundle only from that identity. The union stays at the negotiated candidate bound; the fallback stays its own lane. Extend the consumer-owned reader protocol and adapt existing reader fakes. Map absent head, missing required fallback and corrupt payload to the existing classified failures.

Verification V-C1: Add test_for_you_uses_one_snapshot_during_publication in tests/unit/test_serving.py. A faithful reader advances from A to B between the old operations; expected literal IDs and disclosed identity must all belong to A or all to B. RED: A metadata with B fallback. Add a real repository integration case in test_personalization_storage_candidates.py racing publication and retention, including a missing fallback and a valid empty union. No merchant-source calls are allowed.

Compatibility and cost: New internal repository operation, no public response or migration. Prefer one round trip over a request-long transaction; merely comparing snapshot IDs and retrying without a bound is rejected. Update technical implementation and personalization ownership docs. Benchmark ordinary versus For You query count and p95 after correctness, using the same database/lanes.

C2 — Similar Items tolerates absent usable feature blocks

Owner and surface: pipeline/content.py:_has_metadata/_metadata_matrix/build_similar_items. The valid insufficient-evidence outcome in REQ-018 must cover feature-poor Catalogs.

Plan: Define feature availability by what the encoder can actually emit. Positive price is a usable price feature; zero remains valid Catalog data without creating a price block. Normalize text presence once, then determine word and character vocabulary availability independently. Keep usable structured features when one text block is empty. Do not catch arbitrary vectorizer failures as insufficient evidence. Return an explicit no-feature result or a correctly shaped empty matrix and bypass multiplication when no positive features exist. Keep candidate keys for every Catalog Item.

Verification V-C2: Parameterize test_similarity_handles_unusable_metadata in tests/unit/test_content.py with two zero-price Items; whitespace text; punctuation; one-character text; valid category plus punctuation; and a mixed Catalog. RED: current IndexError/empty-vocabulary ValueError. Oracle: empty results without features, and same-category neighbors when structured features remain. Add an end-to-end Training Run in test_end_to_end.py proving complete publication with valid empty metadata lanes, and failure preservation for a genuinely unexpected encoder error.

Compatibility: No dependency or schema change. Character-only usable features may produce recommendations where the old code crashed; disclose that in the technical docs. Compare scores for ordinary usable metadata within existing numerical tolerance. Implement this before C5 and before touching ANN representation reuse of these features.

C3 — Filter eligibility before bounded generation

Owner and surface: pipeline/run.py:SignalBuilder/_score_candidates, both workstore.py and polars_workstore.py behavioral queries, provider assembly.

Plan: Pass one immutable Catalog eligibility view to candidate generation. In behavioral queries, filter candidate endpoints before ROW_NUMBER/top-N; retain historical support and pair counts as currently defined rather than discarding their source rows. Apply equivalent filters in Polars and before popularity sorting/selection. Preserve anchor coverage for the complete strategy manifest; an ineligible candidate cannot occupy a retained slot. Check category and geographic variants through the same policy. Keep the final publication filter as an invariant guard, not the primary selection step.

Verification V-C3: test_eligible_candidates_survive_generation_cutoff in tests/unit/test_workstore.py: A has 200 higher-ranked ineligible neighbors and lower eligible C with qualifying support. RED: C missing. Compare to a small exhaustive filter-then-sort reference. Add equivalent Polars, global popularity and category/local cases with zero, exactly K, and more than K exclusions. test_end_to_end.py must observe correct primary provenance and avoid fallback when eligible evidence exists.

Compatibility and performance: This corrects recommendation contents and evaluation metrics; bump model identity through the normal release policy so corrected output is not drift-compared as the same implementation. No source-row staging. Inspect query plans and joins at representative Catalog sizes; benchmark high exclusion fractions. Update algorithm and eligibility documentation, not the definition of pair support.

C4 — Distinguish malformed policy resolution from unmatched traffic

Owner and surface: serving_limits/core.py:LimitPolicy/AdmissionController, ASGI translation only if a new reason is exposed. This remains framework-independent.

Plan: Represent resolution as a closed result with matched limits, no selector match, or missing required partition. Permit allow_unmatched only for no selector match. A matching malformed rule denies before state acquisition; parent limits must not be consumed. Reuse the existing unavailable response where compatible, otherwise add a reviewed stable missing-descriptor reason and map it to deployment-unavailable 503. Do not expose the descriptor's raw value in errors or metric labels.

Verification V-C4: test_missing_partition_fails_closed_with_allow_unmatched in tests/unit/test_serving_limits.py, for both flag values and with a global parent rule. RED: allowed under True. A second request with correct dimensions proves the denied request did not debit its parent. Preserve legitimate unmatched allowance and matching requests. Add ASGI/API contract assertions for status, body and absent lease.

Compatibility: A formerly allowed malformed request becomes denied, intentionally. Document the flag's scope in the admission design/config example. No database change. Do this before S1/S2 so indexed state never compensates for malformed policy identity.

C5 — Exact lexical ordering at the metadata top-N cutoff

Owner and surface: pipeline/content.py bounded selection boundary; DEC-7/REQ-019.

Plan: Give exact selection the total order score descending, Product ID ascending, with self excluded before selecting K. First check whether the pinned native kernel supports an explicit secondary comparator; do not assume input sorting controls heap ties. If it does not, use bounded row/column sparse-product blocks and a per-anchor K-entry selector with that comparator. Bound both block axes using the existing work resource budget. Create candidate objects only for survivors. Keep exact scores; no epsilon perturbation, full N-by-N matrix or approximate ANN substitution.

Verification V-C5: test_similarity_cutoff_keeps_lexical_tie_winners in tests/unit/test_content.py: 210 identical-category vectors, K=200, anchor 000. RED: first returned ID 009 instead of 001. Require literal 001–200 where applicable, shuffle input order and vary native threads/block sizes. A small exhaustive cosine reference checks mixed ties and near-equal untied scores; self and zero-score edges must retain established behavior.

Trade-off and qualification: Exact dense-overlap similarity can still require quadratic compute; bounded storage alone does not prove scalable runtime. Measure sparse and highly overlapping Catalogs, report peak buffers and throughput, and reject a default algorithm that cannot meet the agreed training budget. A native tie-comparator extension is preferable if block fallback is too slow, but is a separate dependency decision with parity tests. Update model identity and algorithm docs after qualification.

C6 — Fence every publication mutation by its lease generation

Owner and surface: storage.py:TrainingRunRepository/SnapshotRepository and worker.py claim/publication handoff; migration and lifecycle tests. High assurance.

Plan: Carry an immutable claim token containing run ID, owner and generation (the existing recovery_count can supply the run's generation). Add publication-generation metadata to building snapshots if needed to distinguish abandoned staging. For creation, every set/feature batch, activation and cleanup: lock the run first, check current owner, generation and lease using current authoritative time, then mutate only that generation within the same transaction. Use a consistent run-then-snapshot lock order. Lost owners raise the existing lease-lost error before deletion or insertion. Cleanup may discard its own snapshot ID; it must never search-and-delete the successor's build.

Verification V-C6: test_recovered_publisher_survives_stale_publisher in tests/integration/test_lifecycle.py promotes the nested-generator reproduction with foreign keys enabled. RED: successor IntegrityError. Coordinate stale writes at build creation, set batch, feature batch and activation. Assert successor completes, old owner gets lease-lost, one complete head exists, and old cleanup cannot delete it. Mirror with two real PostgreSQL connections in test_postgresql_storage.py; also test worker identity reuse and failures at each transaction boundary.

Migration and rollout: Add fields/indexes additively, backfill existing available snapshots without changing heads, and classify pre-upgrade building snapshots as abandoned only under a valid claim. Drain old worker binaries before enabling new publication; old writers cannot obey new fencing. Run migration round trips and make test-postgres. On regression, stop new publication while serving the last head; do not roll back to destructive unfenced writers. Update operations and implementation docs with fencing and lock ownership. No long transaction across full snapshot staging.

C7 — Genuine repeatable SQLite source reads

Owner and surface: source.py:SqlAlchemySourceAdapter/SourceReadSession, worker.py:build_source_adapters; REQ-009, REQ-014 and consistent-read boundary.

Plan: Centralize source-engine construction and explicitly configure sqlite3 transaction control. Supported Python versions are 3.12–3.14, so autocommit=False is available; qualify it with isolation-level handling. Alternatively use a single reviewed SQLAlchemy BEGIN event strategy. Never mix both mechanisms. Keep WAL setup outside active transactions and read-session rollback/connection cleanup deterministic. Document the requirements for directly injected engines; reject unsupported SQLite transaction settings rather than advertising snapshot consistency for them.

Verification V-C7: test_sqlite_source_read_is_repeatable_across_queries in tests/contract/test_source_adapter.py: temporary WAL file, two connections, read A, commit B on writer, read again within the same source context. RED: B appears despite unchanged token. Oracle: reader sees A throughout, then a new context sees A+B. Exercise Catalog-to-interaction reads, transaction cleanup after stream failure and writer progress. Run the supported-Python matrix and PostgreSQL canonical-read case.

Compatibility: SQLite writers may encounter locking behavior hidden by legacy mode; qualify busy timeout and WAL, rather than retrying a source stream mid-snapshot. No database schema migration and no merchant schema changes. Preserve dialect defaults where their driver actually provides the documented semantics. Update adapter onboarding and technical implementation; configuration must fail before acquiring long-lived resources.

C8 — Enforce retention per Commerce Scope

Owner and surface: personalization/capability.py, config.py, storage.py:ServiceStore composition, personalization/state_store.py, API and worker roots.

Plan: Supply the repository with an immutable scope-to-retention resolver derived from configured capabilities, with the existing maximum 90 days and explicit fallback for previously configured state. Resolve once per operation. Use it for interaction expiry, idempotency expiry, projection boundary and maintenance eligibility. Both API and worker must receive the same policy, including when personalization admission is disabled but maintenance continues. Avoid replacing the entire repository with one global duration when different scopes have different policies.

Verification V-C8: test_configured_scope_retention_reaches_persistence in tests/integration/test_personalization_storage.py: scopes configured for 30 and 90 days, same synthetic event time. RED: both ledger expiries at day90. At day31, the 30-day scope has no usable signal and the other retains it, with monotonic headers per C10. Verify the actual deployment composition path, not parsing alone. Include disabled capability maintenance, idempotency expiry and unknown policy handling.

Rollout: For shortened policies, existing expiry timestamps require bounded reconciliation using occurred_at plus current scope policy. Define its progress and exposure barrier with S3; do not wait for old 90-day timestamps. Increasing retention does not reconstruct deleted data. Version configuration identity where policy changes alter evidence. Confirm metadata-retention treatment in C10; document shortened-policy rollout and one-sided rollback (erased history cannot be recovered).

C9 — Do not project events outside the allowed history window

Owner and surface: personalization/health.py, interaction service and repository projection; maximum-history contract. Depends on the retention resolver and C10 headers.

Proposed behavior: Preserve durable next-sequence acknowledgement for a valid late business interaction but omit its expired feature contribution. This avoids breaking ordering simply because event occurrence is old. Record only the minimal idempotency and acknowledgement state required for replay; do not insert an already expired feature-ledger payload. Lifecycle transitions still apply immediately according to authorization and commit time; an old timestamp must not neutralize opt-out/deletion. The alternative is a typed late-event rejection before any commit. The maintainer owns this public-behavior choice; the design recommends acknowledgement without stale signal.

Verification V-C9: test_expired_interaction_does_not_affect_profile in tests/integration/test_personalization_storage.py: event exactly at boundary, one microsecond before and after, under 30/90-day policies. RED: expired affinity present. For the proposed choice, assert next sequence advances once, exact retry replays the same acknowledgement, profile signals are unchanged, and causal read observes the commit version. Include late lifecycle commands and rejection under absent authorization.

Documentation and risk: Update Interaction API late-event behavior, maximum history, idempotency and causal examples together. A mere health-layer check is insufficient because the repository is an independent persistence boundary. Preserve safe error shape and never include Shopper IDs or rejected payloads in metrics/logs.

C10 — Separate monotonic acknowledgement state from disposable history

Owner and surface: personalization/state_store.py, state metadata/migration in storage.py, profile read/causal service and deletion reconciliation. High assurance.

Plan: Make a minimal per-Shopper state header authoritative for last committed sequence, acknowledged version and lifecycle barrier. The existing profile row can retain this header with an empty payload, or an additive watermark table can hold it; recommend the empty-profile header first unless bounded maintenance needs a separate publication version. Stop deriving these fields from the last retained feature event. Expiry clears/rebuilds features but preserves committed sequence; a changed projection gets a version no lower than any acknowledgement, and preferably increments for its own publication. Return insufficient-history fallback for an empty active profile.

Verification V-C10: test_expiry_preserves_acknowledged_watermarks in tests/integration/test_personalization_storage.py: recent seq1, expired seq2, then maintenance and seq3. RED: version/watermark2 becomes1 and seq3 rejects. Test full expiry, retries before/after idempotency expiry, prior valid causal tokens, concurrent append/rebuild and suppression/deletion. A request begun after the acknowledged commit must never observe a lower header version, even with no features.

Compatibility and privacy: Keep feature history at the configured maximum. The header contains no interaction payload, but still contains a scoped identity; document its lifecycle and minimal retention basis rather than assuming indefinite retention is free. Reuse the existing deletion tombstone semantics and avoid resurrecting erased profiles. If a new table is selected, backfill maxima from profile, suppression and retained idempotency rows in bounded batches; already-lost watermarks require an explicit recovery procedure, not fabricated sequences. Run migration and PostgreSQL race gates.

C11 — Reap ANN background work independently of the requested head

Owner and surface: ann.py:AnnSnapshotRetriever, API lifecycle if executor cleanup is added.

Plan: At entry to the locked scheduler, reap all completed pending entries before testing capacity or checking the requested identity. Completion handling stores a valid result or bounded negative-cache entry with the existing resident limit. Do not hold the lock while doing artifact I/O, and do not wait on an unfinished Future in a serving request. A completed old generation cannot prevent a new one from being scheduled. Keep one loader and bounded pending/resident state; add deterministic close behavior for shutdown if the scheduler acquires lifecycle responsibility.

Verification V-C11: test_ann_completed_old_load_does_not_block_new_head in tests/integration/test_ann_snapshot.py: controlled reader and Event/Future, finish A, request B without revisiting A. RED: reader never receives B. Repeat for success, missing/corrupt artifact, raised failure, another Commerce Scope, and a running load that legitimately consumes the slot. Assert serving falls back immediately, no duplicate load exists and residency/pending counts remain bounded.

Compatibility: No artifact schema or response change. Optional ANN disabling still provides its current snapshot fallback. Measure cold-request latency and turnover; avoid unbounded queues as a starvation fix. Document scheduler/residency ownership.

C12 — Separate purchase exclusion history from neighborhood seeds

Owner and surface: swimlane_service.py:resolve_swimlane context assembly.

Plan: Keep the complete bounded recent-purchase tuple from the profile for exclusion. Derive a separate first-ten tuple only for retrieval neighborhood seeds. Pass explicit full exclusion sets into configured steps requesting exclusion; do not exclude recent purchases on steps where policy does not request it. Preserve recent-view terminal fill and ordered step provenance. Name these fields by responsibility so another retrieval budget cannot silently become an exclusion budget.

Verification V-C12: test_swimlane_excludes_all_retained_recent_purchases in tests/unit/test_swimlane_service.py: 11 and 100 purchases, a Best Sellers list containing the eleventh purchase and a fresh eligible Item. RED: purchase-10 returned. Assert only the fresh Item survives where exclusion is enabled; disabling exclusion retains the prior lane behavior. The bounded reader records at most ten purchase seed reads. Include deduplication across steps and missing seed metadata.

Compatibility: No schema, query or public shape change. Corrected selections may reduce fill when all candidates were purchased; retain truthful insufficient fill rather than expanding history or source-reading. Update Swimlane examples and rerun the geographic/Swimlane integration path.

C13 — Validate native ANN distance semantics

Owner and surface: ann.py:LoadedAnnIndex artifact validator, ANN lifecycle contract.

Plan: Validate metric_type against the representation's declared inner-product metric before admitting the loaded index. Require native dimensions, entry count, eligible mapping and representation version to agree as already defined. Bind metric identity into any new artifact envelope version/checksum if it is not explicitly recorded; infer supported legacy metric only from the validated native index, never from a permissive default. An incompatible artifact follows the existing safe fallback.

Verification V-C13: test_ann_rejects_incompatible_native_metric in tests/integration/test_ann_snapshot.py: create a checksummed L2 HNSW artifact with otherwise valid shape/IDs. RED: load succeeds. Require typed validation failure and snapshot fallback in serving. Valid inner-product artifacts must preserve native/exact similarity semantics, eligible filtering and same-scope binding; exercise learned and metadata representation paths.

Compatibility and rollout: Existing legitimately incompatible artifacts will stop loading and need rebuilding; the published snapshot fallback remains available. If envelope metadata changes, use an additive migration and explicit representation version, then deploy readers before writers. No best-effort reinterpretation of L2 as cosine. Update the lifecycle spec's test mapping and artifact troubleshooting guidance.

C14 — Validate run IDs before telemetry construction

Owner and surface: api.py:inspect_training_run, transport contract tests.

Plan: Validate bounded path input before constructing TelemetryContext or querying persistence. Use a bounded validator appropriate to existing run identities; avoid requiring UUID syntax if current repository callers legitimately use other IDs. Return the reviewed invalid-identifier 422 response; retain 404 for a valid unknown identity. Reject credential/URL/control-shaped values without echoing them. Keep only bounded route/outcome observations for rejected input.

Verification V-C14: test_inspect_run_rejects_invalid_identifier_before_telemetry in tests/contract/test_api.py: 256 characters, control characters through proper URL encoding, and credential-shaped text. RED: HTTP500. Assert 422, safe error envelope, no rejected content in captured telemetry and no repository access. Valid unknown ID remains404; existing submitted run remains200. Test enabled and disabled observability.

Compatibility: This intentionally changes malformed requests from500 to4xx. Update OpenAPI through its generator, not generated output edits, and document the accepted bound. No durable-state migration or blanket narrowing of unrelated identifiers.

C15 — Make configuration identity match runtime parameter selection

Owner and surface: config.py:DeploymentConfig.config_version, pipeline/run.py popularity/trending parameter selectors; comparable-run evidence.

Plan: Treat search grids as sets, matching their current sorted fingerprint. Canonicalize candidates once and define a total tie order: metric objective first, distance from existing preferred defaults second, then the canonical parameter tuple. No-holdout fallback selects the default if available, otherwise the same deterministic closest candidate and final tuple order. Apply it to all relevant popularity/trending grids and check behavioral-search ties, without changing the primary metric objective.

Verification V-C15: test_parameter_selection_is_invariant_to_grid_order in tests/unit/test_config.py or a focused selector test module: [2,8] versus [8,2], no holdout, no default and tied holdout objectives. RED: same hash, half-life2 versus8. Enumerate permutations for small multidimensional trending grids and compare selected parameters/output, not merely fingerprints. Verify genuinely different grid sets still have different versions and ordinary default selection remains characterized.

Compatibility: Canonicalization is preferred to preserving accidental caller order in the hash. Some previously ambiguous grids change output; bump implementation/model identity and document tie policy so historic runs are not falsely comparable. No hash algorithm migration is required if the set semantics and payload are unchanged.

C16 — Reject nonfinite polling intervals during configuration

Owner and surface: config.py:DeploymentConfig.__post_init__/from_mapping.

Plan: Require a finite positive numeric interval and reject booleans as a numeric interval if the existing parsing helpers permit them. Apply validation to direct construction as well as serialized config. Reuse the repository's finite-positive helper where it actually fits; no general validation framework is needed. Failure must happen before opening repositories or starting the worker loop, with a safe configuration error naming the field.

Verification V-C16: test_worker_poll_interval_must_be_finite in tests/unit/test_config.py: NaN, positive/negative infinity, zero, negative and a valid fractional value; include the accepted JSON parser's NaN form. RED: NaN constructs a config and reaches a failing sleep. Assert all invalid inputs fail construction, the valid interval survives, and no runtime sleep is needed in the test.

Compatibility: Previously malformed config fails early instead of crashing idle workers. No schema or public endpoint change. Update configuration reference and doctor guidance; rerun existing configuration and worker lifecycle tests.

S1 — Bound rate-bucket residency without restoring spent tokens

Owner and surface: serving_limits/core.py:InMemoryLimitState, optional serialized state-bound configuration and denial translation. Token semantics remain unchanged.

Plan: Store a bucket's next semantically safe removal time: the instant at which it would be fully replenished. Maintain an indexed expiry structure with one live entry per bucket; refresh it in place when tokens are consumed. Reclaim only fully refilled buckets. Add an explicit maximum resident bucket count, validated positive and independently qualified for deployment. If allocation would exceed the bound, deny the new partition without mutating any parent's rate/concurrency state. Expose capacity exhaustion as unavailable, with bounded reason metrics, rather than silently evicting a depleted bucket or allowing unlimited partitions.

Verification V-S1: test_idle_rate_buckets_are_reclaimed_without_resetting_debt in tests/unit/test_serving_limits.py: spend bucket, request just before/full refill, advance a fixed clock and touch another partition. RED: all10,000 idle buckets remain. Assert no early capacity restoration, bounded residency at saturation, correct refill after removal/recreation, and atomic parent/child denial. The expiry index itself must remain bounded after repeated refreshes; a lazy heap with unlimited stale nodes fails.

Trade-off and operations: A hard cap can deny an otherwise affordable new consumer, so the default/capacity policy is a maintainer-owned deployment choice. Recommend explicit local-state configuration and503 on state exhaustion. Benchmark high churn, few hot keys and version rollovers; retain the current single-process limitation. Document why ordinary LRU eviction is unsafe for depleted token buckets.

S2 — Index concurrency expiry and retry capacity

Owner and surface: the same state adapter as S1; no distributed enforcement claim.

Plan: Replace full _leases reconstruction with an expiry index and direct lease-to-permit ownership. Index permits by their bucket for retry calculation and release. Remove each expired permit once, decrementing its own bucket idempotently. Use a bounded indexed structure; explicit release removes entries rather than leaving unlimited tombstones. Weighted retry selection must use bucket-local expiry/cumulative cost, not sorting every lease in the process. Preserve partial expiry when one lease contains rules with different TTLs and atomic all-rule acquisition under the lock.

Verification V-S2: test_admission_work_does_not_scan_unrelated_live_leases in tests/unit/test_serving_limits.py: increasing synthetic lease populations and an unrelated bucket request. RED: operations scale with all live permits. Through public acquire/release, compare decisions, retry delays and permit conservation to a simple small reference using a fixed clock, varied costs/TTLs and repeated releases. Include same-deadline expiry bursts, denial storms and generated duplicate lease IDs.

Performance acceptance: Each nonexpiry operation should avoid global Theta(P) work; ordinary admit/release should be O(R log P), for R matched rules and P permits, plus expired entries actually processed. Total creation/expiry work for N retained admissions should be O(N log N), replacing the measured quadratic scan. Expiry bursts still cost work proportional to expired permits; explicitly cap state and qualify tail latency rather than claim constant worst-case time. Repeat the1k/2k/4k probe and larger bounded profiles, recording signatures, residency and lock duration.

S3 — Resumable bounded personalization maintenance

Owner and surface: personalization/state_store.py:expire/_rebuild_profile, worker maintenance orchestration; depends on C10 and scope retention C8.

Plan: Separate logical signal expiry from physical deletion. Use keyset pagination for affected keys and deterministic row limits for delete batches; eliminate global .all() and unbounded all-key transactions. Rebuild one projection through a streamed sequence reducer with bounded feature state. For a single very active Shopper, support checkpointed rebuild chunks or a time/work budget; merely limiting keys is insufficient. Publish rebuilt features only if the source/header version still matches, otherwise retry with a bound. Maintain a durable dirty/rebuild marker if deletion precedes final projection publication, so interruption cannot strand stale features as usable.

Verification V-S3: test_expiry_batches_are_resumable_and_preserve_causal_state in tests/integration/test_personalization_storage.py: multiple expired keys, one large key, retained/new interactions and forced interruption after a delete batch. RED: all keys/rows fetched in one transaction or stale projection lacks a resumable marker. Assert batch row/key/work limits, eventual exact reference projection, monotonic headers, isolation, and no lost concurrent append. PostgreSQL tests must measure lock duration and race CAS publication; SQLite alone is inadequate here.

Migration and operations: Add only necessary progress/dirty fields or a bounded work table, with indexes on scope/expiry/key/sequence. Writers maintain invalidation markers transactionally. Serving omits logically expired or dirty signals and returns documented safe fallback, preserving causal acknowledgement header visibility. Report aggregate backlog, batches and duration without Shopper labels. Rollback must retain dirty-state protection; pausing physical cleanup must not make expired signals usable. Qualify bounded maintenance progress under shortened retention and continuous ingestion.

S4 — Stream exact holdout evaluation instead of retaining the pair graph

Owner and surface: pipeline/workstore.py, polars_workstore.py, pipeline/evaluation.py, run.py parameter selection and WorkEvidence handoff.

Plan: Introduce a streamed holdout evaluator over derived, deduplicated directional edges ordered by anchor. For each anchor, keep predicted top-K membership/rank (bounded), count distinct relevant neighbors, hits and DCG contribution while streaming truth. Compute ideal DCG from min(K,relevant_count); accumulate cohort means and coverage using their existing definitions. Do not build relevant-neighbor sets for every anchor or materialize them again in retained WorkEvidence. Keep the current pure evaluator as the small-data reference. Both stores implement the same narrow evaluation contract.

Verification V-S4: test_streamed_holdout_matches_exact_reference_with_bounded_state in tests/unit/test_workstore.py and Polars equivalent: tiny literal graphs cover duplicates, empty predictions, sparse/cold cohorts, coverage and temporal partitions. RED structural evidence: full pair fetch/materialization. A growing dense graph must not grow Python truth state with total edges; compare exact aggregate metrics at different fetch sizes and parameter-grid reuse. Add end-to-end evidence parity and stream failure/cleanup tests.

Trade-off: Exact streaming is selected over sampling to preserve REQ-013 and evidence schema meaning. It may scan truth more often during parameter search; reuse derived queries or bounded batched candidate evaluation rather than cache the entire graph. Record CPU, peak RSS and spill at high edge density, maintaining no raw-row staging. Update evaluation flow docs; schema version changes only if metric semantics change.

S5 — Share bounded complementary-category candidates

Owner and surface: pipeline/run.py:_category_pair_candidates and provider assembly; reuse pipeline/fallbacks.py:KeyedCandidateProvider.

Plan: Sort each destination category's eligible Items once by the established purchase-score/Product-ID order, retaining the existing generation bound. For each source category, compute its paired-category sequence once with the current support ordering and truncate after the required candidates. Share immutable category-keyed candidate tuples across anchors through the existing keyed provider; avoid constructing identical RecommendationCandidate objects and lists for every Item. Preserve the current ranking/provenance order unless C3 separately requires eligibility correction.

Verification V-S5: test_category_pair_candidates_are_shared_without_changing_rankings in tests/unit/test_category_strategies.py: two categories with several anchors, support ties, purchase-score ties and multiple paired categories. RED structural evidence: the destination category is sorted once per anchor. Compare complete outputs to a small exhaustive reference, across input permutations, evaluation/published partitions and missing categories. Do not use object-identity assertions as the only oracle; bound sort/construction work as additional evidence.

Performance: Destination sorting should occur once per evidence partition/category; category-equivalent anchors should cost an O(1) provider lookup plus unavoidable final output work. Measure 1k/2k/4k balanced-category workloads, candidate allocations and total snapshot output, with exact signatures. No public schema or migration change. Update implemented data-flow docs and rerun geographic/category end-to-end tests.

S6 — Enforce request bytes before JSON parsing

Owner and surface: a small ASGI transport guard assembled in api.py, deployment configuration, contract tests and API error documentation.

Plan: Add a positive finite byte limit per supported POST body class, with one reviewed default sufficient for the maximum legitimate interaction and Swimlane payload. Count actual bytes in wrapped receive; reject a declared oversized body before reads and a streamed body immediately when cumulative bytes exceed the limit. Apply it regardless of Content-Length, chunked transfer, whitespace padding or invalid JSON. On limit violation, return one stable413 before endpoint/state execution; ensure middleware unwinding does not also send500. Preserve protected-route authentication before body consumption and keep safe request observations.

Verification V-S6: test_request_body_limit_stops_chunked_input_before_parse in tests/contract/test_api.py: exact limit, limit+1, oversized Content-Length, no length, incorrect length, multiple chunks, unauthorized protected request and legitimate maximum payload. RED:1/8MiB padded JSON accepted. An actual ASGI receive counter proves input is not drained without bound after rejection; persistence remains unchanged. Include disconnect and cancellation before/after the boundary, with exactly one response.

Compatibility and operations: Proposed limit is an explicit configuration contract; the maintainer sets the numeric default from legitimate payload bounds, not benchmark guesswork. Deploy/document413 and proxy settings together, while the app remains its own enforcement boundary. Qualify memory under concurrent bounded requests; input caps do not replace admission concurrency or downstream timeout controls. Never retain body bytes or auth headers in diagnostic evidence. Add OpenAPI error responses and config tests.

S7 — Bounded exact ANN selection under dense ties

Owner and surface: ann.py:LoadedAnnIndex._exact; score semantics remain unchanged.

Plan: Reuse its fixed vector chunks and maintain only the best K eligible, nonexcluded (score, Product ID) positions, with a deterministic bounded selector. Filter exclusions before retention and instantiate AnnMatch only after final selection. A per-chunk partition followed by explicit lexical threshold selection is also viable if its temporary storage is bounded by chunk size; ensure it includes lexical winners among all ties in that chunk. Avoid a Catalog-sized Python match list and sorting it. The score array can be removed once streaming selection proves parity; do not conflate that optimization with a change to native ANN behavior.

Verification V-S7: test_ann_exact_ties_bound_matches_and_keep_lexical_order in tests/integration/test_ann_snapshot.py:10k equal vectors, K100, exclusions including early lexical winners, learned negative scores and metadata zero scores. RED:10k match objects for100 results. Compare selected IDs/scores to a small exhaustive reference; vary chunk sizes and exercise the native-underfill fallback through public search.

Acceptance and risk: Additional Python selection state is O(K) plus bounded chunk buffers; matching object count is at most K. Repeat10k/100k allocation and latency measurements, including untied and exclusion-heavy workloads. Native vector/index residency is separately budgeted; no promise of O(K) total loaded-index memory. Update ANN performance docs and preserve current representation-specific ranking.

E1 — Count independent evaluation units and reject duplicates

Owner and surface: personalization_evaluation.py:PersonalizationEvaluationContract.

Plan: Define one independent observation per ShopperEvaluationKey and strategy for a given evaluation invocation/split. Reject duplicates before metric accumulation, including same identity with altered cohort labels. Report suppression counts from distinct permitted units, not replayed rows. For a report that combines strategies, keep per-strategy denominators and count each Shopper only once for privacy suppression; do not silently collapse differing strategy outcomes into an unspecified mean. The maintainer owns whether multi-strategy top-level reports are allowed; default to one strategy per report until its aggregate semantics are explicit.

Verification V-E1: test_duplicate_evaluation_units_cannot_bypass_suppression in tests/unit/test_personalization_evaluation.py: one observation repeated10 times at minimum10. RED: available count10. Require a clear duplicate-input failure, or the documented suppressed distinct-unit result if idempotent deduplication is chosen. Test conflicting duplicate labels,9/10 distinct Shoppers, same Shopper in different Catalogs, feedback exclusions and cross-property rejection. Use literal Recall/NDCG oracles; no raw identities in the report.

Compatibility: Offline helper has no live callers, so strict duplicate rejection is recommended over best-effort bias correction. O(U) validation state for U declared units is acceptable only under explicit evaluation input bounds; large deployments use streamed external grouping, not an unbounded service registry. Document the unit and version evidence meaning when observation_count semantics change.

E2 — Attribute only to the exact exposed assignment

Owner and surface: personalization_evaluation.py:ExperimentExposureGuard; implement alongside E3's identity encoding.

Plan: Include variant in the canonical exposure identity and compare the complete stored assignment to the requested assignment before outcome attribution. Experiment, config version, unit type, property and unit key must also match. Preserve first-exposure immutability and timezone-aware exposure-before-outcome ordering. Record two legitimately distinct assignments separately only if the experiment contract permits it; do not use one variant's exposure as authorization for another.

Verification V-E2: test_outcome_requires_exact_variant_exposure in tests/unit/test_personalization_evaluation.py: record holdout exposure, request personalized attribution. RED: accepted. Require absent-exposure or assignment-mismatch failure, then expose the requested variant and prove valid attribution. Test version, unit, property, outcome timing and repeated identical exposure. Stored first exposure must not be replaced by a conflicting assignment.

Compatibility: Previously invalid offline attribution now fails. No production database migration because the current registry is local memory. Future durable state must enforce uniqueness on complete versioned identity, not just a mutable dictionary lookup. Document the change with E3; old ambiguous records cannot establish a new exposure.

E3 — Canonical versioned HMAC material

Owner and surface: StableExperimentAssigner and exposure correlation helper in personalization_evaluation.py. Preserve SHA3-512/HMAC-SHA3-512 policy.

Plan: Encode the identity as a canonical ordered JSON array or explicit length-prefixed UTF-8 fields, with a domain/version prefix distinguishing assignment from exposure. Choose JSON array encoding to reuse existing deterministic JSON conventions. Include every identity component; exposure also includes variant per E2. Hash exactly those bytes and never concatenate unescaped delimiter-separated values. Do not add fallback verification of the old ambiguous format.

Verification V-E3: test_experiment_identity_encoding_is_unambiguous in tests/unit/test_personalization_evaluation.py: property a:b / experiment c versus property a / experiment b:c. RED: the second identity inherits the first exposure. Test delimiter/Unicode values in each field, distinct variants/units/versions and deterministic repeat assignment. Validate encoding with fixed canonical byte literals and independently calculated published synthetic HMAC vectors; never include real keys.

Compatibility and rollout: Assignment buckets can change under a new encoding. Introduce an explicit assignment-algorithm/config version, start a new experiment cohort and do not pool exposures/outcomes across versions. Offline evidence replay must regenerate assignments from authorized synthetic inputs or label prior evidence incompatible. This change is not a rotation of production secrets. Future adoption requires S9 storage/retention and complete Commerce Scope ownership.

E4 — Nonfinite experiment measurements are unavailable evidence

Owner and surface: ExperimentGuardrailThresholds, ExperimentOperationalMetrics and experiment_guardrail_breaches in personalization_evaluation.py.

Plan: Validate finite thresholds and metrics at their value-object boundary: rates within[0,1], maximum latency strictly positive, observed latency nonnegative. Reject NaN and both infinities. Invalid or missing measurements produce a clear invalid evidence result/error; a rollout evaluator must hold exposure rather than interpret an empty breach tuple as approval. Keep the existing stable breach code ordering for valid measurements. Do not replace NaN with zero or infinite thresholds.

Verification V-E4: test_guardrails_reject_nonfinite_latency_evidence in tests/unit/test_personalization_evaluation.py: NaN/infinities in thresholds and observations; finite values exactly at and immediately beyond thresholds. RED: NaN returns no breach. Assert explicit invalid-evidence failure and preserved finite guardrail outcomes. A prospective rollout consumer test must demonstrate hold/stop on invalid evidence before this helper is used operationally.

Compatibility: Offline validation becomes stricter. Document unavailable-evidence semantics in experimentation guidance; no current live rollout system is claimed or added. Keep evidence aggregates bounded and audience-safe.

E5 — Enforce complete metric-summary state

Owner and surface: personalization_evaluation.py:RankingMetricSummary.

Plan: Treat the summary as a closed state: AVAILABLE requires both metrics finite and in[0,1] with a meaningful positive count; SUPPRESSED and NOT_APPLICABLE require both metrics absent, with the existing count semantics. Reject any one-sided value, invalid status/value combination or nonfinite metric. Use direct dataclass validation; there is no need for a hierarchy of status subclasses unless later consumers need it. The normal summarizer should continue constructing the same safe aggregate shape.

Verification V-E5: test_summary_status_controls_all_metric_values in tests/unit/test_personalization_evaluation.py: a decision table covering both-present, both-absent and each single-present case for every status, plus NaN/infinities, negative counts and boundary0/1 metrics. RED: suppressed recall0.5/ndcgNone accepted. Assert invalid combinations fail construction and safe summarizer outputs serialize without hidden values. Test a literal suppressed JSON projection where applicable.

Compatibility: Stricter value-object contract, no database migration. Preserve public report field names and update evidence documentation. Perform this before E1 changes so new suppression behavior cannot create inconsistent summaries.

M1 — Qualify and fix duplicate two-tower training pairs

Status: Duplicate pairs reproduced; downstream relevance degradation remains a model hypothesis. Surface: two_tower_torch.py:TwoTowerArtifactBuilder._pairs/_train.

Plan: Establish the intended pair unit with the documented unique directed-pair contract. Recommended change: deduplicate directed Item-index pairs across co-view and co-purchase before applying the maximum pair budget. Retain deterministic ordering; do not deduplicate reverse directions unless the contract says so. If cross-signal weighting is intended, express it as an explicit weight rather than repeated conflicting in-batch targets. Multi-positive loss changes require a separate controlled comparison.

Verification V-M1: test_two_tower_pairs_deduplicate_across_signals in tests/unit/test_two_tower_torch.py: identical A-to-B evidence in both mappings, reverse B-to-A, and a distinct A-to-C at the cap boundary. RED: duplicate directed tuple. Verify cap counts unique pairs and candidate ordering is invariant to input maps. Then run seeded baseline/fix experiments on independent temporal holdout with identical splits, pair budgets and hardware; report loss, retrieval Recall/NDCG, training time, memory and unique-pair coverage. No conversion or uplift claim follows from this test.

Compatibility: Exported model behavior may change; bump representation/model identity if training semantics change and reject incompatible old/new learned evidence comparisons. Keep NumPy-only serving and qualify clean subprocess training/imports. The disproved OpenMP probe abort is not a reason to install a runtime workaround.

S8 — Bound compatibility fan-out accumulation

Status: Code-level scalability candidate; operational fan-out still needs qualification. Surface: pipeline/compatibility.py:build_compatibility_candidate_sets.

Plan: Preserve its ordered source stream and two independent score partitions. Deduplicate adjacent candidate IDs for each anchor using the source ordering guarantee, then keep only top-K publication and evaluation candidates in separate bounded selectors. Construct candidate objects only for final survivors. Validate monotonic source order at the source adapter as today; do not use a lifetime seen-ID set as the deduplicator. If callers outside the ordered-stream contract exist, give them an explicit bounded adapter rather than silently changing their behavior.

Verification V-S8: test_compatibility_fanout_uses_bounded_selection in tests/unit/test_compatibility.py: one anchor with1k/10k/more rules, repeated adjacent IDs, tie scores and differing publication/evaluation rankings. Compare to a tiny exhaustive dictionary-and-sort oracle. Assert retained working state O(K), not O(M) fan-out, and unchanged eligible filtering/provenance. Qualify temporary Python peak memory and runtime in a fresh process before assigning deployment severity.

Compatibility: Same outputs and public types, no migration. Selection cost should be O(M log K) with two bounded selectors; output still grows with anchors times K. Retain this as a proposed optimization until the contract and parity evidence support it.

S9 — Bound the offline exposure registry before operational adoption

Status: Adoption prerequisite; no live service caller established. Surface: personalization_evaluation.py:ExperimentExposureGuard.

Plan: Keep the current helper explicitly batch-scoped with a finite exposure count or introduce a narrow exposure-store port for real adoption. Define an attribution window, retention, maximum resident entries and no-capacity behavior. A fully qualified durable adapter needs atomic first-exposure insertion and exact assignment identity from E2/E3, plus deletion policy. Expired/absent evidence must reject attribution; eviction must never manufacture exposure. Scope must include Data Source as well as property/catalog wherever property IDs are not globally unique.

Verification V-S9: test_exposure_registry_respects_capacity_and_retention in tests/unit/test_personalization_evaluation.py: fixed clock, count boundary, expired and active exposure, idempotent repeat and identity variants. RED structural evidence: unlimited dictionary growth. Verify unchanged first exposure, capacity-safe rejection, no unit identifiers in exported evidence and no accidental cross-scope lookup. Test concurrent first insert only when a concurrent/durable adapter is introduced.

Decision owner: Maintainer selects attribution window and storage scope before operational use. This plan does not create a distributed experiment service or infer an authorized production rollout. Document lifecycle explicitly.

H1 — Resolve opt-out overlap semantics before strengthening reads

Status: Resolved by the maintainer's “1. ok” decision: retain the existing post-acknowledgement guarantee. A read begun before an opt-out/deletion acknowledgement may complete using the prior profile. Reads begun after acknowledgement must observe suppression. No post-acknowledgement breach was found.

Plan: State the required consistency property: all reads started after an acknowledged barrier must observe suppression; overlapping reads may follow the agreed linearization point. Recommended hardening is one repository statement/read view joining profile, header and suppression, with the applicable suppression/version taking precedence over older active features. Do not hold a database lock across downstream ranking. Reconcile this with C10 header and S3 dirty projection semantics rather than inventing a separate authorization cache.

Verification V-H1: test_profile_read_after_opt_out_acknowledgement_is_suppressed in tests/integration/test_personalization_storage.py: commit opt-out fully, then start read; expected safe fallback and no usable affinities. GREEN BASELINE unless a real violation exists. A barrier-controlled overlapping test specifies the chosen oracle and tests observed newer suppression overriding an old profile, without labeling every overlap a defect. Repeat for deletion and reauthorization on PostgreSQL.

Disposition: Keep the current rule; semantic strengthening is not part of this change. The barrier-controlled PostgreSQL test records the accepted prior-state overlap and checks the acknowledged state on subsequent reads, including reauthorization. Existing immediate post-acknowledgement suppression remains mandatory.

Q1 — Close performance and capacity evidence gaps

Status: Qualification gap, not an assertion that capacity has already failed. The review's normal suite and microbenchmarks do not establish REQ-023/REQ-024.

The maintainer subsequently requested “2. reduce it” after the proposed twelve-hour/256-GiB run was explained. The final verification for this implementation is therefore reduced to 2,000 Catalog Items, 250,000 views and 50,100 purchases over 731 days, with a ten-minute wall stop and an eight-GiB monitored storage stop threshold. It uses owned PostgreSQL source/control schemas and the existing TCP qualifier with explicit reduced volume minima. This supersedes full-profile execution for the present handoff. The original capacity and availability targets remain unqualified operational targets; this reduced run does not establish a Qualification Claim. No production requirement or named qualification preset is reduced.

The reduced execution finished in 62.24 seconds, including 38.67 seconds of training, and published all eleven strategies. Its declared Also Viewed pair-inclusion observation failed: four source contexts fall below the selected minimum support of five. The failed result is retained in the handoff alongside latency and resource measurements. Neither ranking nor the frozen generator was changed to force that observation to pass. This execution provides bounded performance evidence and denies a correctness/qualification claim for the custom profile. Full-scale capacity and availability remain unqualified, owned by the maintainer and operator as subsequent operational work.

Plan: Use the existing generated-source/simulation and qualification scripts after correctness fixes. First run bounded smoke profiles with signature checks; then scale Catalog size, evidence rows, group width, overlap, eligibility fraction, parameter-grid size, geography and backend independently. Finish with the approved profile: at least 200k Products,100M views,100M purchases and two years of history. Generate data only in an explicitly disposable isolated source; the service must stream it without raw staging. Check complete atomic snapshots, evidence semantics and failure head preservation.

Evidence V-Q1: Retain fresh-process repeated end-to-end training and serving results, resource ceilings, CPU/RSS/temp disk, source load/query plans, cold/warm ANN behavior, concurrent tenant isolation and maintenance backlog progress. Existing DEC-11 targets are training within12h; submit/status p95≤500ms; serving p95≤90ms,p99≤150ms. Monthly99.9% serving/99.5% training availability requires sustained operational evidence, not a local test pass. Record unsupported targets as unqualified, with owners.

Operational boundary: Expensive full-profile runs require declared host/disk/time budgets and disposable database authority. This planning request does not authorize merchant data mutation, deployment or hosted settings. No guessed capacity threshold or extrapolated microbenchmark is an acceptance substitute.

Q2 — Align architecture-check claims with enforced boundaries

Status: Assurance gap. Existing validator correctly checks its documented narrow rules; its success does not prove all architecture properties.

Plan: Keep the current guarantees explicit. From the documented dependency direction, derive reviewed package/module allow rules for stable policy, application orchestration, adapters and composition roots. Extend the existing AST import validator only for those agreed boundaries; add cycle detection over repository modules where cycles are prohibited. Do not replace valid composition with artificial layers or forbid legitimate domain vocabulary sharing. Dynamic imports require explicit permitted seams or separate inspection; import checks cannot certify runtime boundedness or consistency.

Verification V-Q2: Extend tests/delivery/test_repository_harness.py or a small validator-focused test: temporary synthetic modules with a forbidden adapter import, allowed composition-root import, relative import and an actual cycle. RED for newly approved rules; GREEN BASELINE for the existing serving_limits/telemetry boundaries. Require diagnostics to name the violating import/rule, with no sensitive content.

Compatibility: Implement new hard rules only after mapping current imports and resolving documented exceptions. Update architecture and agent-readiness docs and run make architecture-check. A design-pattern census is not a substitute for these rules.

Open decisions, failure containment and handoff

Decision Proposed default Owner / consequence
Late business events, C9 Acknowledge ordering, omit expired signals Maintainer; changes Interaction API semantics
Minimal header lifetime, C10 Keep header independent of feature retention; honor deletion tombstones Maintainer; privacy/lifecycle policy needs explicit treatment
Local limiter capacity, S1/S2 Explicit finite cap; deny allocation when exhausted Operator/maintainer; numeric cap needs workload qualification
Body limit, S6 Finite route-class limit derived from legitimate maximum payload Maintainer; new413/config contract
Offline evaluation unit, E1 One key/strategy observation; one strategy per report until aggregation defined Maintainer; suppression/report semantics
Experiment encoding/version, E3/S9 New canonical version and separate cohorts Maintainer; preserve evidence comparability and attribution window
Opt-out overlaps, H1 Post-ack reads suppressed; explicit overlap linearization Maintainer; do not assert stronger semantics without decision

These are concrete proposed defaults and bounded alternative branches, not requests to stop planning for confirmation. Implementation of the affected behavioral branch needs its decision resolved; unrelated local fixes can proceed independently.

Containment follows existing contracts: preserve the last available snapshot after publication/training failure; degrade optional corrupt/unavailable ANN to snapshot retrieval; refuse malformed admission rather than debit partial limits; omit expired, dirty or suppressed personalization; and reject invalid evaluation evidence rather than approve exposure. None of these plans authorizes deployment or external writes.

Every completed implementation slice must update its governing documentation, record red/green case results and every command's exit/result, name unrun checks and reasons, and inspect its final diff. Use agent handoff across sessions. For migrations, preserve additive mixed-version compatibility where safe and rehearse rollback on disposable state; erasure and changed assignment cohorts are not reversible by restoring old configuration.

Historical planning completion audit

Category IDs Planned deliverable
Confirmed correctness C1–C16 16 dedicated change/test/compatibility plans
Confirmed boundedness/performance S1–S7 7 dedicated plans with structural resource oracles
Confirmed offline evidence E1–E5 5 dedicated plans with controlled negative cases
Provisional model/scalability M1,S8,S9 Qualification first, then a concrete conditional fix
Provisional concurrency H1 Explicit consistency decision and discriminating race tests
Assurance gaps Q1,Q2 Capacity evidence and architecture-rule plans
Rejected probe conclusion OpenMP combined-probe abort No production fix; clean production builder succeeded

All listed issues map to named verification IDs, surfaces and dependency order above. At planning completion the artifact was Draft; the user subsequently requested implementation. Execution evidence is now recorded in the remediation handoff rather than inferred from this table. Documentation/link checks for this planning deliverable are recorded in the review record; previous passing tests remain baseline evidence only.