Design: fixes for repository review findings¶
Control and scope¶
| Field | Value |
|---|---|
| Status | Implementation ready for review; reduced execution measured with failed Oracle observation |
| Revision | R3, 2026-09-30; maintainer accepted H1 and reduced Q1 execution |
| Code baseline | 1dfb34a |
| Author | Codex |
| Decision owner and reviewer | Repository maintainer; implementation requested by user |
| Authority | Initial planning request, followed by $implement implement the plan |
| Evidence | Review record and its reproducible probes |
| Deliverable | Implement and verify the remediation; no deployment |
This design covers all 28 confirmed findings: C1–C16, S1–S7 and E1–E5. It also covers M1, S8, S9, H1, Q1 and Q2 as qualification or investigation plans. Those six concerns are not promoted to confirmed production defects. The OpenMP probe abort was disproved as a production-builder defect and requires no fix.
Current implementation evidence, completed slices, and remaining qualification are tracked in the remediation handoff. The plans below remain the full scope; completing one slice does not complete the implementation goal.
The governing contracts are Commerce Scope and domain language, service requirements, personalization, ANN lifecycle, admission design, and repository verification. Their checked-in revisions at the code baseline are the test basis. Review finding IDs are local risk identifiers, not newly approved product requirements.
Planning is complete when every finding has a change surface, an invariant owner, an implementation sequence, discriminating verification, and applicable compatibility and operations treatment. Implementing these plans and qualifying production are subsequent work. Draft recommendations below do not manufacture approval of new API semantics, retention policy, deployments, or capacity claims.
Shared decisions and verification rules¶
- Keep all reads and writes scoped by Data Source, Tracking ID and Catalog ID. Shopper state adds its existing Shopper key. Serving remains source-independent.
- Preserve exact ranking and evidence definitions. Sampling, score perturbations, approximate replacement of exact metadata ranking, and silent cohort reweighting are consequential contract changes, not shortcuts to a performance fix.
- Put invariants inside existing repositories, selection routines and value objects. Introduce a port or state object only where it hides a real consistency, lifecycle or backend boundary. The reusable limiter never imports recommendations.
- For each verification ID below, promote the synthetic scratch counterexample to the named test, run it before the fix, record the specified red signal, implement that behavior, then rerun the same case and its existing file. Test names below are proposed names, not claims that tests already exist or pass.
- Use fixed aware clocks, explicit seeds and small literal or independent reference oracles. Coordinate races with barriers or controlled iterators, never sleeps. Use actual disposable databases for transaction claims. Assertions about SQL count, body consumption or object residency are appropriate when that bound is the contract.
- Use
make test-focused TEST=path::test_namefor one case, orTEST=pathfor its existing file. The target accepts one quoted path; do not pass multiple filenames. Before Python toolchain work, obtain the module's PyCharm environment as required. - Finish each implementation slice with
make test. Addmake migration-checkfor migrations,make smokefor installed-package changes, andmake compose-checkfor Compose changes. PostgreSQL evidence requires an explicitly disposableTEST_CONTROL_DATABASE_URLandmake test-postgres; ordinary SQLite tests do not prove PostgreSQL locking. Final integrated release candidate usesmake verify. - Benchmarks run in fresh processes with fixed workload, backend/thread budgets and correctness signatures. Retain repetitions, median and tail times, peak RSS, traced allocations where useful, temporary disk, environment and dependency versions. Do not assert arbitrary timing limits in ordinary unit tests. Assert structural bounds there, then evaluate deployment SLOs with controlled measurements.
Dependency order and useful patterns¶
| Slice | Findings | Main responsibility and pattern | Depends on |
|---|---|---|---|
| A | C6 | Fenced publication unit of work | None |
| B | C7, C1 | Consistent source adapter and snapshot-bound repository read | None; C1 benefits from A |
| C | C10, C8, C9, S3, H1 | Monotonic state header, retention policy and resumable maintenance | C10 before maintenance changes |
| D | C2, C3, C5, S5, S8 | Feature-block composition and bounded deterministic selection | C2 before C5; C3 before optimizing candidates |
| E | C11, C13, S7, C12, M1 | Bounded load scheduler, artifact validator and retrieval selection | Correctness before retrieval tuning |
| F | C4, S1, S2 | Typed policy resolution and indexed permit lifecycle | C4 before state changes; S1/S2 together |
| G | S6, C14, C15, C16 | Transport guard and canonical configuration policy | None |
| H | E1–E5, S9 | Evaluation-unit validation and exposure identity/state adapter | E3/E2 together; E5 before report changes |
| I | Q1, Q2 | Capacity evidence and dependency rules | Qualification after relevant fixes |
These are cohesive implementation slices, not a requirement to land one large PR. Split migrations from consumers where additive rollout permits it. Each finding below has its own observable acceptance signal even when it shares a slice with others.
C1 — One snapshot for For You candidates and fallback¶
Owner and surface: storage.py:SnapshotRepository,
serving.py:SnapshotReader/ForYouCandidateLoader; ordinary lane loading remains separate.
REQ-012, REQ-018 and the personalization one-head contract govern consistency.
Plan: Add one bounded repository operation returning a snapshot, the deduplicated global union and its Best Sellers fallback. Resolve head identity and required stored lanes through one SQL statement, or a genuinely repeatable read view if multiple statements are necessary. Build the returned immutable bundle only from that identity. The union stays at the negotiated candidate bound; the fallback stays its own lane. Extend the consumer-owned reader protocol and adapt existing reader fakes. Map absent head, missing required fallback and corrupt payload to the existing classified failures.
Verification V-C1: Add test_for_you_uses_one_snapshot_during_publication in
tests/unit/test_serving.py. A faithful reader advances from A to B between the old
operations; expected literal IDs and disclosed identity must all belong to A or all
to B. RED: A metadata with B fallback. Add a real repository integration case in
test_personalization_storage_candidates.py racing publication and retention, including
a missing fallback and a valid empty union. No merchant-source calls are allowed.
Compatibility and cost: New internal repository operation, no public response or migration. Prefer one round trip over a request-long transaction; merely comparing snapshot IDs and retrying without a bound is rejected. Update technical implementation and personalization ownership docs. Benchmark ordinary versus For You query count and p95 after correctness, using the same database/lanes.
C2 — Similar Items tolerates absent usable feature blocks¶
Owner and surface: pipeline/content.py:_has_metadata/_metadata_matrix/build_similar_items.
The valid insufficient-evidence outcome in REQ-018 must cover feature-poor Catalogs.
Plan: Define feature availability by what the encoder can actually emit. Positive price is a usable price feature; zero remains valid Catalog data without creating a price block. Normalize text presence once, then determine word and character vocabulary availability independently. Keep usable structured features when one text block is empty. Do not catch arbitrary vectorizer failures as insufficient evidence. Return an explicit no-feature result or a correctly shaped empty matrix and bypass multiplication when no positive features exist. Keep candidate keys for every Catalog Item.
Verification V-C2: Parameterize test_similarity_handles_unusable_metadata in
tests/unit/test_content.py with two zero-price Items; whitespace text; punctuation;
one-character text; valid category plus punctuation; and a mixed Catalog. RED: current
IndexError/empty-vocabulary ValueError. Oracle: empty results without features, and
same-category neighbors when structured features remain. Add an end-to-end Training
Run in test_end_to_end.py proving complete publication with valid empty metadata
lanes, and failure preservation for a genuinely unexpected encoder error.
Compatibility: No dependency or schema change. Character-only usable features may produce recommendations where the old code crashed; disclose that in the technical docs. Compare scores for ordinary usable metadata within existing numerical tolerance. Implement this before C5 and before touching ANN representation reuse of these features.
C3 — Filter eligibility before bounded generation¶
Owner and surface: pipeline/run.py:SignalBuilder/_score_candidates, both
workstore.py and polars_workstore.py behavioral queries, provider assembly.
Plan: Pass one immutable Catalog eligibility view to candidate generation. In behavioral queries, filter candidate endpoints before ROW_NUMBER/top-N; retain historical support and pair counts as currently defined rather than discarding their source rows. Apply equivalent filters in Polars and before popularity sorting/selection. Preserve anchor coverage for the complete strategy manifest; an ineligible candidate cannot occupy a retained slot. Check category and geographic variants through the same policy. Keep the final publication filter as an invariant guard, not the primary selection step.
Verification V-C3: test_eligible_candidates_survive_generation_cutoff in
tests/unit/test_workstore.py: A has 200 higher-ranked ineligible neighbors and lower
eligible C with qualifying support. RED: C missing. Compare to a small exhaustive
filter-then-sort reference. Add equivalent Polars, global popularity and category/local
cases with zero, exactly K, and more than K exclusions. test_end_to_end.py must
observe correct primary provenance and avoid fallback when eligible evidence exists.
Compatibility and performance: This corrects recommendation contents and evaluation metrics; bump model identity through the normal release policy so corrected output is not drift-compared as the same implementation. No source-row staging. Inspect query plans and joins at representative Catalog sizes; benchmark high exclusion fractions. Update algorithm and eligibility documentation, not the definition of pair support.
C4 — Distinguish malformed policy resolution from unmatched traffic¶
Owner and surface: serving_limits/core.py:LimitPolicy/AdmissionController, ASGI
translation only if a new reason is exposed. This remains framework-independent.
Plan: Represent resolution as a closed result with matched limits, no selector
match, or missing required partition. Permit allow_unmatched only for no selector
match. A matching malformed rule denies before state acquisition; parent limits must
not be consumed. Reuse the existing unavailable response where compatible, otherwise
add a reviewed stable missing-descriptor reason and map it to deployment-unavailable
503. Do not expose the descriptor's raw value in errors or metric labels.
Verification V-C4: test_missing_partition_fails_closed_with_allow_unmatched in
tests/unit/test_serving_limits.py, for both flag values and with a global parent
rule. RED: allowed under True. A second request with correct dimensions proves the
denied request did not debit its parent. Preserve legitimate unmatched allowance and
matching requests. Add ASGI/API contract assertions for status, body and absent lease.
Compatibility: A formerly allowed malformed request becomes denied, intentionally. Document the flag's scope in the admission design/config example. No database change. Do this before S1/S2 so indexed state never compensates for malformed policy identity.
C5 — Exact lexical ordering at the metadata top-N cutoff¶
Owner and surface: pipeline/content.py bounded selection boundary; DEC-7/REQ-019.
Plan: Give exact selection the total order score descending, Product ID ascending, with self excluded before selecting K. First check whether the pinned native kernel supports an explicit secondary comparator; do not assume input sorting controls heap ties. If it does not, use bounded row/column sparse-product blocks and a per-anchor K-entry selector with that comparator. Bound both block axes using the existing work resource budget. Create candidate objects only for survivors. Keep exact scores; no epsilon perturbation, full N-by-N matrix or approximate ANN substitution.
Verification V-C5: test_similarity_cutoff_keeps_lexical_tie_winners in
tests/unit/test_content.py: 210 identical-category vectors, K=200, anchor 000.
RED: first returned ID 009 instead of 001. Require literal 001–200 where applicable,
shuffle input order and vary native threads/block sizes. A small exhaustive cosine
reference checks mixed ties and near-equal untied scores; self and zero-score edges
must retain established behavior.
Trade-off and qualification: Exact dense-overlap similarity can still require quadratic compute; bounded storage alone does not prove scalable runtime. Measure sparse and highly overlapping Catalogs, report peak buffers and throughput, and reject a default algorithm that cannot meet the agreed training budget. A native tie-comparator extension is preferable if block fallback is too slow, but is a separate dependency decision with parity tests. Update model identity and algorithm docs after qualification.
C6 — Fence every publication mutation by its lease generation¶
Owner and surface: storage.py:TrainingRunRepository/SnapshotRepository and
worker.py claim/publication handoff; migration and lifecycle tests. High assurance.
Plan: Carry an immutable claim token containing run ID, owner and generation (the existing recovery_count can supply the run's generation). Add publication-generation metadata to building snapshots if needed to distinguish abandoned staging. For creation, every set/feature batch, activation and cleanup: lock the run first, check current owner, generation and lease using current authoritative time, then mutate only that generation within the same transaction. Use a consistent run-then-snapshot lock order. Lost owners raise the existing lease-lost error before deletion or insertion. Cleanup may discard its own snapshot ID; it must never search-and-delete the successor's build.
Verification V-C6: test_recovered_publisher_survives_stale_publisher in
tests/integration/test_lifecycle.py promotes the nested-generator reproduction with
foreign keys enabled. RED: successor IntegrityError. Coordinate stale writes at build
creation, set batch, feature batch and activation. Assert successor completes, old
owner gets lease-lost, one complete head exists, and old cleanup cannot delete it.
Mirror with two real PostgreSQL connections in test_postgresql_storage.py; also test
worker identity reuse and failures at each transaction boundary.
Migration and rollout: Add fields/indexes additively, backfill existing available
snapshots without changing heads, and classify pre-upgrade building snapshots as
abandoned only under a valid claim. Drain old worker binaries before enabling new
publication; old writers cannot obey new fencing. Run migration round trips and
make test-postgres. On regression, stop new publication while serving the last head;
do not roll back to destructive unfenced writers. Update operations and implementation
docs with fencing and lock ownership. No long transaction across full snapshot staging.
C7 — Genuine repeatable SQLite source reads¶
Owner and surface: source.py:SqlAlchemySourceAdapter/SourceReadSession,
worker.py:build_source_adapters; REQ-009, REQ-014 and consistent-read boundary.
Plan: Centralize source-engine construction and explicitly configure sqlite3 transaction control. Supported Python versions are 3.12–3.14, so autocommit=False is available; qualify it with isolation-level handling. Alternatively use a single reviewed SQLAlchemy BEGIN event strategy. Never mix both mechanisms. Keep WAL setup outside active transactions and read-session rollback/connection cleanup deterministic. Document the requirements for directly injected engines; reject unsupported SQLite transaction settings rather than advertising snapshot consistency for them.
Verification V-C7: test_sqlite_source_read_is_repeatable_across_queries in
tests/contract/test_source_adapter.py: temporary WAL file, two connections, read A,
commit B on writer, read again within the same source context. RED: B appears despite
unchanged token. Oracle: reader sees A throughout, then a new context sees A+B.
Exercise Catalog-to-interaction reads, transaction cleanup after stream failure and
writer progress. Run the supported-Python matrix and PostgreSQL canonical-read case.
Compatibility: SQLite writers may encounter locking behavior hidden by legacy mode; qualify busy timeout and WAL, rather than retrying a source stream mid-snapshot. No database schema migration and no merchant schema changes. Preserve dialect defaults where their driver actually provides the documented semantics. Update adapter onboarding and technical implementation; configuration must fail before acquiring long-lived resources.
C8 — Enforce retention per Commerce Scope¶
Owner and surface: personalization/capability.py, config.py,
storage.py:ServiceStore composition, personalization/state_store.py, API and worker roots.
Plan: Supply the repository with an immutable scope-to-retention resolver derived from configured capabilities, with the existing maximum 90 days and explicit fallback for previously configured state. Resolve once per operation. Use it for interaction expiry, idempotency expiry, projection boundary and maintenance eligibility. Both API and worker must receive the same policy, including when personalization admission is disabled but maintenance continues. Avoid replacing the entire repository with one global duration when different scopes have different policies.
Verification V-C8: test_configured_scope_retention_reaches_persistence in
tests/integration/test_personalization_storage.py: scopes configured for 30 and 90
days, same synthetic event time. RED: both ledger expiries at day90. At day31, the
30-day scope has no usable signal and the other retains it, with monotonic headers
per C10. Verify the actual deployment composition path, not parsing alone. Include
disabled capability maintenance, idempotency expiry and unknown policy handling.
Rollout: For shortened policies, existing expiry timestamps require bounded reconciliation using occurred_at plus current scope policy. Define its progress and exposure barrier with S3; do not wait for old 90-day timestamps. Increasing retention does not reconstruct deleted data. Version configuration identity where policy changes alter evidence. Confirm metadata-retention treatment in C10; document shortened-policy rollout and one-sided rollback (erased history cannot be recovered).
C9 — Do not project events outside the allowed history window¶
Owner and surface: personalization/health.py, interaction service and repository
projection; maximum-history contract. Depends on the retention resolver and C10 headers.
Proposed behavior: Preserve durable next-sequence acknowledgement for a valid late business interaction but omit its expired feature contribution. This avoids breaking ordering simply because event occurrence is old. Record only the minimal idempotency and acknowledgement state required for replay; do not insert an already expired feature-ledger payload. Lifecycle transitions still apply immediately according to authorization and commit time; an old timestamp must not neutralize opt-out/deletion. The alternative is a typed late-event rejection before any commit. The maintainer owns this public-behavior choice; the design recommends acknowledgement without stale signal.
Verification V-C9: test_expired_interaction_does_not_affect_profile in
tests/integration/test_personalization_storage.py: event exactly at boundary, one
microsecond before and after, under 30/90-day policies. RED: expired affinity present.
For the proposed choice, assert next sequence advances once, exact retry replays the
same acknowledgement, profile signals are unchanged, and causal read observes the
commit version. Include late lifecycle commands and rejection under absent authorization.
Documentation and risk: Update Interaction API late-event behavior, maximum history, idempotency and causal examples together. A mere health-layer check is insufficient because the repository is an independent persistence boundary. Preserve safe error shape and never include Shopper IDs or rejected payloads in metrics/logs.
C10 — Separate monotonic acknowledgement state from disposable history¶
Owner and surface: personalization/state_store.py, state metadata/migration in
storage.py, profile read/causal service and deletion reconciliation. High assurance.
Plan: Make a minimal per-Shopper state header authoritative for last committed sequence, acknowledged version and lifecycle barrier. The existing profile row can retain this header with an empty payload, or an additive watermark table can hold it; recommend the empty-profile header first unless bounded maintenance needs a separate publication version. Stop deriving these fields from the last retained feature event. Expiry clears/rebuilds features but preserves committed sequence; a changed projection gets a version no lower than any acknowledgement, and preferably increments for its own publication. Return insufficient-history fallback for an empty active profile.
Verification V-C10: test_expiry_preserves_acknowledged_watermarks in
tests/integration/test_personalization_storage.py: recent seq1, expired seq2, then
maintenance and seq3. RED: version/watermark2 becomes1 and seq3 rejects. Test full
expiry, retries before/after idempotency expiry, prior valid causal tokens, concurrent
append/rebuild and suppression/deletion. A request begun after the acknowledged commit
must never observe a lower header version, even with no features.
Compatibility and privacy: Keep feature history at the configured maximum. The header contains no interaction payload, but still contains a scoped identity; document its lifecycle and minimal retention basis rather than assuming indefinite retention is free. Reuse the existing deletion tombstone semantics and avoid resurrecting erased profiles. If a new table is selected, backfill maxima from profile, suppression and retained idempotency rows in bounded batches; already-lost watermarks require an explicit recovery procedure, not fabricated sequences. Run migration and PostgreSQL race gates.
C11 — Reap ANN background work independently of the requested head¶
Owner and surface: ann.py:AnnSnapshotRetriever, API lifecycle if executor cleanup is added.
Plan: At entry to the locked scheduler, reap all completed pending entries before testing capacity or checking the requested identity. Completion handling stores a valid result or bounded negative-cache entry with the existing resident limit. Do not hold the lock while doing artifact I/O, and do not wait on an unfinished Future in a serving request. A completed old generation cannot prevent a new one from being scheduled. Keep one loader and bounded pending/resident state; add deterministic close behavior for shutdown if the scheduler acquires lifecycle responsibility.
Verification V-C11: test_ann_completed_old_load_does_not_block_new_head in
tests/integration/test_ann_snapshot.py: controlled reader and Event/Future, finish A,
request B without revisiting A. RED: reader never receives B. Repeat for success,
missing/corrupt artifact, raised failure, another Commerce Scope, and a running load
that legitimately consumes the slot. Assert serving falls back immediately, no duplicate
load exists and residency/pending counts remain bounded.
Compatibility: No artifact schema or response change. Optional ANN disabling still provides its current snapshot fallback. Measure cold-request latency and turnover; avoid unbounded queues as a starvation fix. Document scheduler/residency ownership.
C12 — Separate purchase exclusion history from neighborhood seeds¶
Owner and surface: swimlane_service.py:resolve_swimlane context assembly.
Plan: Keep the complete bounded recent-purchase tuple from the profile for exclusion. Derive a separate first-ten tuple only for retrieval neighborhood seeds. Pass explicit full exclusion sets into configured steps requesting exclusion; do not exclude recent purchases on steps where policy does not request it. Preserve recent-view terminal fill and ordered step provenance. Name these fields by responsibility so another retrieval budget cannot silently become an exclusion budget.
Verification V-C12: test_swimlane_excludes_all_retained_recent_purchases in
tests/unit/test_swimlane_service.py: 11 and 100 purchases, a Best Sellers list containing
the eleventh purchase and a fresh eligible Item. RED: purchase-10 returned. Assert
only the fresh Item survives where exclusion is enabled; disabling exclusion retains
the prior lane behavior. The bounded reader records at most ten purchase seed reads.
Include deduplication across steps and missing seed metadata.
Compatibility: No schema, query or public shape change. Corrected selections may reduce fill when all candidates were purchased; retain truthful insufficient fill rather than expanding history or source-reading. Update Swimlane examples and rerun the geographic/Swimlane integration path.
C13 — Validate native ANN distance semantics¶
Owner and surface: ann.py:LoadedAnnIndex artifact validator, ANN lifecycle contract.
Plan: Validate metric_type against the representation's declared inner-product metric before admitting the loaded index. Require native dimensions, entry count, eligible mapping and representation version to agree as already defined. Bind metric identity into any new artifact envelope version/checksum if it is not explicitly recorded; infer supported legacy metric only from the validated native index, never from a permissive default. An incompatible artifact follows the existing safe fallback.
Verification V-C13: test_ann_rejects_incompatible_native_metric in
tests/integration/test_ann_snapshot.py: create a checksummed L2 HNSW artifact with
otherwise valid shape/IDs. RED: load succeeds. Require typed validation failure and
snapshot fallback in serving. Valid inner-product artifacts must preserve native/exact
similarity semantics, eligible filtering and same-scope binding; exercise learned and
metadata representation paths.
Compatibility and rollout: Existing legitimately incompatible artifacts will stop loading and need rebuilding; the published snapshot fallback remains available. If envelope metadata changes, use an additive migration and explicit representation version, then deploy readers before writers. No best-effort reinterpretation of L2 as cosine. Update the lifecycle spec's test mapping and artifact troubleshooting guidance.
C14 — Validate run IDs before telemetry construction¶
Owner and surface: api.py:inspect_training_run, transport contract tests.
Plan: Validate bounded path input before constructing TelemetryContext or querying persistence. Use a bounded validator appropriate to existing run identities; avoid requiring UUID syntax if current repository callers legitimately use other IDs. Return the reviewed invalid-identifier 422 response; retain 404 for a valid unknown identity. Reject credential/URL/control-shaped values without echoing them. Keep only bounded route/outcome observations for rejected input.
Verification V-C14: test_inspect_run_rejects_invalid_identifier_before_telemetry
in tests/contract/test_api.py: 256 characters, control characters through proper URL
encoding, and credential-shaped text. RED: HTTP500. Assert 422, safe error envelope,
no rejected content in captured telemetry and no repository access. Valid unknown
ID remains404; existing submitted run remains200. Test enabled and disabled observability.
Compatibility: This intentionally changes malformed requests from500 to4xx. Update OpenAPI through its generator, not generated output edits, and document the accepted bound. No durable-state migration or blanket narrowing of unrelated identifiers.
C15 — Make configuration identity match runtime parameter selection¶
Owner and surface: config.py:DeploymentConfig.config_version,
pipeline/run.py popularity/trending parameter selectors; comparable-run evidence.
Plan: Treat search grids as sets, matching their current sorted fingerprint. Canonicalize candidates once and define a total tie order: metric objective first, distance from existing preferred defaults second, then the canonical parameter tuple. No-holdout fallback selects the default if available, otherwise the same deterministic closest candidate and final tuple order. Apply it to all relevant popularity/trending grids and check behavioral-search ties, without changing the primary metric objective.
Verification V-C15: test_parameter_selection_is_invariant_to_grid_order in
tests/unit/test_config.py or a focused selector test module: [2,8] versus [8,2], no
holdout, no default and tied holdout objectives. RED: same hash, half-life2 versus8.
Enumerate permutations for small multidimensional trending grids and compare selected
parameters/output, not merely fingerprints. Verify genuinely different grid sets still
have different versions and ordinary default selection remains characterized.
Compatibility: Canonicalization is preferred to preserving accidental caller order in the hash. Some previously ambiguous grids change output; bump implementation/model identity and document tie policy so historic runs are not falsely comparable. No hash algorithm migration is required if the set semantics and payload are unchanged.
C16 — Reject nonfinite polling intervals during configuration¶
Owner and surface: config.py:DeploymentConfig.__post_init__/from_mapping.
Plan: Require a finite positive numeric interval and reject booleans as a numeric interval if the existing parsing helpers permit them. Apply validation to direct construction as well as serialized config. Reuse the repository's finite-positive helper where it actually fits; no general validation framework is needed. Failure must happen before opening repositories or starting the worker loop, with a safe configuration error naming the field.
Verification V-C16: test_worker_poll_interval_must_be_finite in
tests/unit/test_config.py: NaN, positive/negative infinity, zero, negative and a valid
fractional value; include the accepted JSON parser's NaN form. RED: NaN constructs
a config and reaches a failing sleep. Assert all invalid inputs fail construction,
the valid interval survives, and no runtime sleep is needed in the test.
Compatibility: Previously malformed config fails early instead of crashing idle workers. No schema or public endpoint change. Update configuration reference and doctor guidance; rerun existing configuration and worker lifecycle tests.
S1 — Bound rate-bucket residency without restoring spent tokens¶
Owner and surface: serving_limits/core.py:InMemoryLimitState, optional serialized
state-bound configuration and denial translation. Token semantics remain unchanged.
Plan: Store a bucket's next semantically safe removal time: the instant at which it would be fully replenished. Maintain an indexed expiry structure with one live entry per bucket; refresh it in place when tokens are consumed. Reclaim only fully refilled buckets. Add an explicit maximum resident bucket count, validated positive and independently qualified for deployment. If allocation would exceed the bound, deny the new partition without mutating any parent's rate/concurrency state. Expose capacity exhaustion as unavailable, with bounded reason metrics, rather than silently evicting a depleted bucket or allowing unlimited partitions.
Verification V-S1: test_idle_rate_buckets_are_reclaimed_without_resetting_debt
in tests/unit/test_serving_limits.py: spend bucket, request just before/full refill,
advance a fixed clock and touch another partition. RED: all10,000 idle buckets remain.
Assert no early capacity restoration, bounded residency at saturation, correct refill
after removal/recreation, and atomic parent/child denial. The expiry index itself must
remain bounded after repeated refreshes; a lazy heap with unlimited stale nodes fails.
Trade-off and operations: A hard cap can deny an otherwise affordable new consumer, so the default/capacity policy is a maintainer-owned deployment choice. Recommend explicit local-state configuration and503 on state exhaustion. Benchmark high churn, few hot keys and version rollovers; retain the current single-process limitation. Document why ordinary LRU eviction is unsafe for depleted token buckets.
S2 — Index concurrency expiry and retry capacity¶
Owner and surface: the same state adapter as S1; no distributed enforcement claim.
Plan: Replace full _leases reconstruction with an expiry index and direct
lease-to-permit ownership. Index permits by their bucket for retry calculation and
release. Remove each expired permit once, decrementing its own bucket idempotently.
Use a bounded indexed structure; explicit release removes entries rather than leaving
unlimited tombstones. Weighted retry selection must use bucket-local expiry/cumulative
cost, not sorting every lease in the process. Preserve partial expiry when one lease
contains rules with different TTLs and atomic all-rule acquisition under the lock.
Verification V-S2: test_admission_work_does_not_scan_unrelated_live_leases in
tests/unit/test_serving_limits.py: increasing synthetic lease populations and an
unrelated bucket request. RED: operations scale with all live permits. Through public
acquire/release, compare decisions, retry delays and permit conservation to a simple
small reference using a fixed clock, varied costs/TTLs and repeated releases. Include
same-deadline expiry bursts, denial storms and generated duplicate lease IDs.
Performance acceptance: Each nonexpiry operation should avoid global Theta(P) work; ordinary admit/release should be O(R log P), for R matched rules and P permits, plus expired entries actually processed. Total creation/expiry work for N retained admissions should be O(N log N), replacing the measured quadratic scan. Expiry bursts still cost work proportional to expired permits; explicitly cap state and qualify tail latency rather than claim constant worst-case time. Repeat the1k/2k/4k probe and larger bounded profiles, recording signatures, residency and lock duration.
S3 — Resumable bounded personalization maintenance¶
Owner and surface: personalization/state_store.py:expire/_rebuild_profile,
worker maintenance orchestration; depends on C10 and scope retention C8.
Plan: Separate logical signal expiry from physical deletion. Use keyset pagination
for affected keys and deterministic row limits for delete batches; eliminate global
.all() and unbounded all-key transactions. Rebuild one projection through a streamed
sequence reducer with bounded feature state. For a single very active Shopper, support
checkpointed rebuild chunks or a time/work budget; merely limiting keys is insufficient.
Publish rebuilt features only if the source/header version still matches, otherwise
retry with a bound. Maintain a durable dirty/rebuild marker if deletion precedes final
projection publication, so interruption cannot strand stale features as usable.
Verification V-S3: test_expiry_batches_are_resumable_and_preserve_causal_state
in tests/integration/test_personalization_storage.py: multiple expired keys, one
large key, retained/new interactions and forced interruption after a delete batch.
RED: all keys/rows fetched in one transaction or stale projection lacks a resumable
marker. Assert batch row/key/work limits, eventual exact reference projection, monotonic
headers, isolation, and no lost concurrent append. PostgreSQL tests must measure lock
duration and race CAS publication; SQLite alone is inadequate here.
Migration and operations: Add only necessary progress/dirty fields or a bounded work table, with indexes on scope/expiry/key/sequence. Writers maintain invalidation markers transactionally. Serving omits logically expired or dirty signals and returns documented safe fallback, preserving causal acknowledgement header visibility. Report aggregate backlog, batches and duration without Shopper labels. Rollback must retain dirty-state protection; pausing physical cleanup must not make expired signals usable. Qualify bounded maintenance progress under shortened retention and continuous ingestion.
S4 — Stream exact holdout evaluation instead of retaining the pair graph¶
Owner and surface: pipeline/workstore.py, polars_workstore.py,
pipeline/evaluation.py, run.py parameter selection and WorkEvidence handoff.
Plan: Introduce a streamed holdout evaluator over derived, deduplicated directional edges ordered by anchor. For each anchor, keep predicted top-K membership/rank (bounded), count distinct relevant neighbors, hits and DCG contribution while streaming truth. Compute ideal DCG from min(K,relevant_count); accumulate cohort means and coverage using their existing definitions. Do not build relevant-neighbor sets for every anchor or materialize them again in retained WorkEvidence. Keep the current pure evaluator as the small-data reference. Both stores implement the same narrow evaluation contract.
Verification V-S4: test_streamed_holdout_matches_exact_reference_with_bounded_state
in tests/unit/test_workstore.py and Polars equivalent: tiny literal graphs cover
duplicates, empty predictions, sparse/cold cohorts, coverage and temporal partitions.
RED structural evidence: full pair fetch/materialization. A growing dense graph must
not grow Python truth state with total edges; compare exact aggregate metrics at
different fetch sizes and parameter-grid reuse. Add end-to-end evidence parity and
stream failure/cleanup tests.
Trade-off: Exact streaming is selected over sampling to preserve REQ-013 and evidence schema meaning. It may scan truth more often during parameter search; reuse derived queries or bounded batched candidate evaluation rather than cache the entire graph. Record CPU, peak RSS and spill at high edge density, maintaining no raw-row staging. Update evaluation flow docs; schema version changes only if metric semantics change.
S5 — Share bounded complementary-category candidates¶
Owner and surface: pipeline/run.py:_category_pair_candidates and provider assembly;
reuse pipeline/fallbacks.py:KeyedCandidateProvider.
Plan: Sort each destination category's eligible Items once by the established purchase-score/Product-ID order, retaining the existing generation bound. For each source category, compute its paired-category sequence once with the current support ordering and truncate after the required candidates. Share immutable category-keyed candidate tuples across anchors through the existing keyed provider; avoid constructing identical RecommendationCandidate objects and lists for every Item. Preserve the current ranking/provenance order unless C3 separately requires eligibility correction.
Verification V-S5: test_category_pair_candidates_are_shared_without_changing_rankings
in tests/unit/test_category_strategies.py: two categories with several anchors,
support ties, purchase-score ties and multiple paired categories. RED structural
evidence: the destination category is sorted once per anchor. Compare complete outputs
to a small exhaustive reference, across input permutations, evaluation/published
partitions and missing categories. Do not use object-identity assertions as the only
oracle; bound sort/construction work as additional evidence.
Performance: Destination sorting should occur once per evidence partition/category; category-equivalent anchors should cost an O(1) provider lookup plus unavoidable final output work. Measure 1k/2k/4k balanced-category workloads, candidate allocations and total snapshot output, with exact signatures. No public schema or migration change. Update implemented data-flow docs and rerun geographic/category end-to-end tests.
S6 — Enforce request bytes before JSON parsing¶
Owner and surface: a small ASGI transport guard assembled in api.py, deployment
configuration, contract tests and API error documentation.
Plan: Add a positive finite byte limit per supported POST body class, with one reviewed default sufficient for the maximum legitimate interaction and Swimlane payload. Count actual bytes in wrapped receive; reject a declared oversized body before reads and a streamed body immediately when cumulative bytes exceed the limit. Apply it regardless of Content-Length, chunked transfer, whitespace padding or invalid JSON. On limit violation, return one stable413 before endpoint/state execution; ensure middleware unwinding does not also send500. Preserve protected-route authentication before body consumption and keep safe request observations.
Verification V-S6: test_request_body_limit_stops_chunked_input_before_parse in
tests/contract/test_api.py: exact limit, limit+1, oversized Content-Length, no length,
incorrect length, multiple chunks, unauthorized protected request and legitimate maximum
payload. RED:1/8MiB padded JSON accepted. An actual ASGI receive counter proves input
is not drained without bound after rejection; persistence remains unchanged. Include
disconnect and cancellation before/after the boundary, with exactly one response.
Compatibility and operations: Proposed limit is an explicit configuration contract; the maintainer sets the numeric default from legitimate payload bounds, not benchmark guesswork. Deploy/document413 and proxy settings together, while the app remains its own enforcement boundary. Qualify memory under concurrent bounded requests; input caps do not replace admission concurrency or downstream timeout controls. Never retain body bytes or auth headers in diagnostic evidence. Add OpenAPI error responses and config tests.
S7 — Bounded exact ANN selection under dense ties¶
Owner and surface: ann.py:LoadedAnnIndex._exact; score semantics remain unchanged.
Plan: Reuse its fixed vector chunks and maintain only the best K eligible,
nonexcluded (score, Product ID) positions, with a deterministic bounded selector.
Filter exclusions before retention and instantiate AnnMatch only after final selection.
A per-chunk partition followed by explicit lexical threshold selection is also viable
if its temporary storage is bounded by chunk size; ensure it includes lexical winners
among all ties in that chunk. Avoid a Catalog-sized Python match list and sorting it.
The score array can be removed once streaming selection proves parity; do not conflate
that optimization with a change to native ANN behavior.
Verification V-S7: test_ann_exact_ties_bound_matches_and_keep_lexical_order in
tests/integration/test_ann_snapshot.py:10k equal vectors, K100, exclusions including
early lexical winners, learned negative scores and metadata zero scores. RED:10k match
objects for100 results. Compare selected IDs/scores to a small exhaustive reference;
vary chunk sizes and exercise the native-underfill fallback through public search.
Acceptance and risk: Additional Python selection state is O(K) plus bounded chunk buffers; matching object count is at most K. Repeat10k/100k allocation and latency measurements, including untied and exclusion-heavy workloads. Native vector/index residency is separately budgeted; no promise of O(K) total loaded-index memory. Update ANN performance docs and preserve current representation-specific ranking.
E1 — Count independent evaluation units and reject duplicates¶
Owner and surface: personalization_evaluation.py:PersonalizationEvaluationContract.
Plan: Define one independent observation per ShopperEvaluationKey and strategy for a given evaluation invocation/split. Reject duplicates before metric accumulation, including same identity with altered cohort labels. Report suppression counts from distinct permitted units, not replayed rows. For a report that combines strategies, keep per-strategy denominators and count each Shopper only once for privacy suppression; do not silently collapse differing strategy outcomes into an unspecified mean. The maintainer owns whether multi-strategy top-level reports are allowed; default to one strategy per report until its aggregate semantics are explicit.
Verification V-E1: test_duplicate_evaluation_units_cannot_bypass_suppression in
tests/unit/test_personalization_evaluation.py: one observation repeated10 times at
minimum10. RED: available count10. Require a clear duplicate-input failure, or the
documented suppressed distinct-unit result if idempotent deduplication is chosen.
Test conflicting duplicate labels,9/10 distinct Shoppers, same Shopper in different
Catalogs, feedback exclusions and cross-property rejection. Use literal Recall/NDCG
oracles; no raw identities in the report.
Compatibility: Offline helper has no live callers, so strict duplicate rejection is recommended over best-effort bias correction. O(U) validation state for U declared units is acceptable only under explicit evaluation input bounds; large deployments use streamed external grouping, not an unbounded service registry. Document the unit and version evidence meaning when observation_count semantics change.
E2 — Attribute only to the exact exposed assignment¶
Owner and surface: personalization_evaluation.py:ExperimentExposureGuard;
implement alongside E3's identity encoding.
Plan: Include variant in the canonical exposure identity and compare the complete stored assignment to the requested assignment before outcome attribution. Experiment, config version, unit type, property and unit key must also match. Preserve first-exposure immutability and timezone-aware exposure-before-outcome ordering. Record two legitimately distinct assignments separately only if the experiment contract permits it; do not use one variant's exposure as authorization for another.
Verification V-E2: test_outcome_requires_exact_variant_exposure in
tests/unit/test_personalization_evaluation.py: record holdout exposure, request
personalized attribution. RED: accepted. Require absent-exposure or assignment-mismatch
failure, then expose the requested variant and prove valid attribution. Test version,
unit, property, outcome timing and repeated identical exposure. Stored first exposure
must not be replaced by a conflicting assignment.
Compatibility: Previously invalid offline attribution now fails. No production database migration because the current registry is local memory. Future durable state must enforce uniqueness on complete versioned identity, not just a mutable dictionary lookup. Document the change with E3; old ambiguous records cannot establish a new exposure.
E3 — Canonical versioned HMAC material¶
Owner and surface: StableExperimentAssigner and exposure correlation helper in
personalization_evaluation.py. Preserve SHA3-512/HMAC-SHA3-512 policy.
Plan: Encode the identity as a canonical ordered JSON array or explicit length-prefixed UTF-8 fields, with a domain/version prefix distinguishing assignment from exposure. Choose JSON array encoding to reuse existing deterministic JSON conventions. Include every identity component; exposure also includes variant per E2. Hash exactly those bytes and never concatenate unescaped delimiter-separated values. Do not add fallback verification of the old ambiguous format.
Verification V-E3: test_experiment_identity_encoding_is_unambiguous in
tests/unit/test_personalization_evaluation.py: property a:b / experiment c versus
property a / experiment b:c. RED: the second identity inherits the first exposure.
Test delimiter/Unicode values in each field, distinct variants/units/versions and
deterministic repeat assignment. Validate encoding with fixed canonical byte literals
and independently calculated published synthetic HMAC vectors; never include real keys.
Compatibility and rollout: Assignment buckets can change under a new encoding. Introduce an explicit assignment-algorithm/config version, start a new experiment cohort and do not pool exposures/outcomes across versions. Offline evidence replay must regenerate assignments from authorized synthetic inputs or label prior evidence incompatible. This change is not a rotation of production secrets. Future adoption requires S9 storage/retention and complete Commerce Scope ownership.
E4 — Nonfinite experiment measurements are unavailable evidence¶
Owner and surface: ExperimentGuardrailThresholds, ExperimentOperationalMetrics
and experiment_guardrail_breaches in personalization_evaluation.py.
Plan: Validate finite thresholds and metrics at their value-object boundary: rates within[0,1], maximum latency strictly positive, observed latency nonnegative. Reject NaN and both infinities. Invalid or missing measurements produce a clear invalid evidence result/error; a rollout evaluator must hold exposure rather than interpret an empty breach tuple as approval. Keep the existing stable breach code ordering for valid measurements. Do not replace NaN with zero or infinite thresholds.
Verification V-E4: test_guardrails_reject_nonfinite_latency_evidence in
tests/unit/test_personalization_evaluation.py: NaN/infinities in thresholds and
observations; finite values exactly at and immediately beyond thresholds. RED: NaN
returns no breach. Assert explicit invalid-evidence failure and preserved finite
guardrail outcomes. A prospective rollout consumer test must demonstrate hold/stop
on invalid evidence before this helper is used operationally.
Compatibility: Offline validation becomes stricter. Document unavailable-evidence semantics in experimentation guidance; no current live rollout system is claimed or added. Keep evidence aggregates bounded and audience-safe.
E5 — Enforce complete metric-summary state¶
Owner and surface: personalization_evaluation.py:RankingMetricSummary.
Plan: Treat the summary as a closed state: AVAILABLE requires both metrics finite and in[0,1] with a meaningful positive count; SUPPRESSED and NOT_APPLICABLE require both metrics absent, with the existing count semantics. Reject any one-sided value, invalid status/value combination or nonfinite metric. Use direct dataclass validation; there is no need for a hierarchy of status subclasses unless later consumers need it. The normal summarizer should continue constructing the same safe aggregate shape.
Verification V-E5: test_summary_status_controls_all_metric_values in
tests/unit/test_personalization_evaluation.py: a decision table covering both-present,
both-absent and each single-present case for every status, plus NaN/infinities,
negative counts and boundary0/1 metrics. RED: suppressed recall0.5/ndcgNone accepted.
Assert invalid combinations fail construction and safe summarizer outputs serialize
without hidden values. Test a literal suppressed JSON projection where applicable.
Compatibility: Stricter value-object contract, no database migration. Preserve public report field names and update evidence documentation. Perform this before E1 changes so new suppression behavior cannot create inconsistent summaries.
M1 — Qualify and fix duplicate two-tower training pairs¶
Status: Duplicate pairs reproduced; downstream relevance degradation remains a
model hypothesis. Surface: two_tower_torch.py:TwoTowerArtifactBuilder._pairs/_train.
Plan: Establish the intended pair unit with the documented unique directed-pair contract. Recommended change: deduplicate directed Item-index pairs across co-view and co-purchase before applying the maximum pair budget. Retain deterministic ordering; do not deduplicate reverse directions unless the contract says so. If cross-signal weighting is intended, express it as an explicit weight rather than repeated conflicting in-batch targets. Multi-positive loss changes require a separate controlled comparison.
Verification V-M1: test_two_tower_pairs_deduplicate_across_signals in
tests/unit/test_two_tower_torch.py: identical A-to-B evidence in both mappings, reverse
B-to-A, and a distinct A-to-C at the cap boundary. RED: duplicate directed tuple.
Verify cap counts unique pairs and candidate ordering is invariant to input maps.
Then run seeded baseline/fix experiments on independent temporal holdout with identical
splits, pair budgets and hardware; report loss, retrieval Recall/NDCG, training time,
memory and unique-pair coverage. No conversion or uplift claim follows from this test.
Compatibility: Exported model behavior may change; bump representation/model identity if training semantics change and reject incompatible old/new learned evidence comparisons. Keep NumPy-only serving and qualify clean subprocess training/imports. The disproved OpenMP probe abort is not a reason to install a runtime workaround.
S8 — Bound compatibility fan-out accumulation¶
Status: Code-level scalability candidate; operational fan-out still needs
qualification. Surface: pipeline/compatibility.py:build_compatibility_candidate_sets.
Plan: Preserve its ordered source stream and two independent score partitions. Deduplicate adjacent candidate IDs for each anchor using the source ordering guarantee, then keep only top-K publication and evaluation candidates in separate bounded selectors. Construct candidate objects only for final survivors. Validate monotonic source order at the source adapter as today; do not use a lifetime seen-ID set as the deduplicator. If callers outside the ordered-stream contract exist, give them an explicit bounded adapter rather than silently changing their behavior.
Verification V-S8: test_compatibility_fanout_uses_bounded_selection in
tests/unit/test_compatibility.py: one anchor with1k/10k/more rules, repeated adjacent
IDs, tie scores and differing publication/evaluation rankings. Compare to a tiny
exhaustive dictionary-and-sort oracle. Assert retained working state O(K), not O(M)
fan-out, and unchanged eligible filtering/provenance. Qualify temporary Python peak
memory and runtime in a fresh process before assigning deployment severity.
Compatibility: Same outputs and public types, no migration. Selection cost should be O(M log K) with two bounded selectors; output still grows with anchors times K. Retain this as a proposed optimization until the contract and parity evidence support it.
S9 — Bound the offline exposure registry before operational adoption¶
Status: Adoption prerequisite; no live service caller established.
Surface: personalization_evaluation.py:ExperimentExposureGuard.
Plan: Keep the current helper explicitly batch-scoped with a finite exposure count or introduce a narrow exposure-store port for real adoption. Define an attribution window, retention, maximum resident entries and no-capacity behavior. A fully qualified durable adapter needs atomic first-exposure insertion and exact assignment identity from E2/E3, plus deletion policy. Expired/absent evidence must reject attribution; eviction must never manufacture exposure. Scope must include Data Source as well as property/catalog wherever property IDs are not globally unique.
Verification V-S9: test_exposure_registry_respects_capacity_and_retention in
tests/unit/test_personalization_evaluation.py: fixed clock, count boundary, expired
and active exposure, idempotent repeat and identity variants. RED structural evidence:
unlimited dictionary growth. Verify unchanged first exposure, capacity-safe rejection,
no unit identifiers in exported evidence and no accidental cross-scope lookup. Test
concurrent first insert only when a concurrent/durable adapter is introduced.
Decision owner: Maintainer selects attribution window and storage scope before operational use. This plan does not create a distributed experiment service or infer an authorized production rollout. Document lifecycle explicitly.
H1 — Resolve opt-out overlap semantics before strengthening reads¶
Status: Resolved by the maintainer's “1. ok” decision: retain the existing post-acknowledgement guarantee. A read begun before an opt-out/deletion acknowledgement may complete using the prior profile. Reads begun after acknowledgement must observe suppression. No post-acknowledgement breach was found.
Plan: State the required consistency property: all reads started after an acknowledged barrier must observe suppression; overlapping reads may follow the agreed linearization point. Recommended hardening is one repository statement/read view joining profile, header and suppression, with the applicable suppression/version taking precedence over older active features. Do not hold a database lock across downstream ranking. Reconcile this with C10 header and S3 dirty projection semantics rather than inventing a separate authorization cache.
Verification V-H1: test_profile_read_after_opt_out_acknowledgement_is_suppressed
in tests/integration/test_personalization_storage.py: commit opt-out fully, then start
read; expected safe fallback and no usable affinities. GREEN BASELINE unless a real
violation exists. A barrier-controlled overlapping test specifies the chosen oracle
and tests observed newer suppression overriding an old profile, without labeling every
overlap a defect. Repeat for deletion and reauthorization on PostgreSQL.
Disposition: Keep the current rule; semantic strengthening is not part of this change. The barrier-controlled PostgreSQL test records the accepted prior-state overlap and checks the acknowledged state on subsequent reads, including reauthorization. Existing immediate post-acknowledgement suppression remains mandatory.
Q1 — Close performance and capacity evidence gaps¶
Status: Qualification gap, not an assertion that capacity has already failed. The review's normal suite and microbenchmarks do not establish REQ-023/REQ-024.
The maintainer subsequently requested “2. reduce it” after the proposed twelve-hour/256-GiB run was explained. The final verification for this implementation is therefore reduced to 2,000 Catalog Items, 250,000 views and 50,100 purchases over 731 days, with a ten-minute wall stop and an eight-GiB monitored storage stop threshold. It uses owned PostgreSQL source/control schemas and the existing TCP qualifier with explicit reduced volume minima. This supersedes full-profile execution for the present handoff. The original capacity and availability targets remain unqualified operational targets; this reduced run does not establish a Qualification Claim. No production requirement or named qualification preset is reduced.
The reduced execution finished in 62.24 seconds, including 38.67 seconds of training, and published all eleven strategies. Its declared Also Viewed pair-inclusion observation failed: four source contexts fall below the selected minimum support of five. The failed result is retained in the handoff alongside latency and resource measurements. Neither ranking nor the frozen generator was changed to force that observation to pass. This execution provides bounded performance evidence and denies a correctness/qualification claim for the custom profile. Full-scale capacity and availability remain unqualified, owned by the maintainer and operator as subsequent operational work.
Plan: Use the existing generated-source/simulation and qualification scripts after correctness fixes. First run bounded smoke profiles with signature checks; then scale Catalog size, evidence rows, group width, overlap, eligibility fraction, parameter-grid size, geography and backend independently. Finish with the approved profile: at least 200k Products,100M views,100M purchases and two years of history. Generate data only in an explicitly disposable isolated source; the service must stream it without raw staging. Check complete atomic snapshots, evidence semantics and failure head preservation.
Evidence V-Q1: Retain fresh-process repeated end-to-end training and serving results, resource ceilings, CPU/RSS/temp disk, source load/query plans, cold/warm ANN behavior, concurrent tenant isolation and maintenance backlog progress. Existing DEC-11 targets are training within12h; submit/status p95≤500ms; serving p95≤90ms,p99≤150ms. Monthly99.9% serving/99.5% training availability requires sustained operational evidence, not a local test pass. Record unsupported targets as unqualified, with owners.
Operational boundary: Expensive full-profile runs require declared host/disk/time budgets and disposable database authority. This planning request does not authorize merchant data mutation, deployment or hosted settings. No guessed capacity threshold or extrapolated microbenchmark is an acceptance substitute.
Q2 — Align architecture-check claims with enforced boundaries¶
Status: Assurance gap. Existing validator correctly checks its documented narrow rules; its success does not prove all architecture properties.
Plan: Keep the current guarantees explicit. From the documented dependency direction, derive reviewed package/module allow rules for stable policy, application orchestration, adapters and composition roots. Extend the existing AST import validator only for those agreed boundaries; add cycle detection over repository modules where cycles are prohibited. Do not replace valid composition with artificial layers or forbid legitimate domain vocabulary sharing. Dynamic imports require explicit permitted seams or separate inspection; import checks cannot certify runtime boundedness or consistency.
Verification V-Q2: Extend tests/delivery/test_repository_harness.py or a small
validator-focused test: temporary synthetic modules with a forbidden adapter import,
allowed composition-root import, relative import and an actual cycle. RED for newly
approved rules; GREEN BASELINE for the existing serving_limits/telemetry boundaries.
Require diagnostics to name the violating import/rule, with no sensitive content.
Compatibility: Implement new hard rules only after mapping current imports and
resolving documented exceptions. Update architecture and agent-readiness docs and run
make architecture-check. A design-pattern census is not a substitute for these rules.
Open decisions, failure containment and handoff¶
| Decision | Proposed default | Owner / consequence |
|---|---|---|
| Late business events, C9 | Acknowledge ordering, omit expired signals | Maintainer; changes Interaction API semantics |
| Minimal header lifetime, C10 | Keep header independent of feature retention; honor deletion tombstones | Maintainer; privacy/lifecycle policy needs explicit treatment |
| Local limiter capacity, S1/S2 | Explicit finite cap; deny allocation when exhausted | Operator/maintainer; numeric cap needs workload qualification |
| Body limit, S6 | Finite route-class limit derived from legitimate maximum payload | Maintainer; new413/config contract |
| Offline evaluation unit, E1 | One key/strategy observation; one strategy per report until aggregation defined | Maintainer; suppression/report semantics |
| Experiment encoding/version, E3/S9 | New canonical version and separate cohorts | Maintainer; preserve evidence comparability and attribution window |
| Opt-out overlaps, H1 | Post-ack reads suppressed; explicit overlap linearization | Maintainer; do not assert stronger semantics without decision |
These are concrete proposed defaults and bounded alternative branches, not requests to stop planning for confirmation. Implementation of the affected behavioral branch needs its decision resolved; unrelated local fixes can proceed independently.
Containment follows existing contracts: preserve the last available snapshot after publication/training failure; degrade optional corrupt/unavailable ANN to snapshot retrieval; refuse malformed admission rather than debit partial limits; omit expired, dirty or suppressed personalization; and reject invalid evaluation evidence rather than approve exposure. None of these plans authorizes deployment or external writes.
Every completed implementation slice must update its governing documentation, record red/green case results and every command's exit/result, name unrun checks and reasons, and inspect its final diff. Use agent handoff across sessions. For migrations, preserve additive mixed-version compatibility where safe and rehearse rollback on disposable state; erasure and changed assignment cohorts are not reversible by restoring old configuration.
Historical planning completion audit¶
| Category | IDs | Planned deliverable |
|---|---|---|
| Confirmed correctness | C1–C16 | 16 dedicated change/test/compatibility plans |
| Confirmed boundedness/performance | S1–S7 | 7 dedicated plans with structural resource oracles |
| Confirmed offline evidence | E1–E5 | 5 dedicated plans with controlled negative cases |
| Provisional model/scalability | M1,S8,S9 | Qualification first, then a concrete conditional fix |
| Provisional concurrency | H1 | Explicit consistency decision and discriminating race tests |
| Assurance gaps | Q1,Q2 | Capacity evidence and architecture-rule plans |
| Rejected probe conclusion | OpenMP combined-probe abort | No production fix; clean production builder succeeded |
All listed issues map to named verification IDs, surfaces and dependency order above. At planning completion the artifact was Draft; the user subsequently requested implementation. Execution evidence is now recorded in the remediation handoff rather than inferred from this table. Documentation/link checks for this planning deliverable are recorded in the review record; previous passing tests remain baseline evidence only.