Skip to content

Repository remediation handoff

Objective and state

  • Objective: $implement implement the plan, covering the entire remediation design.
  • Status: ready for review under the maintainer's revised execution scope. Remediation regression gates pass; the reduced experiment finished with a failed planted-result observation. Full capacity and sustained availability remain unqualified.
  • Baseline: 1dfb34a; no branch, commit, PR, deployment, or hosted settings changed.
  • Existing new planning/evidence files preserved. All source changes in this execution belong to the remediation; no unrelated changes were present at implementation start.

Current implementation evidence

The table preserves each slice's initial evidence. Later release checkpoints below supersede its pending gate entries; the final timestamp-factory candidate passes the integrated, PostgreSQL, static and supported-runtime gates. The final maintainer-decision checkpoint below resolves H1 and records the reduced Q1 result; it supersedes earlier pending decisions.

Finding Change and observable evidence Remaining qualification
C6 Claim generation captured under the claim transaction and carried by worker; live lease/generation checked under run lock before creation, every set/feature batch and activation; cleanup restricted to its own building ID. SQLite recovery/expiry cases and PostgreSQL separate-connection recovery races pass, including worker-ID reuse. Full PostgreSQL gate and final integrated release checks; no schema migration needed because recovery_count already exists.
C7 SQLite source context verifies a real database transaction and issues BEGIN only when the driver has not already begun it; unsupported isolation/driver fails safely. Real WAL writer and stream-failure cleanup tests pass for legacy/modern sqlite3 control, plus PostgreSQL repeatable streams. Supported Python matrix.
C1 Global union and Best Sellers fallback returned by the existing single joined read; no second head lookup. Missing/corrupt fallback classified. Controlled publication/retention test retains one original identity. Deployment query latency qualification with the same lanes and database.
C14 Run-ID validation before persistence or telemetry binding; safe 422 error and unchanged valid lookup behavior. Eight malformed-input cases pass with and without observability; generated OpenAPI inspection confirms the ErrorResponse 422 schema. Final integrated gates.
C16 Reuses finite-positive numeric validation at direct and serialized deployment boundaries, including boolean rejection. Configuration file passes 63 cases. Final integrated gates.
C15 Canonical grids and total parameter ordering; closest preferred fallback, quality-first selection, explicit tuple tie breakers. Permutation/quality cases pass; model identity recommendations-v2 preserves historical comparison boundaries. Supported runtime matrix and final integrated checks.
C10 Empty feature payload retains the authoritative sequence/version header; rebuild increments its version and never recreates erased state. Three expiry/idempotency regressions pass on SQLite. PostgreSQL partial/full expiry and a barrier-controlled concurrent rebuild/append pass. Final integrated checks, and bounded maintenance addressed separately by S3.
C8 Both deployed roots compose an immutable per-scope retention policy, including disabled capabilities. Real persisted 30/90-day expiry, idempotency cleanup, unknown-scope fallback, and shorter-policy legacy reconciliation are covered. Injected API stores reject policy disagreement. S3 now supplies bounded reconciliation and the read-time barrier. Final release qualification.
C9 Expired business interactions advance the acknowledged sequence/version without storing expired ledger payload or updating signals. Boundary cases at ±1 microsecond under 30/90 days and late lifecycle barriers pass. Legacy-expiry regressions retain explicit old-writer rows. S3 now hides existing expired features before cleanup. Final release qualification.
S3 Expiry/deletion budgets, durable bounded rebuild checkpoints, version-checked publication, and read-time dirty/expiry/policy protection implemented. SQLite interruption/restart and 1,201-row legacy-key evidence pass; PostgreSQL concurrent append preserves the tail with bounded two-row chunks. Migration 0008 invalidates legacy features while retaining headers and refuses unsafe downgrade. Aggregate capped backlog, batch outcomes, expired counts and duration are emitted without Shopper dimensions. Final integrated gates; production query plans, throughput and lock/latency qualification under Q1 remain unproven.
C2 Feature-poor metadata emits explicit empty lanes; positive-price and text-presence checks avoid empty matrices. Word and character vocabularies are independent, preserving usable character and structured features without swallowing encoder errors. Six metadata boundary cases and two backend publication cases pass; unexpected encoder failure preserves the head. Eighty ordinary Items retain exactly equal scores and rankings against HEAD. Final release qualification.
C3 The immutable Catalog eligibility view reaches behavioral generation, tuning, raw-pair baselines, local/category lanes and global popularity before top-N selection. Original support remains unchanged. Twenty-four backend/lane/exclusion cases and three popularity cases pass; real worker tests retain primary evidence and valid global popularity after 200 excluded neighbors. Representative query plans and high-exclusion throughput under Q1; final release qualification.
C5 Pinned native API has no secondary comparator. Bounded products retain all positive matches within each column block; a K-entry selector uses exact score and lexical rank, excluding self first. Literal 210-Item cutoff cases pass across permutations, threads and block budgets; exhaustive close-score/zero/self and product/object bounds pass. Full-profile dense-overlap runtime and feature/output memory qualification remain unproven.
S5 Destination categories sort once per evidence partition; source categories retain shared immutable candidate tuples through the keyed provider. Five regression cases preserve ordering/provenance and stop construction at 200. Balanced 1k/2k/4k probes preserve complete ranked-output signatures while constructing 400 shared candidates instead of 200k/400k/800k. Production category distribution and final release qualification.
S8 The only production caller consumes the adapter's validated ordered stream. Adjacent deduplication uses one previous ID; two independent K-entry heaps retain score/lexical winners, constructing only survivors. Six fan-out/tie cases up to 50k unique rules pass; fresh 100k-rule probe preserves exact output with approximately 95 KB temporary Python allocation. Actual operational fan-out and full-profile qualification; no deployment severity inferred from the local probe.
C11 Completed loads are reaped before checking cache/capacity regardless of requested head. Success, missing, corrupt, reader and executor failure cases permit a successor across the same or another Scope. One-loader reservation also bounds synchronous direct callers; artifact I/O runs outside the scheduler lock. Explicit close joins outside the lock and is owned by the deployed API root on normal/failed exit. Sustained cold-load and multi-process latency/residency qualification; no request latency SLO claim from event-controlled tests.
C13 Both metadata and learned HNSW artifacts must declare native inner product. Checksummed, correctly shaped L2 payloads are rejected before publication; existing bad artifacts retain same-Snapshot HTTP fallback. Native serialization already binds metric identity, so no envelope or migration changes. Cross-platform artifact compatibility and production qualification; incompatible artifacts require rebuilding.
C12 Ten purchase IDs remain neighborhood seeds, while all retained purchase IDs feed step-level exclusion. Eleven/100-purchase cases cover exclusion enabled/disabled, absent seed metadata/neighborhoods, deduplication and truthful underfill. Final release qualification.
S7 Exact ANN scores bounded 8,192-vector chunks and retains a K-entry score/lexical heap. A bounded per-chunk threshold skips strictly losing scores while retaining all ties. Exclusions apply before retention; matches are constructed only for final survivors. Thirty-two 10k-eligible-Item cases cover dense ties, metadata zero, learned negative and untied scores, heavy exclusions, reverse ID mapping, varied chunks and public native-underfill fallback. Fresh 10k/100k probes preserve exact baseline output. Untied scans have higher latency than baseline native partitioning; production fallback latency and native vector/index residency remain unqualified.
M1 Pair budget counts unique directed pairs at their strongest existing weighted signal, preserving reverse directions and deterministic tuple order. Budget-sized map and compacted heap bound selection state. Eight cutoff/permutation cases, nine exhaustive-reference cases, repeated updates, unique training minimum, v1 rejection and clean serving imports pass. Six fresh MPS runs use identical temporal inputs/projections/budgets per arm, removing 218–252 duplicate slots. Representation is now torch-two-tower-v2. Controlled comparison reports mixed quality changes; no general relevance/causal improvement, production resource or broad-enablement claim.
E5 Closed metric-summary state rejects single-sided hidden values, nonfinite/out-of-range available rates, nonpositive available counts, invalid count types and unknown status. Literal suppressed/not-applicable JSON retains null metrics. Final release qualification; no database migration or existing metric field rename.
E4 Thresholds and operational measurements reject nonfinite/range-invalid values at construction. Missing measurements raise explicit unavailable-evidence failure. Finite boundary ordering is preserved; a prospective consumer test holds exposure on validation failure. Offline-only contract; no live rollout controller claimed or added.
E1 One observation per scoped Shopper key/strategy is validated before any metric accumulation, including conflicting cohort labels and feedback mode. One strategy per report; each input sequence has an explicit positive integer unit bound (default 10k). Distinct 9/10-unit literal oracles, Catalog identity, feedback isolation and bound failures pass. Report labels strategy and evaluation_schema_version 2, distinguishing unit counts from prior unversioned replayable row counts. Large evaluations require external streamed grouping. Future multi-strategy aggregates need explicit semantics; no unbounded service registry introduced.
E2 Exposure identity includes the exact variant; attribution verifies complete stored assignment equality. First exposure remains immutable and forced-digest collisions fail closed. Variant/config/experiment/unit/key and all three Commerce Scope dimensions have independent regression evidence. Offline helper only; no live attribution service or causal experiment claim.
E3 Domain-separated UTF-8 JSON arrays replace ambiguous delimiter concatenation for assignment and exposure HMACs. Assignments bind full Commerce Scope and explicit hmac-sha3-512-json-v2 identity. Independent OpenSSL synthetic Unicode/delimiter vectors verify literal digests and allocation bucket. Prior algorithm versions are rejected. Identity change requires a fresh experiment/config cohort; no ambiguous legacy verifier or secret rotation.
S9 Single-owner offline exposure batches require explicit Scope, positive capacity, attribution window, retention and aware clock. Fully retained entries cannot be evicted to admit another identity; operation-time expiry frees slots, repeated exposure preserves the first timestamp, and close clears evidence. Deadline, re-exposure, saturation, malformed time and 1,000-operation bounded-residency cases pass. No implicit live retention policy, durable/concurrent adapter or restart guarantee; maintainers must select these before live use.
C4 Policy resolution distinguishes a nonempty matched limit tuple, no selector match and missing required partition. Only no-selector traffic can use allow_unmatched; missing partitions deny before state access. Both flag values retain parent tokens/concurrency and issue no lease. API contract retains the existing unavailable 503 body with no Retry-After. Local admission only; no distributed enforcement claim.
S1 Explicit positive integer maximum_rate_buckets; capacity denies new partitions without parent debit. A removable indexed heap retains one full-refill deadline per bucket, rechecks rounding before removal and never evicts depleted state. Idle 10k, hot refresh, denial storm, rollover, thread-contention, exact refill and floating-rounding cases pass. Cap is caller-selected, not an application deployment default. State reset/restart cannot preserve rate debt. Production memory and expiry-burst latency remain unqualified.
S2 Explicit maximum_concurrency_permits counts retained rule permits. A global indexed deadline heap, direct per-lease ownership and bucket-local AVL cumulative weights replace global scans/rebuilds/sorts. Partial expiry, weighted retry, repeated/late release, active duplicate IDs, saturation, same-deadline bursts and 2k-operation independent-reference decisions pass. Global-scan traps pass with 100/1k/4k unrelated live leases. Extra per-permit memory is measured below. Local lock and expiry bursts remain; no cross-process, durable or production latency claim.
S4 Both stores stream unique directional holdout edges in canonical order. Evaluation keeps top-K matching state, distinct relevance counts, constant-size aggregate/cohort sums and Catalog-bounded coverage. WorkEvidence borrows repeatable streams within the store context instead of retaining full truth; tuning, category filters and baselines use them. Literal parity, fetch/Arrow-budget bounds, interrupted reads, scratch cleanup and preserved-head failures pass. Repeated native sorting/scanning costs more CPU. Dense local measurements below do not qualify full-profile runtime or whole-process memory; Polars native sorting remains outside its batch budget.
S6 Positive strict route-class byte caps assembled after protected authentication and before parsing. Declared oversize rejects before reads; wrapped receive stops on the crossing chunk. One stable 413, unchanged persistence, exact boundaries, inaccurate/no lengths, invalid JSON, disconnect/cancellation and OpenAPI/config/wiring/observation evidence pass. Interaction default assumes 2 MiB pending any maintainer preference. No independent purchase Item-count maximum. Local concurrent accepted-body measurements below exclude server buffers and do not qualify production memory or latency.
H1 Maintainer accepted the existing post-acknowledgement rule. SQLite/PostgreSQL post-ack reads suppress history; three barrier-controlled PostgreSQL overlap cases preserve the prior lifecycle until that read finishes, then observe acknowledged opt-out, deletion or reauthorization on subsequent reads. No semantic strengthening or cancellation guarantee. The accepted overlap rule is documented and verified.
Q2 Import inventory mapped 80 shipped modules. Validator now enforces six stable-policy modules, four application service seams, incoming composition-root restrictions, known relative/from-import submodules and literal dependency cycles. One explicitly deferred storage-to-state-store edge remains allowed. Temporary-source RED/GREEN cases and production source gate pass. Static literal-import guarantees only; dynamic imports, implicit parent initializer execution and runtime/resource behavior remain separate review/evidence.

H1's decision is resolved. Q1's execution was reduced by the maintainer; its measured failed observation and remaining operational qualification are recorded in the final checkpoint. Q2 implementation evidence is recorded below. No production-capacity claim is established by these tests. Final release checks are recorded in the latest checkpoints below.

Commands and results

Python invocations used the configured Python 3.14.7 project environment and uv, with a PyCharm environment preflight before each invocation. Routine read-only cat, sed, rg, git diff and git status inspected contracts and changes. Some guessed test/module paths were absent; the actual lifecycle, API and observability seam paths were used.

  • make test-focused TEST=tests/integration/test_lifecycle.py::test_recovered_publisher_survives_stale_publisher: RED, exit 2 from Make; successor insert failed foreign-key integrity after stale build deletion.
  • make test-focused TEST=tests/integration/test_lifecycle.py: GREEN, first 10 then 15 cases passed as recovery-phase and live-clock cases were added.
  • make test-focused TEST=tests/contract/test_source_adapter.py::test_sqlite_source_read_is_repeatable_across_queries: RED, legacy mode exposed writer B; modern control passed. Full source file GREEN, 11 passed.
  • make test-focused TEST=tests/unit/test_serving.py::test_for_you_uses_one_snapshot_during_publication: RED, original metadata with successor fallback. Full serving file GREEN, 14 passed before adding four classification/empty-fallback cases; the widened gate passed all 18.
  • make test-focused TEST=tests/integration/test_personalization_storage_candidates.py: GREEN, five passed including real WAL publication and retention during the read.
  • Docker daemon was initially unavailable. open -a Docker succeeded. Started a dedicated disposable recommendations-remediation-postgres container using postgres:17-alpine, a 2 GiB limit and loopback-only ephemeral port. Readiness and port inspection succeeded. No existing databases or containers were reused. It has no production credentials/data.
  • TEST_CONTROL_DATABASE_URL=<disposable local URL> make test-focused TEST=tests/integration/test_postgresql_storage.py: GREEN, eight passed, including all four new recovery phases.
  • Focused Ruff correction first reported missing exception/parameter documentation and long lines; corrections made. Focused format applied, then .build/restore_unchanged_formatting.py restored unchanged AST blocks to baseline spelling to avoid unrelated formatting drift.
  • make test after the first publication/source/serving slice: exit 0; Ruff, mypy (78 source files), documentation (62 Markdown files), and architecture checks passed; 525 unit passed, six skipped; 46 contract/delivery passed; 53 integration passed, 26 PostgreSQL deselected. Existing unit fixtures emitted 18 SQLite datetime-adapter deprecation warnings.
  • make test-focused TEST=tests/unit/test_config.py::test_worker_poll_interval_must_be_finite: RED, NaN/infinity/boolean accepted. Full configuration file GREEN, 63 passed.
  • make test-focused TEST=tests/contract/test_api.py::test_inspect_run_rejects_invalid_identifier_before_telemetry: RED, eight HTTP 500 responses. Full API file GREEN, 24 passed.
  • uv run --frozen python scripts/export_openapi.py .build/remediation-openapi.json: exit 0.
  • make test-static after C14/C16 and claim-capture refinement: exit 0; all static gates passed.
  • make test-focused TEST=tests/unit/test_popularity.py: RED, four permutation/tie cases failed; after selection fix GREEN, nine then ten cases passed including quality-before-default selection.
  • Second make test: exit 0; all static gates passed (63 Markdown files); 541 unit passed, six skipped; 54 contract/delivery passed; 53 integration passed, 26 PostgreSQL deselected.
  • TEST_CONTROL_DATABASE_URL=<disposable local URL> make test-postgres: exit 2; 25 passed, one failed. The unchanged synthetic CLI append fixture constructs today's cutoff, then requests 2026-01-10, which is before its base source. The generator correctly rejects that interval. git diff confirmed both that test and synthetic production code are untouched. This is a baseline date-dependent qualification failure, not evidence of a storage regression. No skip or production workaround was introduced; the complete PostgreSQL gate remains unproven.
  • make test-focused TEST=tests/integration/test_personalization_storage.py: RED, all three new cases lost acknowledged state; after preserving headers GREEN, nine cases passed.
  • Later TEST_CONTROL_DATABASE_URL=<disposable local URL> make test-focused TEST=tests/integration/test_postgresql_storage.py: 11 then 12 cases passed after source repeatability, header expiry and concurrent rebuild/append coverage was added. Full source-adapter file repeated after failure-cleanup coverage: 11 passed.
  • Third make test initially stopped on Ruff's nested-context rule in the new cleanup test; the contexts were combined and the gate rerun: exit 0; all static gates passed, 541 unit passed with six existing skips, 54 contract/delivery passed, and 56 integration passed with 30 PostgreSQL cases deselected. Total 651 passed; 18 existing SQLite adapter warnings.
  • Final diff inspection restricted idempotency-header recovery to missing legacy profiles, avoiding an unnecessary acknowledgement scan during ordinary maintenance. The focused personalization storage file repeated: nine passed. make test-static repeated: exit 0.
  • The normal gate was rerun after that refinement: exit 0, again 651 passed and six skipped.
  • Composition-root inspection found the API entry point still used package version as model identity. Added test_api_entry_point_submits_current_model_identity: RED, returned 0.1.0 instead of recommendations-v2. The entry point now uses the explicit algorithm identity; telemetry service_version continues to identify the installed package.
  • After the entry-point correction, the API contract file passed 25 cases and final make test exited 0: all static gates passed, 541 unit passed with six existing skips, 55 contract/delivery passed, 56 integration passed, and 30 PostgreSQL cases were deselected. Total 652 passed.
  • git diff --check: exit 0. Source diffs inspected for lease ownership, current-clock checks, source transaction cleanup, same-head fallback, monotonic headers, and parameter tie policy. Unrelated formatting was restored; no generated tracked file drift remains.
  • C8 composition-root regression: RED, both API and worker persisted 90-day deadlines for the configured 30-day scope. After wiring the immutable policy, the storage file passed 11 cases.
  • C9 boundary regression: RED, four of six cases retained expired affinity; two inside-window cases passed. After acknowledgement-only handling and preserving explicit legacy rows for C10, the storage file passed 20 cases, including shortened-policy reconciliation and old lifecycle events.
  • make test-static first failed new documentation/import/line/default lint checks; after focused corrections it exited 0 (Ruff, mypy 78 files, documentation 63 files, architecture).
  • The PostgreSQL storage file repeated after C8/C9: 12 passed. The normal make test gate passed 663 tests (541 unit, 55 contract/delivery, 67 integration), with six existing skips.
  • Added preservation of current signals across all four late business types, absent-authorization rejection for late business/lifecycle commands, and injected-store policy mismatch coverage. A test-only misspelled feature field was corrected; the focused storage file then passed 27 cases.
  • Final make test after those additions and a test-only line-length correction: exit 0; all static gates passed, 541 unit passed with six existing skips, 55 contract/delivery passed, 74 integration passed, and 30 PostgreSQL cases were deselected. Total 670 passed.
  • Final git diff --check: exit 0. No migration was added in C8/C9.
  • S3 budget regression: RED, expire lacked batch/work limits. Added dirty projection state, scoped checkpoints, bounded erasure, and scanned-row sequence continuation. Focused storage file first passed 28, then 32/33/34 cases as interruption, append, erasure, telemetry, and the 1,201-row legacy-key tests were added. Test-only context/API field mistakes were corrected.
  • make migration-check: first two existing round trips passed, then three passed after adding old-profile invalidation, header preservation, refusal to downgrade dirty work, rebuild and safe downgrade coverage. The three-case gate repeated against the final tree: pass.
  • Focused Ruff import/format commands applied to affected files. The ignored AST restoration script restored unchanged baseline blocks. make test-static passed after lint corrections.
  • PostgreSQL storage file widened to 13 cases: pass. Its coordinated two-row checkpoint test resumed across a concurrent append on independent connections, bounded every reducer query, and recorded coordination wait and a measured upper bound for header-lock duration. The final focused concurrency case repeated after refining duration measurement: pass.
  • First widened make test passed 541 unit tests with six skips, then failed one delivery grammar test: Ruff's py314 target had removed exception tuple parentheses from the worker. Restored all four compatible handlers and aligned Ruff's target with supported Python 3.12. The minimum-grammar focused test and make test-static then passed. No supported version changed.
  • A subsequent gate stopped at one new test's line length; corrected that line. Final make test: exit 0, all static gates pass, 541 unit passed with six existing skips, 55 contract/delivery passed, 82 integration passed, 31 PostgreSQL cases deselected. Total 678 passed. PostgreSQL storage file repeated separately: 13 passed.
  • Editable rebuild caused only generated SOURCES.txt terminal-newline drift; inspected and restored that generated artifact to HEAD. git diff --check: pass.

  • C2 make test-focused TEST=tests/unit/test_content.py::test_similarity_handles_unusable_metadata: RED, four failed and two passed (empty matrix and absent word vocabulary). After the fix, the complete content file passed ten cases. End-to-end file passed 23 cases, including both aggregation backends publishing valid empty metadata lanes and unexpected encoder failure preserving the head. make test-static: exit 0.

  • C3 work-store cutoff regression initially failed 24 cases because the eligibility contract did not exist. The real worker regression failed on both backends: valid B appeared through metadata fallback instead of co-view primary evidence. After the eligibility joins/semi-joins, the fixture's group order was corrected; all 24 lane/exclusion cases and both worker cases passed. Three global/category popularity cases were added. The first make test stopped on two helper docstrings and an import order; corrected. Final gate: exit 0, all static gates, 574 unit passed with six existing skips, 55 contract/delivery and 87 integration passed. Total 716 passed; 31 PostgreSQL cases deselected.
  • C5 cutoff regression RED: four failed and two passed across input permutations/threads, including literal 009 instead of 001. Local inspection of pinned sparse-dot-topn's API found no secondary comparator. After bounded exact selection, the content file passed 16 then 26 cases as block-budget, exhaustive close-score and construction-bound evidence widened. Added two invalid-budget cases; the full gate passed all 28 content cases.
  • .venv/bin/python .build/metadata_remediation_probe.py parity: exit 0; 80 ordinary Items match HEAD's complete rankings and scores, maximum score error 0.0. Fresh sparse/dense modes each used 2,000 Items, one thread and the default 100,000-key block budget. Both exited
  • Sparse: 0.035 seconds, 632 maximum returned product entries, 6,324 maximum returned-product bytes, 199,557,120 process peak RSS bytes. Dense: 1.326 seconds, 99,856 entries, 800,116 returned product bytes, 265,846,784 peak RSS bytes and 400,000 retained candidates. Both axes were at most 316. Returned-product bytes do not measure all simultaneous native buffers; process RSS includes imports, features and output. No full-profile extrapolation or capacity claim.
  • C5 make test: exit 0, all static gates; 592 unit passed, six existing skips, 55 contract/delivery and 87 integration passed, 31 PostgreSQL deselected. Total 734 passed.
  • S5 make test-focused TEST=tests/unit/test_category_strategies.py::test_category_pair_candidates_are_shared_without_changing_rankings: RED, four cases proved repeated destination sorting. After keyed sharing, the category file passed 15 cases; a three-category construction-cutoff case was then added. Fresh .venv/bin/python .build/category_remediation_probe.py 1000, 2000, 4000: all exit 0, identical complete ranked-output signatures including confidence/provenance and respectively 100k/200k/400k final FBT entries. Baseline candidate allocations were 200k/400k/800k; current allocations were 400 in every probe. Local build times were 0.142/0.444/1.393 seconds before and 0.000275/0.000414/0.000647 after; these small concurrently launched probes establish parity and allocations, not deployment throughput. Gate first stopped on an import order; corrected. Final make test: exit 0, 597 unit passed, six skips, 55 contract/delivery, 87 integration; total 739 passed, all static gates and geographic/category integration cases passed.
  • S8 fan-out regression RED: six failures from per-row candidate construction. After bounded selection, corrected the constructor spy to accept both positional and keyword calls. make test-focused TEST=tests/unit/test_compatibility.py: exit 0, seven passed. Fresh .venv/bin/python .build/compatibility_remediation_probe.py 1000, 10000, 100000: all exit 0, exact output parity on duplicated ordered streams with separate partitions. With K=200, current temporary Python peaks were 92,615/92,743/94,727 bytes versus 371,271/3,615,559/39,668,031 baseline bytes. Current elapsed times were 0.003/0.024/0.266 seconds versus 0.010/0.081/0.914 baseline. Inputs were allocated before tracemalloc; this measures temporary Python selection/output allocations, not total process RSS or native memory. No operational severity or full-profile qualification claim.
  • Focused uv run --frozen ruff format calls on generation source/tests succeeded. The ignored baseline-AST restoration helper preserved unrelated formatting after each call. Source, test and algorithm/operations diffs were inspected. git diff --check: exit 0.
  • Final S8 make test: exit 0; Ruff, mypy (78 source files), docs (63 Markdown files), and architecture checks passed. 603 unit passed with six existing skips, 55 contract/delivery and 87 integration passed, 31 PostgreSQL cases deselected: total 745 passed. The 18 existing SQLite datetime-adapter warnings remain. No migrations changed in this generation slice; migration and PostgreSQL evidence from S3 remains recorded above. Full release make smoke and make verify are deferred until the remaining implementation scope is handled.
  • C11 focused ANN file passed 15 cases, including ten completed-head/failure cases, blocked synchronous loading and joining shutdown. API file passed 27 cases, including normal and failed server-exit ownership. C13 metric regression RED: two correctly checksummed L2 representations did not raise; after validation, the ANN file passed 17 cases.
  • C12 purchase-exclusion regression RED: the two enabled-exclusion cases retained purchase 11; disabled cases passed. After separating exclusions/seeds, the Swimlane file passed ten cases.
  • The first retrieval make test stopped on Ruff's suppressible-exception rule; corrected. A subsequent full gate and make test-integration reached a native SIGSEGV in the first end-to-end metadata product after the ANN file. The new C13 test had introduced eager top-level Faiss loading before the existing metadata native libraries. A fresh three-Item subprocess in .build/test_ann_native_import_order.py reproduced this import-order failure without loader/API state: Faiss-first crashed in the OpenMP barrier; metadata-first passed (one failed, one passed). Moved the test's Faiss import after the actual artifact builder, preserving the existing production lazy order. No production workaround/thread-budget or dependency change. make test then passed all static gates, 607 unit (six existing skips), 57 contract/delivery and 101 integration; total 765, 31 PostgreSQL deselected.
  • S7 regression RED: all twelve initial tie cases failed, constructing 9,963 matches for K100. After chunk/heap selection, the ANN file passed 29 cases. Added untied/heavy-exclusion cases; it passed 49. The full gate passed 797 tests before the cutoff optimization. An untied probe showed the first loop was slower; bounded threshold filtering reduced its 100k runtime from approximately 88 ms to 12 ms. The unchanged rankings and 49-case ANN file passed again.
  • .venv/bin/python .build/ann_exact_remediation_probe.py <10000|100000> <tied|untied|exclusion_heavy>: all six fresh final subprocesses exited 0 with complete baseline ranking/score parity against HEAD's exact selector, K100, two-dimensional float32 vectors, one native build thread and reversed Product IDs. Temporary Python peaks at 10k were 125,320/125,344/125,320 bytes current versus 2,305,864/181,384/245,193 baseline. At 100k they were 247,584/177,640/247,584 current versus 22,820,776/1,620,488/2,116,864 baseline. Current times were 27.7/10.3/6.6 ms at 10k, 341.3/12.3/111.3 ms at 100k; baseline 44.5/1.1/9.7 and 446.3/2.9/94.1 ms respectively. Inputs and native artifact residency were allocated before tracemalloc; it measures temporary Python allocations, not native buffer memory. Whole-process peak RSS ranged 219–339 MB and includes imports, indexes, artifacts, baseline and current runs. Single local observations establish boundedness/parity and expose the latency tradeoff; no throughput/SLO/capacity claim.
  • Focused Ruff formatting/checking passed after correcting import order and long lines; the ignored restoration helper preserved baseline spelling of unchanged AST blocks. Source/API, tests and governing documentation diffs inspected; git diff --check passed.
  • Final retrieval make test after the threshold optimization: exit 0; Ruff, mypy (78 source files), docs (63 Markdown files) and architecture passed. 607 unit passed with six existing skips and 18 existing datetime-adapter warnings, 57 contract/delivery and 133 integration passed; total 797, 31 PostgreSQL deselected. After adding the final performance/handoff prose, make test-static and git diff --check both passed. No migration changes in this slice; full PostgreSQL, runtime matrix, installed-wheel smoke, make smoke/make verify and formal capacity/availability qualification remain pending as described above.
  • M1 make test-focused TEST=tests/unit/test_two_tower_torch.py::test_two_tower_pairs_deduplicate_across_signals: RED, six failed/two passed; duplicate A→B displaced reverse B→A and distinct A→C at the cap. After bounded unique selection, the file passed 12 then 21 cases as exhaustive-reference and update-state evidence widened, finally 23 including the unique minimum and fresh serving import. Representation changed to torch-two-tower-v2; correctly checksummed v1 models are rejected.
  • .venv/bin/python .build/m1_two_tower_comparison.py <baseline|unique> <17|29|41>: six fresh subprocesses exited 0 on MPS, 32 dimensions, model seed 17, pair budget 10k, three epochs/batch 128, unmodified smoke source and a fixed 28-day temporal holdout ending 2026-09-30. The first scratch attempt supplied offline before online purchases and correctly failed ordered-group validation; corrected to the adapter's online-first channel-aware order. Only derived aggregates persist inside the bounded temporary work store. Baseline replays HEAD's _pairs; both arms use the current unchanged model/loss/projection/index builder. Six small aggregate JSON records remain ignored under .build, with no raw source/Shopper/context identifiers.
  • A configured-interpreter JSON audit verified equal derived input/projection hashes, split, budget, hardware/library versions and evaluated-anchor denominators for each source seed. Baseline selected 9,782/9,773/9,748 unique pairs from 10k slots; current 10,000 for every seed. The comparison record preserves loss, Recall/NDCG, training/build time, Python allocations, process RSS and pair coverage, including all limitations. Quality is mixed, not evidence of a general improvement. No production speed, memory, conversion or causal uplift claim. Some local verification overlapped part of seed 29; timings are single observations with instrumentation, not an isolated benchmark.
  • M1 make test: exit 0, all static gates, 628 unit (six existing skips), 57 contract/delivery, 133 integration; total 818, 31 PostgreSQL deselected. The new comparison Markdown passes the 64-file documentation gate. Source/tests/docs diffs inspected and git diff --check passed.
  • E4/E5 make test-focused TEST=tests/unit/test_personalization_evaluation.py: RED, 32 failed/26 passed. NaN latency returned no breaches and single-sided suppressed metrics passed construction. After closed/finite validation the file passed 63; finite rate/latency range cases widened it to 69. Final make test: exit 0, all static gates, 691 unit (six skips), 57 contract/delivery, 133 integration; total 881, 31 PostgreSQL deselected. After linking final comparison guidance, make test-static passed. No schema migration or live rollout controller.
  • E1 focused RED: 12 failed/70 passed. Replaying one observation ten times produced available count 10; label changes and multiple strategies were accepted; bounds/schema identity absent. After validation/bounded single-strategy evidence, the file passed 82, preserving literal Recall/NDCG and Catalog-scoped units. Prior unversioned row-count evidence is incompatible with evaluation_schema_version 2. Focused Ruff checks first flagged imports/long lines; formatting corrected them and the ignored restoration helper preserved unchanged baseline AST spelling.
  • Final E1 make test: exit 0; all static gates passed (78 mypy source files, 64 Markdown files). 704 unit passed, six existing skips and 18 existing datetime-adapter warnings; 57 contract/delivery and 133 integration passed. Total 894 passed, 31 PostgreSQL cases deselected. No migrations changed in this slice. Final handoff prose was checked again with make test-static and git diff --check; both passed. Full release/runtime/PostgreSQL/capacity checks remain pending for the reasons recorded above; this goal is not complete.

  • E2/E3 focused RED: two failures and 82 passes, reproducing cross-variant attribution and delimiter-collision identity. After canonical scope-bound frames and complete assignment checks, focused tests passed; additional scope, Unicode/vector and lifecycle cases widened the file to 115 passing cases. An intermediate test export lacked datetime serialization (108 passed, one fixture failure); explicit ISO serialization corrected the test export.

  • Two independent /opt/homebrew/opt/openssl@3/bin/openssl dgst -sha3-512 -mac HMAC synthetic JSON vectors exited 0 and supplied literal assignment/exposure digest oracles. No real keys or Shopper identities were used. The assignment vector independently selects bucket 6998.
  • .venv/bin/python .build/exposure_capacity_remediation_probe.py 1000 and 10000: exit 0. Trusted HEAD retains 1,000/10,000 entries (343,504/3,345,194 temporary Python peak bytes); the explicit three-entry batch retains three (4,543/7,041 bytes). These fixed-clock synthetic probes establish bounded residency, not durable/live semantics, RSS or latency objectives.
  • E2/E3/S9 static checks initially flagged new docstring returns/raises and import/style issues; corrected them. make test-static and git diff --check passed. Full make test: exit 0, all static gates; 737 unit passed with six existing skips and 18 existing datetime warnings, 57 contract/delivery and 133 integration passed. Total 927 passed; 31 PostgreSQL deselected. No migration changed in this offline slice.

  • C4 make test-focused TEST=tests/unit/test_serving_limits.py: RED, one failure/12 passes; allow_unmatched=True admitted a request missing a required partition. After closed resolution, the file passed all 13 cases, including unchanged unmatched exemption and parent conservation. make test-focused TEST=tests/contract/test_api.py: 29 passed, with both malformed flag values producing the existing 503 unavailable body before downstream work, no lease or retry header. make test-static: exit 0, all static gates. Resolution now returns a nonempty matched tuple or the exported PolicyResolutionFailure enum; no schema, configuration default or denial code changed.

  • C4 full make test: exit 0; Ruff, mypy (78 source files), documentation (64 Markdown files) and architecture gates passed. 739 unit passed, six existing skips and 18 existing datetime warnings; 59 contract/delivery and 133 integration passed. Total 931 passed; 31 PostgreSQL deselected. No migrations or Compose artifacts changed. Final git diff --check passed; full release/runtime/PostgreSQL/capacity checks remain pending as recorded above.

  • S1 focused RED: the 10k idle-bucket test failed with 10,001 resident buckets after full refill. After indexed reclamation and explicit bounds, the admission file passed 24 tests; API file passed 30 including unavailable state-capacity semantics. Static checks first found four long lines, corrected; all static gates then passed (79 source files). No local constructor exists in deployed source: caller-selected caps replace an implicit unbounded optional adapter.

  • S2 global-scan RED: all three 100/1k/4k populations failed the global registry iteration trap. Indexed permits/weighted retry passed the widened file's 39 cases. Ruff then flagged self-type annotations/property wording; future annotations and property wording fixed them. Focused formatting was followed by the ignored AST restoration helper to preserve unchanged spelling. Static checks passed (80 source files). A synthetic arithmetic probe found the literal 1/1.9 rounding edge (projected 0.9999999999999999), now covered with nextafter-boundary evidence.
  • .venv/bin/python .build/admission_state_comparison.py <baseline|current> <hot|partitions|rate> <N>: 22 fresh-process invocations exited 0. Both arms ran each mode at 1k/2k/4k; current hot/partitions also ran at 16k, and both rate arms ran at 16k. A separate JSON assertion invocation verified identical complete decision signatures for all ten paired profiles and complete lease cleanup. HEAD is the trusted baseline; only synthetic descriptors, fixed clocks and deterministic IDs are used. Timings include tracemalloc and a lock wrapper; Python peaks include harness-owned duration and history lists. One sample per arm/profile does not establish RSS, confidence intervals, deployment capacity or production SLOs. No 16k baseline concurrency comparison is claimed.
Local comparison Population HEAD/current total seconds HEAD/current temporary Python peak bytes HEAD/current lock p95 microseconds
Hot concurrency: admit, denial, release 1k 1.266 / 0.119 695,032 / 1,048,807 768.9 / 32.9
Hot concurrency: admit, denial, release 2k 4.902 / 0.240 1,389,568 / 2,085,247 1,572.4 / 33.6
Hot concurrency: admit, denial, release 4k 30.546 / 0.496 2,606,168 / 4,163,031 5,714.7 / 35.5
Partitioned concurrency: admit, release 1k 0.504 / 0.075 689,507 / 1,190,292 426.3 / 32.0
Partitioned concurrency: admit, release 2k 1.838 / 0.140 1,378,867 / 2,364,979 822.5 / 30.5
Partitioned concurrency: admit, release 4k 8.210 / 0.288 2,710,212 / 4,720,931 2,075.2 / 32.1
Rate churn: fully refilled partition reclamation 16k 0.796 / 1.047 7,619,130 / 1,022,268 16.5 / 33.8

Current-only 16k hot and partitioned concurrency took 1.993/1.152 seconds, with 16,639,479/18,869,715 peak Python bytes and 37.4/32.7 microsecond lock p95. Both retained exactly 16k live leases at maximum and zero after release. Rate churn maxima were 500–501 current buckets versus all 1k/2k/4k/16k HEAD buckets; after the final idle interval, current retained one newly spent bucket versus HEAD's N+1. Indexed concurrency trades higher per-permit memory for eliminating unrelated work; safe rate expiry adds CPU. Cap/headroom selection and same-deadline expiry-burst latency still require deployment qualification.

  • Final make test-focused TEST=tests/unit/test_serving_limits.py: 42 passed; extra evidence covers the rounding boundary, thread contention and tree-height/deletion bounds. make test-focused TEST=tests/contract/test_api.py: 31 passed, including both rate-bucket and permit saturation as the same safe 503 code/message with no Retry-After. Full make test: exit 0; Ruff, mypy (80 source files), documentation (64 Markdown files) and architecture passed. 768 unit passed with six existing skips and 18 existing datetime-adapter warnings; 61 contract/delivery and 133 integration passed. Total 962 passed, 31 PostgreSQL deselected. Final source/ASGI/index diffs inspected and git diff --check passed. No migrations, Compose artifacts or hosted configuration changed; broader release/runtime/PostgreSQL/capacity checks remain pending as already recorded. The whole goal remains active.

  • S4 integration RED: make test-focused TEST=tests/integration/test_end_to_end.py::test_training_evaluates_without_materializing_holdout_truth failed at the full graph load in behavioral tuning. After borrowed ordered streams and an extended work-store context it passed for both aggregate engines, covering all grid/final/category and baseline paths. The store now closes before ANN construction; serving/public result DTOs retain no stream or graph handle.

  • Twelve literal store/partition/fetch cases initially failed because the new test rows were not in canonical group order; sorted the synthetic fixture without changing production ingestion. All twelve then passed, preserving duplicate relevance, crossing-group temporal assignment, missing/empty predictions, rank duplicates, cohorts and Catalog coverage. The pure evaluator remains unchanged and independent. Pure evaluation file passed nine cases, including 10k/100k generated-edge residency under 100 KB, unsorted/failing stream closure and no partial report.
  • Focused work-store file: 97 passed/two existing skips before adding the final disk-budget variant. It includes a cursor trap proving only three-row fetches, early close/reuse, Polars no-collect guards and batch-budget cleanup. Four interrupted-stream RuntimeError/MemoryError engine cases passed and preserved the old head. Full end-to-end file: 30 passed, including all existing backend/parallel metric and Snapshot parity checks. Final Polars batch/artifact budget cases pass.
  • Focused Ruff checks flagged imports, long lines, lambda assignment and a test loop binding; fixed these and used the ignored AST restoration helper after focused format. make test-static then passed all gates (80 source files, 64 Markdown files); no public metric/schema changes.
  • .venv/bin/python .build/holdout_stream_comparison.py <duckdb|polars> <reference|stream> <Catalog size> <fetch size>: 13 fresh-process invocations exited 0. Both engines/arms ran at 128 and 512 Items (16,256 and 261,632 directional edges); streaming also ran at 1,024 Items (1,047,552 edges) and at fetch 32 for 512 Items. The thirteenth repeated the Polars small-fetch probe after fixing Parquet row groups at 8,192 independently of fetch size. Twelve final profiles remain. Two separate JSON oracle commands verified all aggregate/cohort/count/coverage parity within 1e-12 and the complete-graph literal Recall=20/N, NDCG=(N-1)/N, coverage=20/N. No raw interactions are generated or staged: the scratch probe inserts synthetic derived pair deltas in 512-row batches.
Dense truth comparison, 512 Items Pure reference / stream wall seconds Reference / stream temporary Python peak bytes Reference / stream full-process peak RSS bytes Base / observed streamed owned scratch bytes
DuckDB, fetch 8,192 0.145 / 0.503 38,723,199 / 2,739,635 285,949,952 / 239,779,840 5,551,040 / 5,551,040
Polars, fetch 8,192 0.149 / 0.580 29,358,937 / 1,857,940 310,280,192 / 285,114,368 722,431 / 902,326

At 1,024 Items/1,047,552 edges, streamed Python peak remains 2,739,795 bytes (DuckDB) and 1,858,077 (Polars), with process RSS 354,238,464/449,331,200 bytes and wall time 2.267/2.278 seconds. These larger streamed profiles have the literal metric oracle, not a claimed unrun reference arm. Polars native allocations exceed the nominal 256 MB batch budget, which is not a process-memory limit. A 32-row fetch gives DuckDB 24,168 peak Python bytes at 1.442 seconds; fixed-group Polars 169,142 bytes at 2.041 seconds and 281,837,568 process RSS bytes. Coupling Parquet groups to that fetch initially grew owned scratch to 6,243,949 and RSS to 374,030,336 bytes; fixed groups restore owned scratch to 902,326 and preserve the reference metrics.

Measurements are one fresh sample per final profile with tracemalloc and periodic owned-file inspection. CPU/RSS/owned-file values are retained in ignored JSON; native temporary spill is not sampled continuously, and no complete peak-spill or production training-budget claim follows. Streaming trades CPU/native ordering and a temporary derived-only Polars artifact for eliminating Python graph residency. Native query, Catalog and prediction-map residency still require Q1.

  • Final S4 make test: exit 0; Ruff, mypy (80 source files), documentation (64 Markdown files) and architecture passed. 799 unit passed with six existing skips and 18 existing datetime warnings; 61 contract/delivery and 138 integration passed. Total 998 passed, 31 PostgreSQL deselected. No migrations or Compose artifacts changed. Final stream/lifetime/diff inspection and git diff --check passed. Final handoff/static prose check repeated successfully.
  • The 512-Item measured CPU seconds were 0.145/0.503 reference/stream for DuckDB and 0.210/0.739 for Polars. Larger stream CPU was 2.265/2.858 seconds at 1,024 Items. These single instrumented samples include query/evaluation but exclude graph preparation; no full-run speed improvement or qualification is implied.

Current maintenance implementation and next seam

S3 uses additive migration 20260930_0008: dirty flag, earliest signal deadline, history policy, and a scoped reducer checkpoint. No deployed schema was changed. Old writers must be drained before migration/runtime rollout; older binaries cannot safely serve these projections. Downgrade refuses dirty work; keep personalization admission disabled for an older-binary rollback.

Expiry deletes at most 1,000 interaction and 1,000 idempotency rows, with at most 100 affected interaction keys. Rebuild visits at most 100 due keys and scans at most 1,000 ledger rows total. The sequence limit includes expired rows, not just contributing rows. Manual rebuild has the same 1,000-row ceiling. Dirty feature payload is omitted while sequence/version/lifecycle state remain visible. Healthy current projections need not be invalidated by deleting already-excluded old payloads. Checkpoints preserve an immutable prefix when a newly acknowledged tail extends it; expired contributing signals, policy change, or a non-append header change require restart. Header-locked publication still checks the captured current version. Erasure removes profile and checkpoint immediately, then deletes at most 1,000 ledger/idempotency payload rows per pass. Receipts remain pending until every payload is erased.

H1's overlap decision remains pending. Its approved design requires distinguishing post-acknowledgement reads from overlaps; the baseline guarantees the former, and exploratory overlap alone is not a breach. Do not silently promise cancellation of ranking begun before a barrier. All twenty-eight confirmed findings now have implementation evidence, along with provisional S8, S9, M1 and Q2. H1, Q1 and final release qualification still belong to the active goal. The S6 interaction default is a stated 2 MiB assumption pending any maintainer preference; an asynchronous question remains optional.

Request-body bounds and local evidence

RequestBodyLimits is a strict runtime-only deployment section excluded from training identity. Training and Swimlane defaults are 8,192 bytes; Interaction defaults to 2,097,152 bytes. Canonical 255-byte fields with maximal ASCII JSON escaping fit the small defaults; 1,000 unique escaped 255-byte purchase identifiers fit the interaction default. There is no finite maximum for all previously accepted purchase shapes or normalization/whitespace representations. The new wire cap therefore explicitly narrows that behavior instead of claiming an existing Item-count bound. The optional preference question has received no answer; 2 MiB is the stated current assumption.

The ASGI guard adds no payload buffer. Declared length comparison uses decimal byte-string ordering, so enormous decimal headers do not depend on Python integer conversion limits. Actual bytes are counted regardless of length headers. FastAPI's registered HTTP-exception handler owns streamed 413 generation, before the server-error layer; authentication and observations wrap the guard. Disconnect/cancellation remain propagated by the transport without draining input. The guard cannot prevent a native server from delivering one oversize chunk; accepted bodies remain assembled by FastAPI up to the configured cap. A fully authorized padded interaction is rejected without a profile/version change; retrying its original body commits sequence one, proving no partial write.

Commands in this slice (all Python/Make commands followed SDK preflight):

  • make test-focused TEST=tests/contract/test_api.py::test_request_body_limit_stops_chunked_input_before_parse: RED, baseline accepted an 8 MiB padded training request with 202. First implementation run failed on an incorrect route-enum name; corrected run GREEN, one passed, two reads and no run created.
  • make test-focused TEST=tests/unit/test_request_body.py: 18 passed.
  • make test-focused TEST=tests/unit/test_config.py: 88 passed.
  • make test-focused TEST=tests/contract/test_api.py: initial fixture-name/content-type mistakes corrected; 47 passed. After adding entry-point/observation evidence, an incorrect recording-wrapper assumption was corrected; 49 passed. These were test construction errors, not product failures.
  • make test-focused TEST=tests/contract/test_api.py::test_oversize_authorized_interaction_leaves_sequence_and_idempotency_unchanged: one passed.
  • Focused .venv/bin/python -m ruff check --select I --fix and ruff format on the six affected source/test files, then on API tests: passed. .venv/bin/python .build/restore_unchanged_formatting.py restored unchanged baseline blocks after each format pass; both runs completed before gates.
  • make test-static: initial import/line-length and then default-constructor-call findings corrected; GREEN, Ruff, mypy (81 source files), documentation (64 Markdown files) and architecture passed.
  • make test: 1,059 passed (842 unit, 79 contract/delivery, 138 integration), six existing skips, 31 PostgreSQL deselections and 18 existing SQLite datetime warnings. Final checks below supersede this run after the authorized-interaction persistence regression was added.
  • Final make test: exit zero, 1,060 passed (842 unit, 80 contract/delivery, 138 integration), six existing skips, 31 PostgreSQL deselections, 18 existing SQLite warnings; all four static gates passed with 81 source and 64 Markdown files. API contract file now has 50 passing cases.
  • .venv/bin/python .build/request_body_residency.py 1, 8, and 32: each exit zero. Fresh processes used the real FastAPI guard and JSON parser, 16 KiB receive chunks, exact 2 MiB bodies and a barrier holding all accepted bodies concurrently. Only aggregate measurements were retained; no bodies or credentials were written. These are single-sample local transport probes, excluding network/proxy buffers, full product schema objects, database work and native serving concurrency.
  • git diff --check: exit zero; source/config/guard changes inspected with the accumulated goal diff.
Concurrent accepted bodies Python allocation peak bytes Process maximum RSS bytes Wall seconds
1 4,245,721 57,868,288 0.0042
8 19,055,848 73,203,712 0.0270
32 69,811,416 125,304,832 0.1019

These results establish local bounded accepted-body residency and its concurrency dependence, not a total RSS guarantee or REQ-024 qualification. Upload duration and concurrency remain separate operational controls. No migrations or Compose artifacts changed in this slice; final migration, installed-wheel smoke, disposable PostgreSQL gate, supported-runtime matrix and Q1/Q2 remained pending at the S6 checkpoint. The Q2 evidence below supersedes that part of the status.

Profile-consistency audit and architecture checks

The post-acknowledgement oracle starts reading only after the lifecycle commit returns. Both SQLite and PostgreSQL initially expose usable history, then return an acknowledged suppressed lifecycle state and no usable profile through PersonalizedRecommendationService.read_authorized_profile. Opt-out can be followed by an independently acknowledged reauthorization; deletion completion returns the persistent deleted tombstone with empty affinities. These are GREEN BASELINE tests, not a newly fixed privacy breach. No production profile-read code changed in this slice.

The ignored .build/profile_read_overlap.py probe ran only after the PostgreSQL test process ended. It used SQLAlchemy query hooks, separate real PostgreSQL connections and bounded thread events to capture the profile statement before another transaction committed. Under READ COMMITTED the later suppression statement ran after commit. Opt-out/deletion overlaps returned old active version 1 despite acknowledged version 2; a reauthorization overlap returned old opted-out version 2 despite acknowledged version 3. Subsequent reads returned version 2 opted-out/deletion-pending and version 3 active, respectively. Only lifecycle states, versions and synchronization facts were retained; no Shopper key or payload was written. The task database schema was cleaned after the probe. The explicit required overlap-rule question is pending. Work depending on semantic strengthening must wait for its answer; do not interpret elapsed time as approval.

The AST inventory found 80 shipped modules and one literal cycle: storage consumes a deferred personalization repository factory while that repository imports storage table definitions. The new graph resolves actual absolute/relative from-import submodules. Policy rules include deferred imports; cycle rules exclude only that specific storage-to-state-store edge when syntactically deferred in a function or literal TYPE_CHECKING guard. Promoting it to immediate import fails. Policy modules allow standard library and the six named stable modules. The four application seams cannot import transport, training, merchant-source or direct HTTP/SQL framework details. No shipped module imports API/worker composition roots. Existing registry/Snapshot/ANN seams remain permitted. Rules, exceptions and guarantee limits are recorded in architecture, testing and agent-readiness docs.

Commands in this slice (SDK preflight before every Python/Make invocation):

  • make test-focused TEST=tests/integration/test_personalization_storage.py::test_profile_read_after_opt_out_acknowledgement_is_suppressed: two passed, unchanged production behavior.
  • Docker inspect revalidated the task-owned PostgreSQL container as running on loopback port 51790.
  • TEST_CONTROL_DATABASE_URL=<task disposable URL> make test-focused TEST=tests/integration/test_postgresql_storage.py::test_postgresql_profile_read_after_barrier_acknowledgement_is_suppressed: two passed. Full tests/integration/test_postgresql_storage.py then passed all 15 cases.
  • .venv/bin/python .build/dependency_inventory.py: exit zero; 80-module literal graph and exact six-policy import inventory retained in ignored aggregate JSON; no application code was executed.
  • make test-focused TEST=tests/delivery/test_architecture_validator.py: first RED eight failures and seven baseline passes; GREEN 15 passed after policy/cycle implementation. Application-rule RED added 28 expected failures; GREEN 45 passed. Incoming-root RED added six expected failures; final focused GREEN 51 passed. Existing generic telemetry/reusable-package rules remain covered.
  • make architecture-check: both interim production-source runs exited zero; the final normal gate also validates the incoming-root rule.
  • Focused Ruff import fixing and formatting on the validator and three test files: passed. .venv/bin/python .build/restore_unchanged_formatting.py restored unchanged baseline blocks and completed before gates. make test-static initially found two long fixture strings; splitting literals fixed them. GREEN: Ruff, mypy (81 source files), docs (64 Markdown files), architecture.
  • .venv/bin/python .build/profile_read_overlap.py: exit zero, all three controlled overlaps and their post-acknowledgement reads characterized in ignored aggregate JSON.
  • git diff --check: exit zero; validator and accumulated source/test changes inspected.
  • Final make test: exit zero, 1,113 passed (842 unit, 131 contract/delivery, 140 integration), six existing skips, 33 PostgreSQL deselections, 18 existing SQLite warnings. All static gates passed with 81 source and 64 Markdown files. The separate disposable PostgreSQL storage file also passed 15 tests; the entire PostgreSQL gate is still pending.

An initial Q1 surface read confirmed qualification drift that requires explicit reconciliation: the current v1 profile has 200k Catalog Items, 200M views and 25M purchases, while its generated-source guide says 100M views/5M purchases and the approved remediation/service capacity target requires at least 100M views/100M purchases. The qualification CLI defaults to the service minima and the source CLI intentionally does not expose the qualification profile. No profile quota was silently changed, and no full-profile load was started. This host has 32 GiB RAM and roughly 683 GiB free disk, but a host/disk/time budget and approved qualification environment are still required before expensive execution. Continue with bounded smoke evidence and final installed/runtime/PostgreSQL gates.

Decisions, artifacts and authority

  • Existing recovery_count provides an immutable publication generation; no snapshot schema change is necessary because staging writes and cleanup are keyed by unique snapshot ID.
  • Publication uses the worker's injected live clock, never the fixed generation timestamp. Direct callers default to UTC wall time and must provide the claimed recovery generation.
  • Drain old workers before relying on fencing. Keep serving the previous head if publication is stopped; restoring unfenced publishers is not a safe rollback.
  • Drain existing old-model pending/running jobs before the model-identity transition; an older run must not be silently computed by a changed implementation under its original identity.
  • PostgreSQL storage cases repeated against the final source tree: 13 passed.
  • .build/ contains ignored red/green scratch and generated OpenAPI; it is not committed source.
  • The dedicated disposable PostgreSQL container remains available for the continuing goal at loopback port 51790. Recheck its live state before use; do not assume a recorded port proves readiness. Stop this task's container after verification is complete.
  • The release checkpoint below supersedes earlier pending PostgreSQL, installed-wheel and runtime checks. Full capacity and availability qualification remain unproven; do not mark the full goal complete without the scope audit and required decisions.
  • User authority covers local implementation and verification. It does not authorize deployment, package publication, merchant data mutation, external messages, or hosted settings.

Release checkpoint, factory preference and bounded Q1 evidence

The user's additional coding preference is to use a factory when a conditional selects an object, message or response to construct. Apply named factory functions at the construction boundary and reuse existing factories. create_personalization_attempt now owns complete-runtime selection between authorized personalization and ordinary fallback; create_app calls it unconditionally. The reusable admission middleware delegates denial status, message, encoding and retry headers to _create_denial_response before sending ASGI events. Both preserve the existing observable contracts; no abstract factory hierarchy or cross-package application dependency was introduced.

The release audit found two test-environment defects:

  • A PostgreSQL append fixture used today's default cutoff before attempting a January 2026 append. Its base cutoff is now explicitly January 1, 2026. Production validation and the existing exact appended-count, artifact recovery and idempotency assertions remain intact.
  • One heartbeat telemetry test used SQLite StaticPool, sharing one physical connection across concurrent worker and heartbeat transactions. The first pinned full gate reported a running run instead of successful publication; focused and twelve-process-case attempts did not reproduce it. A separate deterministic transaction probe showed another logical connection's rollback discarding the uncommitted write on the shared connection. File-backed SQLite with independent physical connections preserved isolation. The test now injects an owned file-backed store through the existing factory and disposes it afterward. Heartbeat interval and all success/correlation/phase/ renewal assertions are unchanged. This follows the documented concurrent-connection limitation of SQLAlchemy StaticPool.

The real sibling telemetry checkout differs from CI's pin and has unrelated dirty work. The normal locked diagnostics therefore rejected its metadata. An exploratory make lock changed five telemetry metadata lines; those changes were inspected and reversed. uv.lock remains unchanged. An ignored isolated layout now copies this working tree beside an archive of telemetry commit e65ad90bd7bbd718536a355060a731a7faecff73, preserving the shared checkout and original SDK. The isolated environment uses the configured Python 3.14.7 executable through uv. CI-equivalent wheel tests use separately created 3.12.10/3.13.15 environments; the IDE SDK was not changed.

Final pinned make verify passes all static checks, 842 unit tests, 131 contract/delivery tests, 140 integration tests, and the installed-wheel smoke check (1,113 tests total; six existing skips, 18 existing SQLite warnings, 33 PostgreSQL deselections). The separate pinned disposable PostgreSQL gate passes all 33 cases. Migration checks pass three cases. The earlier installed wheel passes 885 unit/contract cases and 59 canonical-source/lifecycle/migration/end-to-end cases separately on Python 3.12 and 3.13, with source-tree imports rejected. The final rebuilt-wheel matrix containing both factories also passes on Python 3.12.10, 3.13.15 and 3.14.7: 885 plus 59 cases per version, six existing skips and 18 existing SQLite warnings. All three isolated installed-package reports reject source-tree imports and exercise both aggregation backends and all console entry points.

Bounded generated-source evidence

The ignored remediation_source_to_serving_smoke.py runs the unmodified v1 smoke profile at seed 23, cutoff September 1, 2026: 1,000 Catalog Items (950 eligible), 25,000 views, 8,000 online and 2,000 offline purchases, 570 compatibility rules, 90 days and 20 categories. Each fresh process generates and verifies a separate owned SQLite merchant source, submits through the actual API, runs the worker and normal parameter grids, verifies the complete 11-strategy manifest and published head, and checks both planted Frequently Bought Together/Also Viewed outcomes through HTTP. Source counts match durable run evidence; source canonical reads use the scope-prefixed order indexes. Each snapshot contains 10,011 sets. Forty in-process HTTP reads per process retain their own snapshot identity. Owned source/control databases and derived scratch are cleaned after each run.

Backend Fresh runs Worker training Total measured wall Process peak RSS Sampled derived peak
DuckDB 3 10.40–10.78 s 13.61–13.98 s 685–711 MiB 23.6 MiB
Polars 5 11.32–12.23 s 14.57–16.09 s 666–735 MiB 3.7 MiB

Configured pipeline budgets are two threads, 512 MB and 1 GB derived disk; each process has a 300-second timeout. Four Polars runs explicitly set POLARS_MAX_THREADS=2; the first is a separate default-pool characterization, not a controlled thread comparison. Total sampled source/control/ derived disk peaks at 109.2 MiB; published control data occupies about 99.9 MiB. Sampling at 50 ms can miss short peaks, and native-library/query allocation lies outside the pipeline memory budget. Peak RSS includes loaded native libraries but excludes process startup from the measured wall/CPU interval. These numbers are not extrapolated to full scale or used as a production acceptance test.

All runs have the same Generation Manifest and stable metric signatures. All three DuckDB runs have identical full serialized snapshot signatures. Polars serialized scores vary: a bounded streamed comparison of two retained derived snapshots found 1,274 differing score entries in 982 sets, with maximum absolute difference 8.881784197001252e-16 and relative difference 3.808290836571061e-16. All 10,011 sets retain exactly equal outcomes, rankings and provenance. The Polars ranking/provenance signatures also match DuckDB. Exact score byte reproducibility is not established for Polars; no rounding or production scoring change was introduced to hide this.

Six additional single-process DuckDB cases vary one workload/configuration dimension at a time. They use the same seed/cutoff/resource budgets, retain separately labeled profile/config identities, and pass receipt/counts, all 11 strategies, both planted HTTP outcomes and scratch cleanup. Catalog/view/purchase variants use explicit bounded profile versions, not a changed named default. Wide-group and single-grid cases retain the baseline source manifest. The wide-group case admits all previously excluded groups: baseline exclusions were two view groups/301 rows and four purchase groups/1,202 rows; limits of 400 reduce both exclusions to zero. This measures group handling on the same source, without claiming a controlled independent overlap or eligibility experiment.

Case Change Worker time Process peak RSS Published sets
Baseline 1k Items, 25k views, 10k purchases; normal grids/limits 10.59 s 713 MiB 10,011
Catalog 2k Items 17.50 s 964 MiB 16,341
Views 50k views 13.65 s 771 MiB 14,626
Purchases 20k purchases 11.58 s 737 MiB 10,938
Group limits View/purchase distinct-Item limits of 400 13.44 s 821 MiB 10,163
Grid size One choice in each parameter-grid family 10.38 s 689 MiB 10,011

These are one run per variant, not repeated statistical comparisons or performance acceptance. API and worker use explicit injected configuration and separate stores; no deployed config identity or comparable-run drift history is asserted across cases. Raw source data is confined to owned temporary merchant databases, which are removed; only aggregate/hashes survive in ignored reports.

Identities retained by the aggregate summary:

  • Source (80 Python files), pyproject and lock SHA3-512: 315043410a366f09745bc09fd11d347ac8225dbfbbaeeba979761979f82134d2f23487bf8febb0883a059d945ca3d4cd89a3709cea115d02d0078d7c0440da34.
  • Generation Manifest SHA3-512: fe504b473763502a374041383bee4c5aa44f197df5e8f29cd01c49fdd46fa4dd5e4c51281fc28bdc50ca786dc419a80f084a8a60258ac7f862c75c3159fd4aaf.
  • Complete ranking/provenance SHA3-512: b89e18a2190251d03b01208ba6362bc705753667309989ac387928a7d86ae3bb3825a244a8cb9b60bfb142b9697bc7993b58ad7d4ad899d7924dd1afbaa0ef39.
  • Final rebuilt wheel SHA3-512 (1,141,962 bytes): 84eb7285994341bfe27f876ad20ca8bfd48bda24c410d3031ca69e68872cc49b4fc528cdeaa62f003bb698434763cee95f88811faca32f75765d9f42439bf232.
  • Final sdist SHA3-512 (1,109,859 bytes): 493567df0731ba2f0f9cfef766d3984c0fa93951806a3bebb4bd0950a2530cbd9053ba64dd5287aabd70d7e489c5dd6717b665097c0c9dbc5efc861599e5d5f4.

The generated-source guide now records actual v1 qualification quotas (200k Catalog Items, 200M views, 25M purchases) and the actual 120M-interaction development profile. It explicitly records that v1 qualification does not meet the approved 100M-purchase capacity target. No generation profile/version was silently changed. A versioned replacement must preserve old manifest replay/ append semantics and qualification eligibility; profile_for currently resolves only by name.

Commands in this release slice

SDK preflight preceded Python/Make toolchain invocations. Routine git, rg, sed, Docker inspect, disk inspection and rsync calls inspected state or copied into the ignored isolated layout.

  • Original make doctor: failed because sibling telemetry metadata disagreed with the lock. make lock: exited zero; its five metadata changes were subsequently reversed. The temporary metadata-aligned doctor passed but is not counted as CI-pin evidence.
  • Original make smoke BUILD_DIR=.build/remediation-wheel-3.14: first failed locked export; after the temporary metadata alignment it built/installed successfully but the custom nested build directory broke the supported smoke script's one-level relative path. The corrected invocation uses the default .build inside the isolated layout, with no Makefile change.
  • make migration-check: three passed. Original disposable make test-postgres: 32 passed, one dated fixture failed. Focused corrected append test: one passed.
  • git archive of the exact telemetry pin plus rsync of this working tree created the isolated layout. Pinned make setup and make doctor: exited zero without lock changes.
  • First pinned make verify: static/unit/contract passed, integration 139 passed and the shared- connection heartbeat test failed. Focused heartbeat characterization then passed once. .build/heartbeat_concurrency_characterization.py: twelve memory and twelve file cases passed; these repeats do not erase the original failure. The deterministic SQLite transaction-isolation characterization exited zero and proved shared-connection rollback interference.
  • Pinned disposable make test-postgres: 33 passed, 140 deselected. The Docker container was revalidated as running at loopback port 51790 before use; no schema-mutating probes overlapped it.
  • After the fixture/factory changes, focused end-to-end: 30 passed; serving unit: 18 passed; API contract: 50 passed; admission unit: 42 passed. Two focused make test-static invocations: Ruff/mypy (81 files)/docs (64 files)/architecture passed. Profile correction make docs-check: passed. git diff --check: passed.
  • Pinned make verify after the first factory: exited zero, 1,113 passed plus installed-wheel smoke. Final pinned make verify after both factories: exited zero, same counts plus installed smoke.
  • uv python list --only-installed: confirmed Python 3.12.10, 3.13.15 and 3.14.7 locally available. .build/remediation-runtime-matrix.sh exported locked extras, created isolated environments, installed the same wheel, rejected source imports and tested it on 3.12/3.13: 885 plus 59 passed on each. The final rebuilt-wheel matrix adds 3.14 and passes 885 plus 59 cases on all three runtimes. Its first launch refused existing task-owned venvs; adding explicit uv venv --clear recreated only those owned environments. The corrected process exited zero; log: .build/remediation-final-installed-matrix-green.log.
  • Eight separate uv run --frozen python .build/remediation_source_to_serving_smoke.py processes (absolute script path from the pinned layout, backend/report arguments per case): exited zero. Three DuckDB and five Polars runs verified all bounded oracles. Four Polars runs and the final DuckDB run explicitly set POLARS_MAX_THREADS=2.
  • .venv/bin/python .build/compare_smoke_snapshot_scores.py: exited zero; streamed aggregate score/ranking comparison recorded. Only two owned derived control snapshots were retained, each under a 150 MiB limit; no raw interaction source or Shopper payload was exported.
  • .venv/bin/python .build/remediation_evidence_summary.py: exited zero; source identity and common manifests/metrics, exact DuckDB output and cross-backend ranking signatures verified.
  • .venv/bin/python .build/remediation_package_identity.py: exited zero; original/mirror source, pyproject and lock identities match the smoke build, and both candidate archive digests recorded.
  • Six additional sequential smoke-script processes with --backend duckdb --case set to catalog-2x, views-2x, purchases-2x, wide-groups, single-grid, baseline: exited zero. Each uses POLARS_MAX_THREADS=2 and the same explicit pipeline budgets/timeout. The scaling summary script exited zero and independently verified unchanged source manifests for the two config variants, changed identities for changed quotas, and the expected exclusion reductions.
  • Generated editable SOURCES.txt changes were inspected and restored to HEAD. The IDE automatically adds an exclusion for the ignored mirror venv and immediately re-adds it after a surgical removal. That single local generated exclusion remains; the configured original SDK and other module settings are unchanged. Remove it when cleaning the owned mirror venv after qualification work.
  • Release handoff make docs-check: passed (64 Markdown files). Final git diff --check: passed.

Remaining work and decisions

  1. Final installed-wheel matrix, docs check and final diff check are complete. The lock and generated SOURCES.txt are unchanged. The IDE's automatic mirror-venv exclusion is described above; clean it with the owned environment at the end of qualification work.
  2. Continue Q1 bounded repeated scaling and remaining dimensions (overlap, eligibility, geography, concurrent isolation and failure/maintenance progress). Single cases now cover Catalog/view/ purchase volumes, group limits and grid size. Cold/warm ANN latency and full-profile source/ control query plans remain unqualified.
  3. Prepare the versioned full-capacity profile/replay contract and a reviewable execution budget. This host has 32 GiB RAM and approximately 682 GiB free disk; Docker has 12 CPUs and about 15.6 GiB RAM. An expensive full run still requires the explicit host/disk/time budget and disposable qualification environment required by Q1's operational boundary. A budget question is pending: defer the full run while continuing bounded scaling, or use this host for at most 12 hours and 256 GiB disk with a separate owned PostgreSQL environment and versioned capacity profile. No expensive run is authorized by elapsed time.
  4. H1's explicit overlap semantics decision remains pending. Post-acknowledgement suppression continues to pass; do not strengthen overlapping-read semantics without that decision.
  5. Record availability and production latency targets as unqualified until sustained operational evidence exists. Local tests, injected clocks and in-process HTTP are not substitutes.

The full goal remains active. No package was published, service deployed, hosted setting changed, merchant data mutated, or message sent externally.

Qualification preparation and additional bounded evidence

This checkpoint supersedes the earlier statement that profile_for resolves only by name. The current qualification preset is v2: 200,000 Catalog Items, 200 million views, 100 million purchases, 730 days and 43 categories. Explicit profile_for("qualification", version="v1") retains the exact historical parameters and eligibility metadata. Smoke/development stay at v1. Only exact registered qualification presets with the existing fixed seed/cutoff/capabilities qualify for eligibility metadata; custom quotas and unknown versions do not. Eligibility remains distinct from a Qualification Claim, and historical v1 still fails today's 100-million-purchase capacity floor.

The versioned factory and registry change is preparation under Q1, not a full-scale run. No generation algorithm, manifest schema, append rule, qualification-script minimum or CLI profile exposure changed. Replaying artifacts also requires their recorded generator version. Focused tests preserve the old literal v1 parameter identity and cover current v2 quotas, explicit selection, unknown versions, changed quotas and fixed-cutoff/override eligibility.

Concurrent Scope evidence and SQLite boundary

Two fresh processes exercised two simultaneous real workers with PostgreSQL 17.11 control storage and separate owned SQLite merchant sources. Sources intentionally share Tracking ID, Catalog ID and Product IDs, while differing in Data Source and evidence volume/seed. Barriers confirmed both source contexts were open, both runs were RUNNING and neither head was published before release. Both workers published distinct complete heads with all 11 strategies. Forty planted reads per Scope, 12 invalid-selector reads, per-Scope counts and ranking/provenance signatures passed. Each source saw five training SELECTs and zero serving SELECTs. Joined head/snapshot/strategy-count polling observed only complete available heads. Polling is not a proof of every intermediate transaction; publication-fencing tests provide that contract evidence.

Training took 23.40/23.45 seconds; whole-process wall time was 29.44/29.49 seconds. Peak RSS was approximately 996/1,029 MiB. The second process measured 35,995,648 control bytes by summing table/materialized-view relation totals. The first size calculation also included indexes as separate relations and double-counted them; do not compare that first byte value. Both owned PostgreSQL schemas were dropped and source/scratch directories removed. Reports are .build/remediation-concurrent-postgres-{1,2}.json; source labels, counts and hashes survive, not raw Browsing Session or Order IDs. This is bounded thread/in-process HTTP evidence with an injected clock, not production-equivalent concurrency, latency or availability qualification.

The initial file-backed SQLite concurrency probe failed, and an instrumented repeat identified two workers claiming the same pending run. A separate deterministic two-connection, source-free characterization coordinated the conditional claim UPDATE and confirmed duplicate claim returns while one other run stayed pending. SQLite omits the PostgreSQL row-lock clause; the current allocation path does not check that UPDATE's affected-row count before returning its claim. Technical implementation documentation now explicitly limits SQLite control storage to a single worker. Production remains PostgreSQL. No production claim semantics or unsupported SQLite concurrency guarantee was introduced to make the probe pass.

Independent source projections

Three new first-run DuckDB cases retain the baseline physical source, fixed seed/cutoff and resource settings while changing declared adapter projections independently:

Projection Observation Worker time Published sets
No geography columns Every set uses global geography 8.64 s 7,603
Effective eligibility 50% 500 eligible Items; no excluded Item in published entries 7.03 s 6,072
Shared view-Item pool Distinct viewed Items decrease from 752 to 212 7.33 s 8,726

The shared-pool case preserves each Browsing Session's distinct-Item width: independently streamed ordered width counts and their checksum agree before/after the projection. Planted signal Items are preserved. This varies cross-session overlap while keeping widths; it is not an unmodified named profile. The eligibility case's effective projection is 500 Items, while the physical Catalog remains 950 eligible Items. Query digests and these distinctions are recorded rather than attributing projected inputs to the unchanged Generation Manifest alone. All three cases passed receipt/counts, complete-strategy, eligibility, planted HTTP and cleanup oracles. These runs and the concurrent probes used the preceding wheel, identified in the prior checkpoint, before the profile-registry change. New-build evidence must be recorded separately.

Reviewable full-run envelope — awaiting the existing budget answer

The proposed ceiling is this Mac (32 GiB RAM), 12 hours total and 256 GiB of owned disk. This is an execution stop budget, not a new service SLO: source generation consumes it, whereas the approved 12-hour Training Run target starts at acceptance. A run stopped by the total budget cannot establish or refute that training target. Use a separate owned PostgreSQL environment, never the existing public test schema or merchant data. Record the chosen PostgreSQL version, container/storage limits and authority before provisioning. Provision the exact released v2 preset through the generation API; verify its Generation/Oracle Manifests and Materialization Receipt before submitting a Training Run. Keep the 200k/100M/100M service minima unchanged.

The fixed released cutoff is January 1, 2026. Today's ordinary service clock would discard part of that source through the rolling two-year window, so a historical qualification environment needs a declared, advancing clock shared by API and worker. A candidate mapping is the released cutoff plus elapsed monotonic time from a common epoch; real wall time still enforces all execution limits. A frozen clock would invalidate lease-expiry evidence. Validate the generated interval against the worker's calendar-year window at claim time and record the 730-day profile's exact historical coverage; do not silently reinterpret it as a separately proven two-calendar-year source. The clock/environment selection remains qualification setup, not a change to the deployed service or authorization to alter persisted run timestamps.

Before an expensive execution, confirm the outstanding host/disk/time and disposable-environment answer, finish the bounded harness's clock/window and stop/cleanup checks, and record exact CPU/RSS/disk ceilings. Serve over loopback TCP with separate API/worker processes and reuse scripts/qualify.py for submission/status/serving observations. Preserve all current checks: volume minima, complete atomic publication, source-independent serving, evidence semantics, failure preservation and checksummed build/environment identities. Record unsupported capacity, production-equivalent latency and sustained monthly availability as unqualified. H1's explicit overlap-rule decision also remains pending; post-acknowledgement behavior is unchanged.

Commands in this checkpoint

SDK preflight preceded Python/Make invocations. Routine state reads, Docker inspection and rsync copied only into the owned ignored layout and preserved the original sibling telemetry changes.

  • Concurrent smoke on SQLite: initial and instrumented attempts exited one. The source-free characterization initially used a nonexistent probe API and failed; corrected .create characterization exited zero and confirmed the duplicate-claim observation above.
  • Two PostgreSQL concurrent probe invocations exited zero. Each created/dropped only its owned unique schema. The second corrected relation-size accounting; first sizes remain labeled.
  • Three installed-wheel projection probe invocations exited zero, with the per-case oracles above.
  • Versioned-profile RED: five failed, 13 passed. First GREEN had one test defect from calling a dictionary property; correcting that assertion yielded 18 passed. The final cutoff eligibility parametrization adds one case, checked by the complete suite below.
  • make test-focused TEST=tests/unit/test_synthetic_schedule.py: 28 passed, checking bounded schedule arithmetic for the new quota without generating the full source.
  • make test-static: Ruff, mypy (81 source files), docs (64 Markdown files) and architecture passed.
  • Pinned make verify: log shows all stages completed: 851 unit tests passed (six skipped), 131 contract/delivery tests passed, 140 integration tests passed (33 PostgreSQL deselected), wheel/sdist built and installed-wheel smoke passed. The process finished; its execution handle was unavailable after context recovery, so no recovered shell exit-code claim is made.
  • Pinned disposable make test-postgres: exited zero; 33 passed, 140 deselected.
  • .build/remediation_profile_v2_identity.py: exited zero; root/mirror source, pyproject and lock identities match. New wheel/sdist hashes are recorded in .build/remediation-profile-v2-package-identity.json; preceding evidence was not overwritten.
  • Profile-v2 installed-wheel runtime matrix exited zero: Python 3.12.10, 3.13.15 and 3.14.7 each passed 894 unit/contract cases (six skipped) plus 59 focused integration cases; installed smoke rejected source-tree imports and exercised both aggregation backends on each runtime.

The full goal remains active. Full-profile execution, sustained availability and the H1 decision are outstanding. No expensive execution was launched, package published, service deployed, external data changed or hosted setting modified.

Additional ANN factories and qualification checkpoint

The user's conditional-construction preference also applies to the ANN seams inspected during qualification. _create_query_encoder now owns metadata/learned encoder selection after the existing supported-representation/vector validation; _create_load_executor owns creation of the optional one-worker background executor. Constructors call these named stateless factories. Validation order, metadata/learned scoring, fallback, cache capacity and shutdown ownership remain the existing contracts. No new class hierarchy, deployment setting or artifact version was added. The focused ANN integration file supplies behavioral coverage for both representations, loader mode/concurrency, corruption, generation transitions and shutdown; no factory-private test added.

These changes produced a separate candidate from the profile-v2 checkpoint:

  • Source, pyproject and lock SHA3-512: b4e7199c0b358aaeceb5c8eecd61dea3b9218dbc3babcac43f0dba19b879fd75d3e207607ee971b54bdfe6acb292269e6c325a919eea18c0dc0d3eabe976951a.
  • Wheel SHA3-512 (1,142,413 bytes): 41f50d4b95d4a5c1958c8751f068af113f72ad38c8657cc388d74a3faa58dfcd583a998bcf9e0c20d08fb04e905e79f789444cbf428d58baa1fff96f1060386c.
  • Sdist SHA3-512 (1,110,296 bytes): 727707ea8473dbd9211b237f1a1c13eb5128fdef92d21f8377032359f12ebb8a4ca99b362a12d8528d017ceb70c7c8f8d96215920d9c1d2dde60275025162785.

The profile-only build identity remains in its own report. Its three fresh baseline processes retain exactly the preceding build's Generation Manifest, stable metrics, full serialized snapshot, rankings and provenance. Each of the three source-projection cases also passes on DuckDB and Polars, with equal projected counts, complete manifests and rankings/provenance across backends. Projection timings are local observations, not an established backend speed comparison.

Published ANN cold/warm evidence

Six sequential fresh processes (three per size) exercised actual published metadata artifacts and the final candidate's background retriever. Each process published two successive generations under a one-resident-generation limit. Cold calls returned the ordinary empty fallback while loading; 200 warm queries per generation performed no further artifact reads. Independent dot products checked returned scores against published vectors; every result obeyed eligibility, anchor exclusion and scope isolation. Foreign Scope calls returned empty without artifact I/O. All 12 generation transitions and 2,400 warm queries passed, and all owned databases were removed.

Catalog Items Processes Cold call return Ready after scheduling Warm p95 per generation Peak process RSS
1,000 (950 eligible) 3 0.011–0.104 ms 3.28–3.84 ms 0.156–0.166 ms 169–175 MiB
10,000 (9,500 eligible) 3 0.027–0.157 ms 17.87–20.93 ms 0.250–0.305 ms 280–282 MiB

Each size retained identical artifact and complete query-ranking signatures across its three processes. Native, OpenMP, OpenBLAS and Accelerate thread settings were explicitly two for these controlled repeats. Build time medians were 0.60/1.27 seconds. An earlier successful 1k probe did not set BLAS environment limits and is retained separately, excluded from this comparison. These measure direct retriever calls using SQLite control storage and metadata representations. The OS page cache was not flushed. Peak RSS includes artifact construction and publication. They do not establish Serving API latency, learned-index behavior, production residency, capacity or availability. Dependency/Python versions, per-generation p50/p95/p99, artifact sizes and exact settings remain in ignored aggregate reports.

Commands and final scope

  • make test-focused TEST=tests/integration/test_ann_snapshot.py: exited zero, 49 passed.
  • Final pinned make verify: exited zero; static checks, 851 unit tests (six skipped), 131 contract/delivery tests, 140 integration tests (33 PostgreSQL deselected), wheel/sdist build and installed-wheel smoke passed. Log: .build/remediation-ann-factories-pinned-verify.log.
  • Final pinned disposable make test-postgres: exited zero; 33 passed, 140 deselected.
  • Final rebuilt-wheel runtime matrix: exited zero; 894 plus 59 passed on each of Python 3.12.10, 3.13.15 and 3.14.7. Log: .build/remediation-ann-factories-installed-matrix.log.
  • Nine profile-only installed-wheel source-to-serving processes: exited zero; three exact baseline repeats and three projection cases on each backend. These predate ANN factory edits and retain the separate profile-only identity, not the final candidate's identity.
  • Native-load probe's first fixture used category instead of category_id and failed before artifact creation. Corrected probe passed; six controlled fresh-process repeats exited zero. No retrieval oracle or production behavior was weakened.
  • New evidence summary exited zero, proving exact profile-only baseline parity, cross-backend projected-input equality, successful native-load checks and repeated concurrent Scope results. It preserves separate build identities and makes no unbounded qualification claim.
  • .build/remediation_ann_factories_identity.py: exited zero; original/mirror source, pyproject and lock match; new wheel/sdist digests recorded separately. git diff --check: passed.
  • Maintenance probe initially passed an unsupported clock keyword to create_service_store. Both initial environment attempts failed before schema creation. After inspecting the actual factory contract, immutable ServiceStore replacement supplied the clock over the same owned engine. Corrected first process exited zero. It removed 400 expired interactions through 58 batches (at most seven each), completed bounded rebuild checkpoints, reopened its store after ten batches, retained all 20 headers/200 live interactions and replayable live receipts, and dropped only its owned PostgreSQL schema. Two fresh repeats also exited zero with exactly 58 batches, at most seven removals per batch, complete resumable profiles and receipt replay. Maintenance time was 0.872–0.953 seconds and peak RSS 70.5–71.0 MiB. These small sequential single-maintainer observations do not establish production throughput or lock latency.
  • Three final-candidate installed-wheel source-to-serving baseline processes exited zero; complete serialized snapshots, rankings/provenance, Generation Manifests and stable metrics exactly match the preceding baseline. The final aggregate checker exited zero, verifying those signatures, all three maintenance repetitions and zero remaining owned concurrent/ maintenance PostgreSQL schemas. Log: .build/remediation-ann-factories-final-probes.log; aggregate: .build/remediation-final-probe-summary.json.
  • Final make docs-check: exited zero, 64 Markdown files checked. Final git diff --check: passed. The lock and generated SOURCES.txt remain unchanged; the IDE-owned mirror-venv exclusion remains as described above. All processes from this checkpoint reached terminal results; no qualification job is running in the background.

No migration or Compose artifact changed in this additional slice; their prior applicable checks remain recorded above. Full-profile source/control query plans, production-equivalent TCP HTTP latency, dense full-profile runtime, monthly availability and H1's chosen overlap semantics remain unqualified or undecided. The host/disk/time question is still pending; elapsed time grants no authority for expensive execution. The exact 730-day historical coverage also needs reconciliation with the calendar-year qualification window before claiming the two-year floor. The full goal is active, not complete.

Calendar window correction and SQLite query factory

This checkpoint supersedes the preceding candidate's 100-million-purchase/730-day qualification v2 shape and the unresolved calendar-window item. Historical v1 stays unchanged and selectable. Current v2 has 200,000 Catalog Items, 200 million views, 101 million purchases and 731 days. At the fixed January 1, 2026 cutoff, two calendar years begin on January 1, 2024 and include the leap day. The purchase margin permits some oldest rows to leave an advancing worker window; actual observed counts must still meet the service minima. No full-profile source was generated. The corrected unpublished preset retains v2; generator, schema and manifest versions are unchanged. Three source guides were reconciled with the actual versioned presets, including stale development quotas. Profile eligibility metadata still does not establish a Qualification Claim.

The advancing-clock TCP rehearsal then exposed an independent SQLite generated-source defect. Stored timestamps have a T separator, six fractional digits and Z. Default sqlite3 datetime bindings have a space, optional fractional digits and an offset. Lexical comparison selected same-day rows before the lower bound (causing source_invalid) and excluded valid upper-bound rows. Eight independently counted midday/microsecond cases reproduced both failure modes. A temporary parameter-only encoding intervention changed eight failures to zero; its engine listener was removed in finally. Diagnostic output contained boundary labels and error types.

The materializer now uses the named _create_sqlite_queries factory. It normalizes the two UTC parameter expressions into the stored six-digit format; it does not parse or transform row columns. MaterializedSource.queries carries the resulting engine-specific contract. Generic merchant adapters, PostgreSQL queries, stored rows, schema and logical digests are unchanged. The original adapter-compatible test retains all its previous assertions. Twelve observable cases now cover lower/upper midday, microsecond and exact endpoints for views and the combined online/offline purchase stream. An independent Python count verifies inclusive history_start and exclusive cutoff through the real adapter. Default sqlite3 datetime-adapter deprecation warnings remain visible.

Commands and intermediate candidates

  • Profile make test-focused TEST=tests/unit/test_synthetic_source.py: RED, two failed and 18 passed; corrected calendar/purchase preset GREEN, 20 passed with six warnings.
  • make test-focused TEST=tests/unit/test_synthetic_schedule.py: exited zero, 28 passed.
  • Calendar candidate pinned make verify: exited zero, 852 unit, 131 contract/delivery and 140 integration tests passed (six skipped, 33 PostgreSQL deselected), build and installed smoke. .build/remediation-calendar-pinned-verify.log retains this intermediate candidate.
  • Calendar candidate disposable make test-postgres: exited zero, 33 passed, 140 deselected.
  • Calendar installed-wheel matrix: exited zero; 895 plus 59 tests passed on each of Python 3.12.10, 3.13.15 and 3.14.7. It predates the timestamp fix and is not the final build evidence.
  • .build/remediation_calendar_identity.py: exited zero; source/lock parity and separate calendar wheel/sdist identities recorded in .build/remediation-calendar-package-identity.json.
  • Timestamp regression make test-focused on the parameterized case: RED, eight failed with 22 warnings. Initial temporary characterization attempts also exposed an insertion mistake: original-test assertions had been displaced into the new test. They were restored before the corrected parameter-only intervention exited zero (eight native failures, zero normalized). .build/remediation-sqlite-parameter-characterization-corrected.log and .build/remediation-sqlite-parameter-characterization.json retain that corrected evidence.
  • Timestamp factory full-file GREEN: 28 passed with 90 warnings; after four exact endpoint cases, make test-focused TEST=tests/unit/test_synthetic_source.py: 32 passed, 126 warnings. .build/remediation-window-exact-green.log is the latest focused result.
  • Timestamp candidate pinned make test-static and make verify: exited zero; 860 unit, 131 contract/delivery and 140 integration tests passed, build and installed smoke. This candidate precedes four exact endpoint cases and the return-docstring correction.
  • Verification correction: Ruff invoked inside the ignored .build mirror reports no Python files, so its apparent pass does not establish linting. Original-root make lint caught four missing Returns sections in the ANN/SQLite factories. After correction, original-root make test-static exited zero with actual Ruff checks, mypy (81 files), documentation (64 Markdown files) and architecture checks. Log: .build/remediation-window-root-static.log. This supersedes prior mirror-only lint claims; their other independently executed gates remain recorded. No lint exclusion or production rule was weakened.
  • Pinned make smoke after the return-docstring correction: exited zero, rebuilt wheel/sdist and installed smoke passed. That wheel was reinstalled with uv pip install --no-deps --reinstall-package commerce-recommendations in the owned Python 3.14 environment, exit zero.
  • Read-only status/diff/log/process inspections passed; a guessed characterization report and summary-helper filename were absent and corrected. One context-only patch was rejected and made no change. Original/mirror test copying succeeded. No lock/SOURCES regeneration occurred.

Repeated bounded TCP evidence

Three fresh rehearsals use actual TCP, separate API/worker processes, SQLite merchant source and unique owned PostgreSQL schemas. An injected epoch begins twelve hours after the fixed cutoff and advances with real monotonic time. The custom 1,000-Item/25,000-view/10,100-purchase, 731-day profile is explicitly ineligible for qualification. Existing scripts/qualify.py runs with reduced volume minima (1,000/24,900/10,000), 100 Serving requests, 40 idempotent Training submissions and concurrency four. Production qualifier defaults and oracles are unchanged.

The first exploratory attempt used an invalid placeholder anchor and failed; its exact HTTP failure was not retained. The corrected Oracle Manifest anchor attempt failed source_invalid after three source reads, leading to the timestamp diagnosis above. After the query factory fix, all three fresh processes passed. Each independently counted its actual worker window: 25,000 views, 8,078 online and 2,020 offline purchases (10,098 total). All eleven strategies published, each worker made five source SELECTs, and twenty additional TCP Serving reads after removing the merchant source retained the snapshot identity and declared Oracle result. Each process cleaned its children, source and owned PostgreSQL schema.

Local training times were 7.007–7.036 seconds; Serving p95 was 10.067–10.691 ms and p99 11.914–13.522 ms. Source Manifest identity and qualifier-script identity matched across all three. Reports are .build/remediation-tcp-window-{1,2,3}.json; success logs are the first -fixed-window.log and the second/third .log. Stale first-attempt .failed.json is separate and must not be interpreted as a fresh failure. These bounded observations use ANN disabled, reduced minima and a logical clock. They do not establish full-capacity or production-equivalent HTTP latency, monthly availability or a full Qualification Claim.

Final timestamp-factory release checkpoint

  • Final pinned make verify: exited zero; 864 unit tests (six skipped, 138 visible warnings), 131 contract/delivery and 140 integration tests passed (33 PostgreSQL deselected), followed by wheel/sdist build and installed smoke. Log: .build/remediation-window-final-verify.log. The separate original-root static gate supplies actual Ruff evidence; mirror Ruff still skips ignored files as documented above. The pinned telemetry checkout and lock remain unchanged.
  • Final disposable make test-postgres: exited zero, 33 passed, 140 deselected. Log: .build/remediation-window-final-postgres.log. No TCP schema-mutating probe was live during it.
  • Final installed-wheel matrix: exited zero. Python 3.12.10, 3.13.15 and 3.14.7 each passed 907 unit/contract tests (six skipped, 138 visible warnings) and 59 canonical-source/end-to-end/ lifecycle/migration tests. Installed smoke rejects source-tree imports on every runtime. Log: .build/remediation-window-final-installed-matrix.log.
  • .build/remediation_window_identity.py: exited zero; original/mirror production source, pyproject and lock identity match. The source/lock SHA3-512 is beff32c1f23a38419555ac9082388a1a45768ac501d662bb06ce4236eac9f99bd44a6a60b272fe09ab56126447d457d0ddb9f5235c3b8e4e274456f67fdd660f. The 1,142,941-byte wheel SHA3-512 is 0512a26782a1649495d05a01974e6bbee718c985c582bad0972cb449e2794f1aed9849c723e3176e1353fdd93ac4fc163892a309e1c9e766bfdecfb8494d956b. The 1,110,831-byte sdist SHA3-512 is d916d1bb1dd3800207afb451ce8b96c2275a65a233c19ea8eb4f6ba2b6126958fd9d1ef3bae75a7c52c3bd9b80883dae4976caa41a7516afc387169bc539220d. .build/remediation-window-package-identity.json preserves these separately from prior builds.
  • One fresh Python 3.12 installed-wheel DuckDB source-to-serving baseline exited zero. Its Generation Manifest, stable metrics, full serialized snapshot and rankings/provenance exactly match the earlier baseline. It overlapped the runtime matrix, so its timing is excluded from comparisons. .build/remediation_window_evidence_summary.py exited zero, asserting that parity, repeated TCP source/window checks, sixty source-independent Serving reads and zero remaining owned Q1, maintenance or calendar-window PostgreSQL schemas. Aggregate: .build/remediation-window-evidence-summary.json. TCP used the preceding smoke wheel with the same production source; final archives were rebuilt after test-only endpoint additions.
  • Final status/log/diff inspections and git diff --check passed. A guessed old probe helper was absent; no replacement run was started from that path. No migration or Compose artifact changed in this slice, so their prior applicable checks remain recorded above.
  • Final original-root make test-static: exited zero; actual Ruff checks, mypy (81 source files), documentation (64 Markdown files) and architecture passed with all twelve endpoint cases and the refreshed handoff. Log: .build/remediation-window-final-root-static.log. A telemetry status command using a pathspec outside this repository failed; status was subsequently inspected with git -C ../telemetry without modifying the sibling checkout.
  • All owned rehearsal, release and runtime processes reached terminal results. No qualification job is running. Final make docs-check and git diff --check passed. The lock and generated SOURCES.txt remain unchanged. The existing IDE-owned mirror-venv exclusion remains for the still-owned qualification layout and is not a configured SDK change.

The full goal remains active. The date-window defect is resolved and the bounded TCP rehearsal is green; full-profile source/control query plans, dense full-profile runtime, production-equivalent HTTP latency and monthly availability remain unqualified. The following decisions still govern the next dependent work. H1 overlap semantics and the expensive-run host/disk/time decision are still pending; no dependent semantic change or expensive execution has been started. No package publication, deployment, hosted change, merchant mutation or external message was authorized or performed.

Native PostgreSQL TCP and query-plan preparation

The preceding goal turn made concrete progress: the calendar/SQLite correction, endpoint tests, three TCP rehearsals and final release/runtime gates changed code and verified behavior. This continuation adds previously missing bounded native PostgreSQL source-to-TCP evidence; it is not a verified wait or a completion claim. Production source and the release candidate are unchanged.

The owned PostgreSQL 17.11 container on loopback port 51790 was revalidated before use. Each process created unique remediation_pg_source_* and remediation_pg_window_* schemas. Source connections have a two-slot pool and a schema-specific search path; control connections translate only service metadata into their own schema. The real PostgreSQL materializer used two COPY workers and 1,000-row batches. Five source tables were analyzed inside their owned schema. No public source table or global database statistics operation was used. The container has a 2 GiB memory cap, 128 MiB shared buffers and 4 MiB work_mem; it has no configured CPU quota. Worker and native library thread settings were explicitly two. It is a local rehearsal environment, not the approved full-capacity environment.

The custom 1,000-Item/25,000-view/10,100-purchase, 731-day profile, advancing historical clock, reduced qualifier minima and unchanged planted Oracle match the preceding SQLite rehearsal. One native pilot passed and captured source plans. Three additional sequential fresh processes also captured control plans and source relation sizes. Every process verified its receipt, observed exactly 25,000 views and 10,098 combined purchases inside the actual worker window, published all eleven strategies and made five worker source SELECTs. The Generation Manifest and counts match the SQLite realization. After the worker exited, the owned merchant schema was dropped. Twenty further real TCP Serving reads per process still returned the same snapshot and declared Oracle result. API/worker processes and both schemas were cleaned in finally.

Actual query plans

EXPLAIN (ANALYZE, BUFFERS, FORMAT JSON) ran the five canonical source queries against their actual bound worker window. Only operator/relation/index names, row/loop/filter counts, sort statistics, buffer counts and timing survive in reports. SQL, filters, conditions, parameters, interaction rows and their identifiers are excluded from retained plans.

After source removal, instrumentation captured the actual SQL of the ordinary published-set and global-candidate repository reads in memory. Each operation issued exactly one SELECT and resolved the same snapshot. Its listener was removed in finally before explaining the read; no SQL/parameter capture was written to an artifact. Ordinary serving returned one joined row; the global union with its Best Sellers fallback returned three lanes. Both plans used the recommendation_set_pkey index. The one-row serving-head/snapshot tables used sequential scans. At this source size the five source queries used sequential scans and in-memory sorting; all source/control plans had zero temporary blocks. These observations do not prescribe the planner choice or query latency at full scale.

Observation Three enhanced repeats
Training time 6.897–7.113 s
TCP Serving p95 / p99 9.787–11.008 / 11.145–12.062 ms
Source views query execution 12.462–20.951 ms
Ordinary-set / global-union query execution 0.029–0.040 / 0.034–0.051 ms
Five source relations including indexes 11,436,032–12,042,240 bytes

These are warmed local observations: no cache flush, ANN disabled, shared source/control server, custom reduced-volume profile and injected advancing clock. Relation sizes omit the generator's state table, WAL, temporary files and control database; they are not whole-process or disk bounds. The pilot is excluded from these three-repeat ranges. No production-equivalent SLO, dense full-profile runtime, capacity or availability claim follows from them.

Commands, identities and remaining work

  • Read-only git status, git diff --check, contract/handoff/source/log reads, process inspection and Docker inspection succeeded. No release or qualification process was live at the start.
  • Copied the preceding ignored TCP helper and added named factories for the native source engine and bounded source/control plan reports. The first native pilot invocation exited zero.
  • Three enhanced .build/remediation_tcp_postgres_window_probe.py invocations, using the final Python 3.14 installed wheel with two native threads, each exited zero. Source: the owned disposable PostgreSQL URL; qualifier: the pinned layout's unchanged scripts/qualify.py. Success reports/logs: .build/remediation-tcp-postgres-window-{2,3,4}.{json,log}.
  • .build/remediation_postgres_window_summary.py: exited zero. It compares every shipped Python file's bytes between the installed wheel, original source and pinned mirror, confirms the current source/pyproject/lock digest against the preceding release identity, validates all source/control plan counts and matching SQLite manifest/counts, and independently verifies zero remaining owned native source/control schemas. Report: .build/remediation-postgres-window-summary.json. The candidate source/lock and wheel identities remain those of the final timestamp-factory checkpoint above. Environment: Python 3.14.7, SQLAlchemy 2.0.54, psycopg 3.3.6, PostgreSQL 17.11, Darwin arm64.
  • Python environment preflight preceded each invocation. Only documentation and ignored local probes changed, so release, runtime, PostgreSQL integration, migration and Compose gates were not repeated; their preceding terminal evidence remains applicable. Final make docs-check exited zero (64 Markdown files), and git diff --check passed. All four native rehearsal processes and the aggregate checker are terminal; no qualification job is running. The lock and generated SOURCES.txt remain unchanged.

Q1 now has bounded native PostgreSQL TCP and source/control plan evidence. Its full-profile query plans, dense runtime, production-equivalent latency and sustained monthly availability remain unqualified. H1's overlap choice and the full run's declared host/disk/time budget and disposable qualification environment remain pending. No dependent semantic strengthening, expensive run, package publication, deployment or hosted change was started. The full goal remains active.

Current acceptance audit and unresolved execution decisions

The preceding continuation made progress by adding native PostgreSQL TCP and query-plan evidence. This audit revalidated the current worktree and terminal artifacts rather than starting another small rehearsal. No active rehearsal, release, runtime or qualification process was found; this is not a verified wait.

  • An AST inventory maps all 34 finding sections to the current tests. All 31 named remediation regression functions exist, as does H1's separate post-acknowledgement baseline. Q2 has its dedicated architecture-validator cases; Q1 is an execution-evidence requirement, not a named unit test. Artifact: .build/remediation-acceptance-inventory.json.
  • Every current test file matches the tested pinned mirror byte for byte. The installed Python 3.14 JUnit file confirms the named unit/contract regressions passed. Integration cases are not part of that XML; their complete files are in the final integrated gate, which selected and passed all 140 non-PostgreSQL integration cases. The separate disposable gate passed all 33 PostgreSQL cases. Test existence is not treated as semantic or capacity proof.
  • The actual lease-generation, resumable-maintenance, native metric rejection, dense ANN selection and background-load tests were inspected against their design oracles. They check successor survival, bounded resumed state, rejection of L2 artifacts, independent exact rankings/allocation bounds and old-generation completion, respectively. H1's current test commits the barrier before beginning the read, verifies suppression and covers lifecycle continuation; it does not promise cancellation of a read that began before the barrier.
  • The six skipped unit cases all require optional MLX/ANE accelerator dependencies. They are not skipped remediation regressions and do not establish accelerator qualification. They cannot be counted toward Q1 or H1. The full-profile targets remain unproven despite the passing default backends, static gates and supported-runtime matrix.
  • Read-only git, contract/source/artifact reads, process inspection and the AST/XML audit command exited zero. Python environment preflight preceded the audit. No source, test, migration, configuration or package identity changed, so already-passing executable gates were not repeated. Final make docs-check exited zero (64 Markdown files), and git diff --check passed. The lock and generated SOURCES.txt remain unchanged.

The full objective is not achieved. The same outstanding H1/full-run decisions are recorded in the last three goal checkpoints, and the independently authorized code and bounded qualification work is exhausted. Further small repeats would not satisfy the actual remaining requirements. There is no live full run to poll and no basis for fabricating capacity or monthly-availability evidence.

The required next inputs remain:

  1. H1: choose the rule for a profile read overlapping an acknowledged opt-out/deletion. The design disposition says: “Only implement semantic strengthening after the maintainer confirms the overlap rule.” Existing post-acknowledgement suppression is already required and verified. No strengthened overlap rule has been authorized.
  2. Q1: resolve the pending full-profile host/disk/time budget and disposable qualification environment. The operational boundary says: “Expensive full-profile runs require declared host/disk/time budgets and disposable database authority.” The reviewable proposal already recorded above is this 32 GiB host, at most twelve hours total and 256 GiB disk, with separate owned PostgreSQL qualification state. The existing 2 GiB local test container and reduced native pilot are not substitutes.

These are design-reserved decisions, not an automatic approval rejection or a new skill-imposed confirmation. Prior unanswered questions remain unanswered; automatic continuation, elapsed time and a style preference do not select either branch. The full objective remains intact and incomplete. Resume dependent work when those inputs arrive.

Maintainer decisions and final reduced execution

This checkpoint supersedes the preceding blocked audit. The maintainer replied “1. ok” to keeping existing opt-out overlap semantics and “2. reduce it” to the proposed twelve-hour, 256-GiB full-profile run. H1 is resolved; full-profile execution is replaced for this handoff by a custom 2,000-Item/250,000-view/50,100-purchase profile over 731 days. Neither the named qualification preset nor production requirements were reduced. All previously implemented remediation slices remain in scope and retain their final release evidence.

Accepted profile-read boundary

test_postgresql_overlapping_profile_read_may_observe_prior_lifecycle covers opt-out, deletion request and reauthorization with real separate connections. A cursor barrier holds an already-started profile read while the lifecycle transition commits and acknowledges. That read may finish with its prior version/lifecycle; a subsequent read observes the acknowledged version and lifecycle. Existing tests independently verify that application reads begun after opt-out/deletion acknowledgement expose no usable profile. Production read semantics were not strengthened. The personalization document and H1 disposition now describe the accepted boundary explicitly.

Reduced measurement and failed observation

The bounded run uses the final installed Python 3.14 wheel, native PostgreSQL merchant and control state in unique owned schemas, separate API/worker processes, and real loopback TCP. It retains the default generation seed, an advancing historical clock, ordinary parameter grids, all eleven strategies and the existing qualifier. Custom volume minima are explicit: 2,000 Items, 249,000 views and 50,000 combined purchases in the actual worker window.

The stop budgets are ten minutes wall time and an eight-GiB sampled storage threshold, with two worker/native threads, 512 MB pipeline memory and 1 GB derived temporary storage. PostgreSQL retains its existing 2-GiB container cap. The one-second storage sampler sums owned source/control relations, incremental WAL and owned scratch files; it can miss short peaks and is not a whole-host disk ceiling. Pipeline memory limits do not cap native-library or total process residency. No stop threshold fired and no monitor failed.

Final process observation Measured result
Total wall / training duration 62.240 / 38.666 s
Worker CPU / peak RSS 54.876 s / 1,829,748,736 bytes
Sampled storage peak 180,912,128 bytes over 62 samples
Actual views / online / offline purchases 249,886 / 40,046 / 10,013
TCP Serving requests / concurrency 100 / 4
TCP Serving p95 / p99 10.562 / 12.134 ms
Idempotent submissions / submission p95 / status p95 40 / 14.122 ms / 9.377 ms
Published strategies / worker source SELECTs 11 / 5
Identical serving reads after merchant schema removal 20

The reduced qualifier's execution and latency checks passed. Its aggregate result does not check the generated Oracle's planted pair. The additional Also Viewed pair-inclusion observation failed at both 20 and 100 returned Items. Independent scoped source SQL counted only four distinct supporting contexts for that pair; training selected minimum support five and shrinkage 20. Both work-store backends exclude pairs below their selected support threshold. The source count is an upper bound on reduced qualifying support, so exclusion is consistent with that ranking contract. The earlier hypothesis that only the response limit caused the failure was disproved.

This is a failed custom-profile Oracle observation, not evidence of a service ranking bug or a successful correctness/qualification run. The manifest's unconditional pair expectation does not account for the selected support threshold. No seed, history, ranking threshold, frozen recipe or expected inclusion was changed to obtain a pass. The final helper writes validation_passed: false and exits one after preserving aggregate evidence and cleaning its processes/schemas. General generator support for arbitrary custom-profile Oracle claims is a separate limitation exposed by this experiment; it has not been silently repaired by changing historical output. The synthetic-generation documentation records that limitation.

The final source queries use canonical-order index scans for views/online purchases and in-memory sorting for Catalog, compatibility and offline purchases. Ordinary/global control reads each issue one SELECT and use the recommendation-set primary-key index. All plans have zero temporary blocks. Plans retain only whitelisted operator/count/buffer/timing fields; SQL, parameters and raw interaction identifiers are not retained. Snapshot identity and the hash of the exact ranked Items remain unchanged across merchant-schema removal.

Commands and final evidence

  • Read-only git status, git diff, git diff --stat, rg, sed, cat, tail, environment, artifact and contract inspections succeeded. Python environment preflight preceded each toolchain invocation. No production source, migration, lock or package identity changed in this checkpoint; only the overlap test, governing documentation and ignored QA helpers changed. Earlier unrelated sibling telemetry/lock/generated-file state remains preserved.
  • TEST_CONTROL_DATABASE_URL=<owned disposable URL> make test-focused TEST=tests/integration/test_postgresql_storage.py::test_postgresql_overlapping_profile_read_may_observe_prior_lifecycle: exit zero, three passed in 1.42 s; log .build/remediation-h1-accepted-overlap.log.
  • Original-root make test-static: first failed two new test line lengths, corrected by wrapping them. The second invocation exited zero: actual Ruff, mypy for 81 source files, 64 Markdown documents and architecture rules. Log .build/remediation-reduced-final-static.log.
  • Three fresh .build/remediation_reduced_benchmark.py invocations with the owned disposable control URL and pinned-layout scripts/qualify.py: all exited one on the declared planted inclusion. Invocation 1 checked 20 results; invocation 2 disproved the response-limit hypothesis with 100 results; invocation 3 also captured source support, selected parameters, CPU/RSS, storage, source-removal signatures and the complete failed aggregate report. Logs .build/remediation-reduced-validation-{1,2,3}.log; final aggregate report .build/remediation-reduced-validation-3.json; diagnostics for runs 2/3 are labeled .oracle-diagnostics.json. Temporary merchant rows were removed in finally every time.
  • TEST_CONTROL_DATABASE_URL=<owned disposable URL> make test-postgres: exit zero, 36 passed/140 deselected in 36.69 s; log .build/remediation-h1-final-postgres.log.
  • Pinned-layout make test: exit zero; 864 unit passed/six existing optional-accelerator skips/138 SQLite adapter warnings, 131 contract/delivery passed, 140 non-PostgreSQL integration passed/36 deselected. Mypy, docs and architecture passed. Its Ruff invocation again found no files inside ignored .build; only the actual original-root Ruff run above is lint evidence. Log .build/remediation-h1-final-test.log. Together with PostgreSQL, 1,171 selected cases passed; the failed custom Oracle experiment is reported separately.
  • .build/remediation_reduced_summary.py using the final installed wheel: exit zero. It verifies the failed aggregate report, independently counts zero remaining owned source/control schemas, compares all shipped Python bytes with the installed wheel and pinned mirror, compares all current test bytes with the tested mirror, and verifies the unchanged final source/pyproject/lock digest. Artifacts .build/remediation-reduced-validation-summary.{json,log} bind the final report/helper identities and preserve validation_passed: false.

The source/lock and wheel identities remain those of the final timestamp-factory checkpoint. The prior installed Python 3.12/3.13/3.14 matrix and build/smoke gates still cover the unchanged production candidate. They were not repeated for this test/documentation-only checkpoint. Migration and Compose checks were not repeated because those surfaces did not change.

Full-profile query plans, dense runtime, production-equivalent SLO/capacity and monthly availability are unqualified; the maintainer/operator own that subsequent operational work. The new custom-profile Oracle limitation also remains explicit. There is no outstanding H1 decision, no full run pending authorization for this revised scope, and no deployment, publication or external data mutation.

  • Final original-root make docs-check: exit zero, 64 Markdown files; log .build/remediation-reduced-final-docs.log. git diff --check passed.
  • Independent cleanup audit found zero remaining owned source/control schemas. docker stop recommendations-remediation-postgres exited zero; subsequent inspection reports no remaining task container (exit one, object absent).
  • An initial cleanup command was rejected before execution because automatic review disallows rm -f-style deletion. The safer reversible cleanup used mktemp and mv, both exit zero, to archive the pinned test environment at /tmp/recommendations-remediation-venv.yCP9iS/pinned-layout-venv. No environment was deleted. Its automatically added IDE exclusion was removed; the original project SDK is unchanged. Reports, source mirror and installed-wheel environments remain available for review. Move the archived environment back to its recorded mirror .venv path before rerunning mirror gates.

The implementation and maintainer-reduced execution are ready for review. The custom experiment remains failed, and the excluded full-capacity/availability work is explicitly unqualified. No additional local repeat can turn those facts into a qualification claim.