Skip to content

Repository review evidence

Status: baseline review evidence; remediation implementation in progress. Baseline: 1dfb34a, 2026-09-30.

The user changed the goal to detailed fixes for every issue found. The linked remediation design is the current deliverable. The coverage limits below remain evidence limits; further repository review is not required to complete this planning task. The subsequent implementation request and current red/green evidence are tracked in the remediation handoff. Findings below describe the review baseline.

The preceding objective was repository-wide correctness, scalability, performance, and architecture review. These passes do not certify every line or production capacity. At that review baseline, production source and committed tests were unchanged. Findings distinguish executable counterexamples from static analysis and from unresolved questions.

Method and acceptance evidence

Review contracts before implementations; trace source, reduction, publication, and serving boundaries; examine tests for counterexamples they omit. Measure scaling over multiple input sizes, preserve inputs and commands, and distinguish asymptotic reasoning from local timing. A passing suite establishes covered behavior only.

Patterns are evaluated against concrete pressures: variation, isolation, consistency, bounded state, and resource lifecycle. Pattern count is not a measure of architectural quality. Prefer a narrow abstraction with an invariant over additional layers without an identified responsibility.

Completion of that review would have required an inventory covering production packages, migrations, scripts, delivery/configuration, and tests; reviewed contracts and failure paths; reproduced high-impact findings; workload-based performance evidence; and explicit remaining risks. That evidence is incomplete today.

Spec and correctness findings

C1 — P1: For You mixes two published snapshots

src/recommendations/serving.py:200 reads global candidates and then independently reads the current Best Sellers head. Each repository operation can be correct while a publication between calls returns different snapshots. The resulting CandidateBundle uses the first snapshot's metadata and the second snapshot's fallback. Snapshot-local serving is consequently violated without any cross-scope query.

Executable probe .build/review_for_you_race.py uses the actual loader and domain values with a deterministic reader advancing from snapshot A to B. Output:

reported_snapshot=snapshot-A
candidate=A-only fallback=B-only
Confirmed: fallback from snapshot B is packaged with snapshot A metadata.

Recommendation: resolve one immutable snapshot identity, then read both lanes by that identity, preferably through a cohesive repository operation. A snapshot-bound read view is a useful pattern here. A race regression must exercise a publication between the reads and verify every entry and disclosed identity belong to one snapshot.

C2 — P1: Valid feature-poor Catalogs can abort Similar Items generation

src/recommendations/pipeline/content.py:167 considers zero price and whitespace text metadata. _price_buckets omits nonpositive prices; document processing omits whitespace. With two such products, _metadata_matrix calls hstack([]) at line 122. Punctuation text reaches the word vectorizer at line 113, which raises empty-vocabulary ValueError, even when structured category features are usable.

Executable reviewer probes reproduced IndexError for two zero-price rows and two whitespace-title rows; punctuation-only titles reproduced ValueError. Treat absent usable feature blocks as insufficient evidence and qualify word/character blocks independently. Tests must distinguish empty metadata, unusable text, and structured metadata with unusable text; ordinary usable features must retain their scores.

C3 — P2: Candidate truncation precedes eligibility filtering

src/recommendations/pipeline/workstore.py:516 ranks and limits behavioral candidates before final eligibility checks. pipeline/run.py:521 supplies no eligibility mapping for unconstrained behavioral lanes. pipeline/fallbacks.py later removes ineligible products, but cannot recover qualifying candidates discarded by the earlier cutoff.

Reviewer probe: anchor A, 200 higher-scoring ineligible neighbors, and eligible C with qualifying lower support; generation limit 200 drops C, yielding no primary recommendations after eligibility filtering. Global popularity has the same ordering risk: _score_candidates at pipeline/run.py:1595 truncates its input mapping before the downstream eligibility filter.

Recommendation: apply publication eligibility before candidate top-N, across DuckDB, Polars, popularity, and applicable category/geographic paths. Verify against an independent reference with more excluded candidates than the generation limit.

C4 — P2: Missing partition identity can bypass applicable limits

src/serving_limits/core.py:274 collapses a matching rule with missing partition descriptors into the same empty result as genuinely unmatched traffic. With allow_unmatched=True, AdmissionController admits it without enforcing any matched rule. This option should exempt unmatched traffic, not malformed matching traffic.

The admission probe prints matching_rule_missing_partition_allowed: true for a globally matching consumer-partitioned rule and a request with only a service descriptor. The existing missing-partition test covers only the default closed policy. Represent malformed resolution separately from no selector match and test both policy settings, including a parent rule alongside a malformed child.

C5 — P2: Similar Items discards lexical tie winners before sorting

The native sparse top-N at pipeline/content.py:50 selects tied candidates before Python applies Product ID ordering. Probe: 210 identical category vectors, top_n=200, anchor 000; the first ten returned IDs are 009 through 018, rather than 001 through 010. Ascending Product ID tie ordering is explicit in specs/commerce-recommendation-service.md:97. Sorting survivors cannot recover discarded winners. Make secondary ordering part of bounded selection and test more tied candidates than the cutoff, including input permutations and native thread counts.

Standards and scalability findings

The next confirmed correctness findings are recorded separately below; finding IDs remain stable as coverage expands.

S6 — P2: Request bodies have no service-enforced byte bound

The POST routes in src/recommendations/api.py register ordinary FastAPI body models without an ASGI byte limiter. Field and tuple bounds take effect after body buffering and parsing. A chunked Training Run request containing the valid 80-byte JSON scope plus whitespace was accepted with either 1 MiB or 8 MiB of padding. The service consumed all bytes and returned 202. Traced Python peak allocations were 3,440,265 bytes and 16,933,459 bytes respectively. Probe: .build/review-body-bound.py.

These are bounded local memory observations, not RSS or throughput qualification. Source inspection establishes the lack of a byte-limit branch. The repository's AGENTS invariant explicitly requires bounded request bodies; an externally configured proxy limit is not enforced by this app. Enforce both declared and streamed body bytes before JSON parsing, including chunked requests and absent or incorrect Content-Length. Preserve authentication-before-body behavior on protected routes and use a stable 413 response. Test early termination of receive, rather than merely eventual rejection.

S7 — P2: Exact ANN tie selection allocates every tied match

src/recommendations/ann.py:416-423 materializes and sorts every match at the selection threshold before retaining at most 100. The real _exact path, probed with equal unit vectors, instantiated 10,000 AnnMatch objects for a 10,000-Item Catalog and 100,000 for a 100,000-Item Catalog; both returned 100 entries. The larger probe used about 22.7 MB of traced temporary allocations and 0.28 seconds, versus about 2.3 MB at 10,000 Items.

The full score array is documented; the additional Theta(N) Python match residency and worst-case Theta(N log N) tie sorting are avoidable. Preserve exact score/Product-ID ordering with bounded selection, and exercise underfilled native retrieval as well as small-index exact search. These diagnostic measurements are not production latency claims. Probe: .build/review-ann-probe.py.

Further spec and correctness findings

C6 — P1: A stale publisher destroys its successor's staged snapshot

src/recommendations/storage.py:1093-1121 deletes any building snapshot for a run before checking lease ownership at lines 1241-1251. A recovered worker starts publication; the old worker resumes, replaces the new owner's staging, and only later fails its ownership check. The legitimate publisher then fails because its own staged parent has disappeared. The last serving head can survive while forward publication fails.

.build/review_stale_publication.py uses actual repositories with SQLite foreign keys enabled, an old lease recovered by a new owner, and a deterministic generator interleaving. Output:

Old owner rejected for lost lease after replacing new owner staging.
Current owner publication failed: IntegrityError
Published snapshots: 0; successor staging was destroyed by stale owner.

Fence ownership before destructive staging operations and make batches/cleanup belong to a fenced generation. The check and mutation must be atomic with lease recovery; an isolated preliminary check alone is insufficient. Existing stale-publisher tests protect an available head but omit overlapping building publishers. Verify this race on PostgreSQL as well; this reproduction is not PostgreSQL locking qualification.

C7 — P1: Default SQLite source transactions are not repeatable reads

src/recommendations/source.py:350-354 calls SQLAlchemy begin with SERIALIZABLE, but worker.py:631 constructs default sqlite3 engines. Under the pinned Python driver, legacy transaction control does not issue BEGIN for SELECT. Fully consuming one stream and opening another can observe a new committed source state without changing the consistency token. This defeats the consistent source-read contract even when every query is correctly scoped and cutoff-filtered.

Probe .build/review-source-consistency.py uses a disposable WAL-mode database and separate reader/writer connections. The default reader returns A, a writer commits B, then the same read returns A and B; its token check passes. A control that explicitly emits BEGIN returns only A on the second read. Both assertions pass.

SQLAlchemy's SQLite transaction documentation explains this driver behavior. This repository's technical documentation acknowledges that token identity alone does not establish snapshot stability. Configure genuine driver transaction semantics and add cross-query concurrent-writer tests to the source contract suite. The existing seven source-adapter tests pass without testing it.

C8 — P1: Scoped personalization retention configuration is ignored

config.py:1305 accepts retention_days per capability, but storage.py:2102-2106 always builds the state repository with its default 90-day policy. Production API wiring at api.py:1521,1547 never forwards the configured scope policies.

The actual app/store probe .build/review_personalization_retention.py configures 30 days, commits a synthetic interaction, and observes a 90-day ledger expiry. At day 31, maintenance removes zero rows and the profile remains. Parsing-only tests do not establish enforcement. Resolve retention by Commerce Scope inside persistence; one global replacement cannot represent multiple scope-specific policies. Verify distinct policies on two isolated scopes, idempotency retention, and shortened-policy rollouts against the advertised maximum-history contract.

C9 — P2: Already expired interactions immediately affect personalization

personalization/health.py:37-47 rejects future skew but permits arbitrarily old interactions. state_store.py:733-744 then adds their affinities despite the expired retention boundary. A 91-day-old viewed interaction receives an acknowledgement and an item affinity of 1.0 under the default 90-day policy. This contradicts the maximum history behavior documented in docs/shopper-personalization.md.

Choose the explicit late-event contract: reject expired interactions or acknowledge their ordering without deriving expired features. Do not rely on a later maintenance pass to establish serving-time retention. Probe: .build/review_expired_interaction.py.

C10 — P1: Expiry rolls acknowledged watermarks backward

personalization/state_store.py:687-700 reconstructs sequence and profile version from the last retained interaction, conflating acknowledgement metadata with disposable history. The interaction contract requires exactly the next sequence and a valid causal token to observe at least the acknowledged version.

Probe: sequence 1 is recent, sequence 2 occurred 91 days earlier, and both are acknowledged. Expiry removes sequence 2 and changes the profile from sequence/version 2/2 to 1/1. The client's valid next sequence 3 is rejected with last_sequence=1. Full expiry can erase the watermark altogether. Existing tests cover increasing occurrence time, missing this reverse-order case. Maintain monotonic sequence/version separately from the retained feature ledger; test late events, full expiry, causal reads and retries. Probe: .build/review_expired_interaction.py.

C11 — P1: Completed ANN loads can block all later generations

ann.py:508-516 only reaps a completed Future when its own snapshot is requested again. If a new head replaces it first, the old pending entry remains; requests for new generations or other scopes always fall back without scheduling another load. Production config enables background_load at api.py:1527-1528.

The deterministic probe explicitly awaits old-head load completion, then requests the new head three times. Only the old snapshot reaches the artifact reader, and the old completed entry remains pending. Reap completed work independently of the requested identity using the scheduler's existing lock. Test completed success, missing artifact, failure, head replacement and another scope. Probe: .build/review-ann-probe.py.

C12 — P2: Swimlane recent-purchase exclusions stop after ten Items

swimlane_service.py:177-178 takes ten purchases as retrieval seeds and reuses them as the exclusion set at line 217. A profile can retain 100 purchases. With eleven recent purchases and an explicit exclude_recent_purchases Best Sellers step, the eleventh purchase is returned. The exclusion behavior in specification lines 535-536 does not define a ten-purchase subset. Separate the full bounded exclusion history from the smaller retrieval seed budget. Probe: .build/review-swimlane-purchase-probe.py.

C13 — P2: ANN artifact validation accepts the wrong distance metric

ann.py:305-317 validates native type and dimensions but omits metric_type. A correctly checksummed HNSWFlat index with METRIC_L2 loads successfully, while serving interprets native distances as inner-product similarity and exact fallback computes inner product. The ANN lifecycle spec explicitly requires metric binding and incompatible-metric tests. This is trusted-artifact compatibility validation, not a claim about malicious artifact authenticity. Reject unsupported metrics at loading and qualify native/exact score parity. Probe: .build/review-ann-metric-probe.py.

C14 — P2: Unvalidated run IDs turn client errors into server errors

api.py:490 constructs TelemetryContext with the unvalidated path run_id. Its 255-character identifier guard raises before the repository's missing-run response. A 256-character run ID, or a credential-shaped ID, produces HTTP 500, even with disabled observability; an ordinary unknown ID produces 404. Validate transport input before constructing telemetry values and return the chosen 4xx contract without echoing rejected material. Probe: .build/review-api-run-id.py; actual TestClient, server exceptions disabled.

C15 — P2: Identical config fingerprints can select different parameters

config.py:1026-1035 sorts search grids in the configuration fingerprint, but no-holdout selection at pipeline/run.py:1050-1051 uses the first grid value when the grid excludes the default. Trending selection similarly uses candidates[0] at lines 1114-1119, and equidistant tie selection can retain iteration order. Configured view half-life grids [2,8] and [8,2] therefore share config_version but choose 2 and 8 days respectively.

This invalidates configuration identity as comparable-run/reproducibility evidence. Either canonicalize runtime choice to match hashing or preserve semantically meaningful order in the fingerprint. Test no-holdout and tied-holdout selection for every grid. Probe: .build/review_config_grid_identity.py.

C16 — P3: Nonfinite polling configuration crashes idle workers

config.py:1001 checks only worker_poll_seconds <= 0. NaN passes that check and reaches the worker's sleep, which raises ValueError. The configured-Python probe .build/review_poll_interval.py reproduced validation acceptance and sleep failure. Validate finite positive intervals before constructing the runtime; cover NaN, both infinities, zero and negative values in configuration tests.

E1 — P2: Duplicate observations bypass minimum-count suppression

personalization_evaluation.py:275-284 accepts repeated ranking observations. Repeating one Shopper observation ten times produces an available threshold-10 report with count 10 and Recall 1.0. Count independent evaluation units and reject duplicate observations. This is an offline evidence-contract issue, not a demonstrated API leak.

E2 — P2: Outcome attribution accepts an unexposed variant

personalization_evaluation.py:613-635 omits variant from exposure correlation and does not compare the stored assignment. A holdout exposure permits personalized-variant attribution. Bind and verify the complete assignment identity.

E3 — P2: Delimiter-joined HMAC inputs have ambiguous scope identity

personalization_evaluation.py:498-507,626-634 joins tuple fields with colon without escaping or length prefixes. Property a:b / experiment c and property a / experiment b:c produce identical exposure correlation material, permitting cross-property attribution in the offline helper. Use unambiguous versioned canonical encoding for assignment and exposure, without compatibility fallback to ambiguous digests.

E4 — P2: Nonfinite latency silently passes experiment guardrails

personalization_evaluation.py:658-659,685-693 permits NaN latency and nonfinite latency thresholds. The guardrail function returns no breach for NaN p95_latency_ms. Require finite measurements and thresholds; missing evidence must not become rollout approval.

E5 — P3: Suppressed metric summaries can retain a hidden value

personalization_evaluation.py:205-206 accepts a suppressed summary containing recall=0.5 and ndcg=None. Require both metrics absent for suppressed/not-applicable status and both finite bounded metrics for available status. Normal summary construction is currently safe; the value-object boundary is not.

All five E findings were reproduced in .build/review-evaluation-probe.py. Searches found no external production callers for these offline helpers. Their severity refers to evidence integrity and prospective adoption, not a current serving exploit.

M1 — Model concern requiring further qualification

two_tower_torch.py:134-145 returns duplicate directed positive pairs when purchase and view evidence contain the same relationship, despite the helper's unique-pair contract. The probe returns ((0,1),(0,1)), consuming two places in the capped training pair budget. Duplicate targets can also become in-batch negatives under diagonal cross-entropy. Missing deduplication is reproduced; relevance impact and intended cross-signal weighting need a controlled model experiment before a stronger claim.

An early combined ANN/torch probe aborted with a duplicate OpenMP runtime error. Clean import and actual production-builder subprocesses succeeded. That abort is excluded from repository findings; it is a probe interaction, not an established production failure.

Additional standards and scalability findings

S1 — P2: Rate-bucket residency grows with lifetime cardinality

src/serving_limits/core.py:451 retains token buckets for every observed partition. No idle expiry or maximum bucket count exists; lease expiry does not remove rate buckets. Descriptor length bounds constrain one key, not the number of keys over time.

Probe: admit 10,000 distinct synthetic consumers, advance time to 1,000,000 seconds, trigger maintenance through acquire. All 10,000 token buckets remain. This is O(U) resident state in lifetime partition cardinality U, including fully replenished buckets. Expire buckets only when replenishment makes removal semantically equivalent to a fresh bucket; apply bounded, incremental maintenance and an explicit cardinality policy. Blind LRU removal of depleted buckets would incorrectly restore capacity.

S2 — P2: Every admission scans all live leases

src/serving_limits/core.py:425 invokes _expire_leases on each admission; line 550 reconstructs the entire active lease dictionary. Release also scans it. For N successful admissions retained within one TTL, total scanning is O(N²), and each operation runs synchronously under a process lock on the calling async path. Denial retry calculation adds another global scan and sorting.

Probe: one concurrency rule, capacity N+1, frozen clock 0, TTL 30 seconds, N admissions without release, three fresh-state repetitions per N. Measured seconds:

N Repetitions Median
1,000 0.06017, 0.06050, 0.06183 0.06050
2,000 0.22672, 0.22479, 0.22337 0.22479
4,000 0.88248, 0.88418, 0.89026 0.88418

Doubling N increased time by approximately 3.72× and 3.93×. This local microbenchmark supports the code-level quadratic mechanism, not a production throughput or latency claim. Use an expiry priority queue and bucket-local permit indexes while preserving atomic multi-rule acquisition, partial permit expiry, and idempotent release.

S3 — P2: Personalization expiry uses an unbounded transaction

Static evidence: src/recommendations/personalization/state_store.py:346 materializes all expired Shopper keys, deletes all expired rows, and rebuilds every affected profile inside one transaction. _rebuild_profile at line 670 additionally materializes the entire retained history for each Shopper. Time retention does not bound interaction count. Large backlogs increase memory, transaction duration, locks, and database work.

Recommendation: a bounded, resumable maintenance operation with limited keys/rows per transaction and streamed reduction. Qualification must cover concurrent ingestion, interruption/retry, preservation of profile versions, and eventual backlog clearance. This finding has not yet been reproduced against a large database.

S4 — P2: Holdout evaluation materializes the entire pair graph

Static evidence: pipeline/workstore.py:633 fetches all holdout pairs and builds two set memberships per edge. pipeline/run.py:984 loads that graph during search and line 420 loads it again for retained evidence. Polars repeats materialization at polars_workstore.py:404. For E distinct pairs, Python memory is Theta(E), with a worst-case quadratic graph in Catalog size, outside DuckDB spill limits. Bounding each interaction group does not bound the accumulated graph. Stream evaluation or introduce an explicitly qualified evaluation sample; preserve exact metric meanings or disclose sampling changes. Large-scale memory reproduction remains outstanding.

S5 — P2: Complementary-category fallback repeats equivalent full sorts

Static evidence: pipeline/run.py:1548 sorts every Item in paired categories separately for every anchor, builds full candidate lists, then truncates to 200. For two equally sized categories, sorting work is Theta(N² log N) although anchors in a category share the same candidate ordering. Precompute one bounded list per category and reuse the existing KeyedCandidateProvider pattern. Benchmark with growing category sizes and verify exact scores, ordering, eligibility and fallback provenance.

S8, additional scalability candidate: pipeline/compatibility.py:52 accumulates every candidate for one anchor in dictionaries before bounded selection. Per-anchor memory is Theta(M) and sorting Theta(M log M). Review the supported fan-out contract before assigning severity or choosing a bounded selection accumulator.

Architectural assessment

The repository already has useful patterns: Strategy via backend/provider contracts, Adapter at source/storage/telemetry boundaries, composition roots in API and worker, Factory for work-store selection, and an ordered provider chain for fallback ranking. The reusable admission package preserves its application-independent port.

The demonstrated architectural opportunities are snapshot-bound read consistency, indexed resource lifecycle, and bounded resumable maintenance. These improve invariants behind existing interfaces. Adding generic managers, a service locator, or factories for pure helpers would introduce coupling without addressing these findings.

scripts/validate_architecture.py checks forbidden test/prototype imports, serving_limits independence, and the telemetry adapter boundary. It does not prove complete inward dependency direction, acyclic modules, bounded execution, or runtime consistency. Its green result must not be interpreted as such.

Algorithmic references: RFC 2697 describes finite-capacity token replenishment; the idle-bucket removal recommendation is an inference from that behavior. Python event-loop documentation explains why synchronous CPU work delays an event loop; the observed scan complexity comes from this repository and the local probe.

Coverage and remaining work

Direct first-pass review includes serving candidate loaders, admission core/config/ASGI, metadata similarity, fallback ranking, architecture validator, and relevant API setup. Parallel reviewers inspected personalization service/interactions/health/lifecycle; selected storage, state-store, authorization, reranking and worker paths; pipeline reduction/generation and associated tests. Selected-section inspection is not full-file coverage. Static tooling type-checked 78 source files, but that is not manual review.

The architecture reviewer fully read pipeline/content.py, cooccurrence.py, ingestion.py, groups.py, fallbacks.py, compatibility.py, evaluation.py, backend.py, workstore_factory.py, and tests/unit/test_content.py. The contract reviewer fully read serving.py, personalization/service.py, interactions.py, health.py, and lifecycle.py. Other named large modules received targeted inspection.

Historical review backlog (outside the current planning goal):

  1. Independently revalidate generation counterexamples against end-to-end Training Runs; assess tied-score top-N behavior against the explicit ranking contract.
  2. Review all storage publication, lease recovery, migrations, source consistency, retries and cancellation; qualify PostgreSQL behavior on a disposable database.
  3. Complete ANN residency/search/artifact validation, two-tower training, geography, personalization authorization/tokens/retention and request body bounds.
  4. Complete simulation, synthetic generation/materialization, observability, CLI scripts, packaging, Docker/Compose and CI review, recording every file's coverage.
  5. Inspect tests for missing observables and reference-model independence; measure representative training/serving workloads with seeds, sizes, resource limits, repeated timings and correctness signatures. Include peak memory and tail latency.
  6. Consolidate severity, contract citations, reproductions, pattern recommendations, and any fixes separately authorized by the user; repeat relevant gates after changes.

Later probes confirmed expired-interaction admission and sequence/version regression as C9 and C10. The remediation design covers all 28 confirmed findings and six provisional concerns individually, including verification, dependencies, compatibility, and rollout. Remaining product decisions are identified in that draft; they do not prevent completion of the requested plans.

Commands and handoff

Python environment preflight reported Python 3.14.7, project .venv/bin/python, uv.

  • git status --short: clean baseline; final tree contains two new documentation files, this evidence record and the remediation design. uv-generated SOURCES.txt newline drift was restored to baseline.
  • git log -5 --oneline: established HEAD and recent changes.
  • Read-only cat, sed, nl, rg --files, git diff and IDE search calls: inspected skills, guidance, files and evidence; one attempted content.py path was absent, corrected to pipeline/content.py.
  • make test-static: exit 0; Ruff, mypy, documentation and architecture checks passed.
  • make docs-check: exit 0, initially 61 Markdown files and then 62 with the design; git diff --check: exit 0. Separate whitespace checks cover both new files.
  • make test-focused TEST=tests/unit/test_serving_limits.py: 11 passed, exit 0.
  • .venv/bin/python .build/review_admission_probe.py: exit 0; retention/bypass counterexamples and timings above. Synthetic descriptors only.
  • .venv/bin/python .build/review_for_you_race.py: exit 0; mixed-snapshot assertions.
  • make test: first attempt denied access to the uv cache before tests, exit 2; approved rerun exit 0: static checks passed; 520 unit tests passed, six skipped; 42 contract/delivery passed; 46 integration passed, 22 PostgreSQL cases deselected. Unit suite emitted 18 SQLite datetime-adapter deprecation warnings.
  • .venv/bin/python .build/review-architecture-probe.py: reviewer exit 0; metadata, tie ordering and behavioral eligibility counterexamples. Exact inputs are preserved in that ignored scratch script; convert them to observable regressions if fixing.
  • make test-focused TEST=tests/contract/test_source_adapter.py: seven passed.
  • Five separate focused checks for personalization evaluation, quality, observability, operations configuration, and the observability adapter: 49 passed. An initial invocation supplying multiple paths in TEST failed collection; the supported single-path invocations passed.
  • Additional synthetic probes for stale publication, source consistency, retention, expired interactions, ANN lifecycle and metric compatibility, purchase exclusions, API run identifiers, configuration identity, poll intervals, body bounds, evaluation, profile barriers, and two-tower imports: final corrected runs exited 0. Their observations and limitations are recorded with the findings above. Harness setup failures and an ad hoc OpenMP collision were excluded as production findings.
  • Remediation coverage audit: 16 C findings, seven S findings, five E findings, and six provisional concerns each have an individual plan and verification entry. Planned regression tests have not been implemented or executed.

Not run: PostgreSQL qualification (no disposable database selected), installed-wheel smoke, full capacity/load benchmarks, and optional backend/deployment qualification. Normal test success does not establish those properties. No migrations, hosted settings, external data, packages, or deployed services were changed.