Repository review evidence¶
Status: baseline review evidence; remediation implementation in progress. Baseline: 1dfb34a, 2026-09-30.
The user changed the goal to detailed fixes for every issue found. The linked remediation design is the current deliverable. The coverage limits below remain evidence limits; further repository review is not required to complete this planning task. The subsequent implementation request and current red/green evidence are tracked in the remediation handoff. Findings below describe the review baseline.
The preceding objective was repository-wide correctness, scalability, performance, and architecture review. These passes do not certify every line or production capacity. At that review baseline, production source and committed tests were unchanged. Findings distinguish executable counterexamples from static analysis and from unresolved questions.
Method and acceptance evidence¶
Review contracts before implementations; trace source, reduction, publication, and serving boundaries; examine tests for counterexamples they omit. Measure scaling over multiple input sizes, preserve inputs and commands, and distinguish asymptotic reasoning from local timing. A passing suite establishes covered behavior only.
Patterns are evaluated against concrete pressures: variation, isolation, consistency, bounded state, and resource lifecycle. Pattern count is not a measure of architectural quality. Prefer a narrow abstraction with an invariant over additional layers without an identified responsibility.
Completion of that review would have required an inventory covering production packages, migrations, scripts, delivery/configuration, and tests; reviewed contracts and failure paths; reproduced high-impact findings; workload-based performance evidence; and explicit remaining risks. That evidence is incomplete today.
Spec and correctness findings¶
C1 — P1: For You mixes two published snapshots¶
src/recommendations/serving.py:200 reads global candidates and then independently
reads the current Best Sellers head. Each repository operation can be correct while a
publication between calls returns different snapshots. The resulting CandidateBundle
uses the first snapshot's metadata and the second snapshot's fallback. Snapshot-local
serving is consequently violated without any cross-scope query.
Executable probe .build/review_for_you_race.py uses the actual loader and domain
values with a deterministic reader advancing from snapshot A to B. Output:
reported_snapshot=snapshot-A
candidate=A-only fallback=B-only
Confirmed: fallback from snapshot B is packaged with snapshot A metadata.
Recommendation: resolve one immutable snapshot identity, then read both lanes by that identity, preferably through a cohesive repository operation. A snapshot-bound read view is a useful pattern here. A race regression must exercise a publication between the reads and verify every entry and disclosed identity belong to one snapshot.
C2 — P1: Valid feature-poor Catalogs can abort Similar Items generation¶
src/recommendations/pipeline/content.py:167 considers zero price and whitespace text
metadata. _price_buckets omits nonpositive prices; document processing omits whitespace.
With two such products, _metadata_matrix calls hstack([]) at line 122. Punctuation
text reaches the word vectorizer at line 113, which raises empty-vocabulary ValueError,
even when structured category features are usable.
Executable reviewer probes reproduced IndexError for two zero-price rows and two whitespace-title rows; punctuation-only titles reproduced ValueError. Treat absent usable feature blocks as insufficient evidence and qualify word/character blocks independently. Tests must distinguish empty metadata, unusable text, and structured metadata with unusable text; ordinary usable features must retain their scores.
C3 — P2: Candidate truncation precedes eligibility filtering¶
src/recommendations/pipeline/workstore.py:516 ranks and limits behavioral candidates
before final eligibility checks. pipeline/run.py:521 supplies no eligibility mapping
for unconstrained behavioral lanes. pipeline/fallbacks.py later removes ineligible
products, but cannot recover qualifying candidates discarded by the earlier cutoff.
Reviewer probe: anchor A, 200 higher-scoring ineligible neighbors, and eligible C with
qualifying lower support; generation limit 200 drops C, yielding no primary
recommendations after eligibility filtering. Global popularity has the same ordering
risk: _score_candidates at pipeline/run.py:1595 truncates its input mapping before
the downstream eligibility filter.
Recommendation: apply publication eligibility before candidate top-N, across DuckDB, Polars, popularity, and applicable category/geographic paths. Verify against an independent reference with more excluded candidates than the generation limit.
C4 — P2: Missing partition identity can bypass applicable limits¶
src/serving_limits/core.py:274 collapses a matching rule with missing partition
descriptors into the same empty result as genuinely unmatched traffic. With
allow_unmatched=True, AdmissionController admits it without enforcing any matched
rule. This option should exempt unmatched traffic, not malformed matching traffic.
The admission probe prints matching_rule_missing_partition_allowed: true for a
globally matching consumer-partitioned rule and a request with only a service
descriptor. The existing missing-partition test covers only the default closed policy.
Represent malformed resolution separately from no selector match and test both policy
settings, including a parent rule alongside a malformed child.
C5 — P2: Similar Items discards lexical tie winners before sorting¶
The native sparse top-N at pipeline/content.py:50 selects tied candidates before
Python applies Product ID ordering. Probe: 210 identical category vectors, top_n=200,
anchor 000; the first ten returned IDs are 009 through 018, rather than 001
through 010. Ascending Product ID tie ordering is explicit in
specs/commerce-recommendation-service.md:97. Sorting survivors cannot recover
discarded winners. Make secondary ordering part of bounded selection and test more
tied candidates than the cutoff, including input permutations and native thread counts.
Standards and scalability findings¶
The next confirmed correctness findings are recorded separately below; finding IDs remain stable as coverage expands.
S6 — P2: Request bodies have no service-enforced byte bound¶
The POST routes in src/recommendations/api.py register ordinary FastAPI body models
without an ASGI byte limiter. Field and tuple bounds take effect after body buffering
and parsing. A chunked Training Run request containing the valid 80-byte JSON scope
plus whitespace was accepted with either 1 MiB or 8 MiB of padding. The service consumed
all bytes and returned 202. Traced Python peak allocations were 3,440,265 bytes and
16,933,459 bytes respectively. Probe: .build/review-body-bound.py.
These are bounded local memory observations, not RSS or throughput qualification. Source inspection establishes the lack of a byte-limit branch. The repository's AGENTS invariant explicitly requires bounded request bodies; an externally configured proxy limit is not enforced by this app. Enforce both declared and streamed body bytes before JSON parsing, including chunked requests and absent or incorrect Content-Length. Preserve authentication-before-body behavior on protected routes and use a stable 413 response. Test early termination of receive, rather than merely eventual rejection.
S7 — P2: Exact ANN tie selection allocates every tied match¶
src/recommendations/ann.py:416-423 materializes and sorts every match at the selection
threshold before retaining at most 100. The real _exact path, probed with equal unit
vectors, instantiated 10,000 AnnMatch objects for a 10,000-Item Catalog and 100,000 for
a 100,000-Item Catalog; both returned 100 entries. The larger probe used about 22.7 MB
of traced temporary allocations and 0.28 seconds, versus about 2.3 MB at 10,000 Items.
The full score array is documented; the additional Theta(N) Python match residency and
worst-case Theta(N log N) tie sorting are avoidable. Preserve exact score/Product-ID
ordering with bounded selection, and exercise underfilled native retrieval as well as
small-index exact search. These diagnostic measurements are not production latency
claims. Probe: .build/review-ann-probe.py.
Further spec and correctness findings¶
C6 — P1: A stale publisher destroys its successor's staged snapshot¶
src/recommendations/storage.py:1093-1121 deletes any building snapshot for a run before
checking lease ownership at lines 1241-1251. A recovered worker starts publication;
the old worker resumes, replaces the new owner's staging, and only later fails its
ownership check. The legitimate publisher then fails because its own staged parent
has disappeared. The last serving head can survive while forward publication fails.
.build/review_stale_publication.py uses actual repositories with SQLite foreign keys
enabled, an old lease recovered by a new owner, and a deterministic generator
interleaving. Output:
Old owner rejected for lost lease after replacing new owner staging.
Current owner publication failed: IntegrityError
Published snapshots: 0; successor staging was destroyed by stale owner.
Fence ownership before destructive staging operations and make batches/cleanup belong to a fenced generation. The check and mutation must be atomic with lease recovery; an isolated preliminary check alone is insufficient. Existing stale-publisher tests protect an available head but omit overlapping building publishers. Verify this race on PostgreSQL as well; this reproduction is not PostgreSQL locking qualification.
C7 — P1: Default SQLite source transactions are not repeatable reads¶
src/recommendations/source.py:350-354 calls SQLAlchemy begin with SERIALIZABLE, but
worker.py:631 constructs default sqlite3 engines. Under the pinned Python driver,
legacy transaction control does not issue BEGIN for SELECT. Fully consuming one
stream and opening another can observe a new committed source state without changing
the consistency token. This defeats the consistent source-read contract even when
every query is correctly scoped and cutoff-filtered.
Probe .build/review-source-consistency.py uses a disposable WAL-mode database and
separate reader/writer connections. The default reader returns A, a writer commits B,
then the same read returns A and B; its token check passes. A control that explicitly
emits BEGIN returns only A on the second read. Both assertions pass.
SQLAlchemy's SQLite transaction documentation explains this driver behavior. This repository's technical documentation acknowledges that token identity alone does not establish snapshot stability. Configure genuine driver transaction semantics and add cross-query concurrent-writer tests to the source contract suite. The existing seven source-adapter tests pass without testing it.
C8 — P1: Scoped personalization retention configuration is ignored¶
config.py:1305 accepts retention_days per capability, but storage.py:2102-2106
always builds the state repository with its default 90-day policy. Production API
wiring at api.py:1521,1547 never forwards the configured scope policies.
The actual app/store probe .build/review_personalization_retention.py configures
30 days, commits a synthetic interaction, and observes a 90-day ledger expiry. At
day 31, maintenance removes zero rows and the profile remains. Parsing-only tests do
not establish enforcement. Resolve retention by Commerce Scope inside persistence;
one global replacement cannot represent multiple scope-specific policies. Verify
distinct policies on two isolated scopes, idempotency retention, and shortened-policy
rollouts against the advertised maximum-history contract.
C9 — P2: Already expired interactions immediately affect personalization¶
personalization/health.py:37-47 rejects future skew but permits arbitrarily old
interactions. state_store.py:733-744 then adds their affinities despite the expired
retention boundary. A 91-day-old viewed interaction receives an acknowledgement and
an item affinity of 1.0 under the default 90-day policy. This contradicts the maximum
history behavior documented in docs/shopper-personalization.md.
Choose the explicit late-event contract: reject expired interactions or acknowledge
their ordering without deriving expired features. Do not rely on a later maintenance
pass to establish serving-time retention. Probe: .build/review_expired_interaction.py.
C10 — P1: Expiry rolls acknowledged watermarks backward¶
personalization/state_store.py:687-700 reconstructs sequence and profile version from
the last retained interaction, conflating acknowledgement metadata with disposable
history. The interaction contract requires exactly the next sequence and a valid
causal token to observe at least the acknowledged version.
Probe: sequence 1 is recent, sequence 2 occurred 91 days earlier, and both are
acknowledged. Expiry removes sequence 2 and changes the profile from sequence/version
2/2 to 1/1. The client's valid next sequence 3 is rejected with last_sequence=1. Full
expiry can erase the watermark altogether. Existing tests cover increasing occurrence
time, missing this reverse-order case. Maintain monotonic sequence/version separately
from the retained feature ledger; test late events, full expiry, causal reads and retries.
Probe: .build/review_expired_interaction.py.
C11 — P1: Completed ANN loads can block all later generations¶
ann.py:508-516 only reaps a completed Future when its own snapshot is requested
again. If a new head replaces it first, the old pending entry remains; requests for
new generations or other scopes always fall back without scheduling another load.
Production config enables background_load at api.py:1527-1528.
The deterministic probe explicitly awaits old-head load completion, then requests the
new head three times. Only the old snapshot reaches the artifact reader, and the old
completed entry remains pending. Reap completed work independently of the requested
identity using the scheduler's existing lock. Test completed success, missing artifact,
failure, head replacement and another scope. Probe: .build/review-ann-probe.py.
C12 — P2: Swimlane recent-purchase exclusions stop after ten Items¶
swimlane_service.py:177-178 takes ten purchases as retrieval seeds and reuses them
as the exclusion set at line 217. A profile can retain 100 purchases. With eleven
recent purchases and an explicit exclude_recent_purchases Best Sellers step, the
eleventh purchase is returned. The exclusion behavior in specification lines 535-536
does not define a ten-purchase subset. Separate the full bounded exclusion history
from the smaller retrieval seed budget. Probe: .build/review-swimlane-purchase-probe.py.
C13 — P2: ANN artifact validation accepts the wrong distance metric¶
ann.py:305-317 validates native type and dimensions but omits metric_type. A correctly
checksummed HNSWFlat index with METRIC_L2 loads successfully, while serving interprets
native distances as inner-product similarity and exact fallback computes inner product.
The ANN lifecycle spec explicitly requires metric binding and incompatible-metric
tests. This is trusted-artifact compatibility validation, not a claim about malicious
artifact authenticity. Reject unsupported metrics at loading and qualify native/exact
score parity. Probe: .build/review-ann-metric-probe.py.
C14 — P2: Unvalidated run IDs turn client errors into server errors¶
api.py:490 constructs TelemetryContext with the unvalidated path run_id. Its 255-character
identifier guard raises before the repository's missing-run response. A 256-character
run ID, or a credential-shaped ID, produces HTTP 500, even with disabled observability;
an ordinary unknown ID produces 404. Validate transport input before constructing
telemetry values and return the chosen 4xx contract without echoing rejected material.
Probe: .build/review-api-run-id.py; actual TestClient, server exceptions disabled.
C15 — P2: Identical config fingerprints can select different parameters¶
config.py:1026-1035 sorts search grids in the configuration fingerprint, but no-holdout
selection at pipeline/run.py:1050-1051 uses the first grid value when the grid excludes
the default. Trending selection similarly uses candidates[0] at lines 1114-1119, and
equidistant tie selection can retain iteration order. Configured view half-life grids
[2,8] and [8,2] therefore share config_version but choose 2 and 8 days respectively.
This invalidates configuration identity as comparable-run/reproducibility evidence.
Either canonicalize runtime choice to match hashing or preserve semantically meaningful
order in the fingerprint. Test no-holdout and tied-holdout selection for every grid.
Probe: .build/review_config_grid_identity.py.
C16 — P3: Nonfinite polling configuration crashes idle workers¶
config.py:1001 checks only worker_poll_seconds <= 0. NaN passes that check and reaches
the worker's sleep, which raises ValueError. The configured-Python probe
.build/review_poll_interval.py reproduced validation acceptance and sleep failure.
Validate finite positive intervals before constructing the runtime; cover NaN, both
infinities, zero and negative values in configuration tests.
E1 — P2: Duplicate observations bypass minimum-count suppression¶
personalization_evaluation.py:275-284 accepts repeated ranking observations.
Repeating one Shopper observation ten times produces an available threshold-10 report
with count 10 and Recall 1.0. Count independent evaluation units and reject duplicate
observations. This is an offline evidence-contract issue, not a demonstrated API leak.
E2 — P2: Outcome attribution accepts an unexposed variant¶
personalization_evaluation.py:613-635 omits variant from exposure correlation and
does not compare the stored assignment. A holdout exposure permits personalized-variant
attribution. Bind and verify the complete assignment identity.
E3 — P2: Delimiter-joined HMAC inputs have ambiguous scope identity¶
personalization_evaluation.py:498-507,626-634 joins tuple fields with colon without
escaping or length prefixes. Property a:b / experiment c and property a / experiment
b:c produce identical exposure correlation material, permitting cross-property
attribution in the offline helper. Use unambiguous versioned canonical encoding for
assignment and exposure, without compatibility fallback to ambiguous digests.
E4 — P2: Nonfinite latency silently passes experiment guardrails¶
personalization_evaluation.py:658-659,685-693 permits NaN latency and nonfinite latency
thresholds. The guardrail function returns no breach for NaN p95_latency_ms. Require
finite measurements and thresholds; missing evidence must not become rollout approval.
E5 — P3: Suppressed metric summaries can retain a hidden value¶
personalization_evaluation.py:205-206 accepts a suppressed summary containing
recall=0.5 and ndcg=None. Require both metrics absent for suppressed/not-applicable
status and both finite bounded metrics for available status. Normal summary construction
is currently safe; the value-object boundary is not.
All five E findings were reproduced in .build/review-evaluation-probe.py. Searches
found no external production callers for these offline helpers. Their severity refers
to evidence integrity and prospective adoption, not a current serving exploit.
M1 — Model concern requiring further qualification¶
two_tower_torch.py:134-145 returns duplicate directed positive pairs when purchase
and view evidence contain the same relationship, despite the helper's unique-pair
contract. The probe returns ((0,1),(0,1)), consuming two places in the capped training
pair budget. Duplicate targets can also become in-batch negatives under diagonal
cross-entropy. Missing deduplication is reproduced; relevance impact and intended
cross-signal weighting need a controlled model experiment before a stronger claim.
An early combined ANN/torch probe aborted with a duplicate OpenMP runtime error. Clean import and actual production-builder subprocesses succeeded. That abort is excluded from repository findings; it is a probe interaction, not an established production failure.
Additional standards and scalability findings¶
S1 — P2: Rate-bucket residency grows with lifetime cardinality¶
src/serving_limits/core.py:451 retains token buckets for every observed partition.
No idle expiry or maximum bucket count exists; lease expiry does not remove rate
buckets. Descriptor length bounds constrain one key, not the number of keys over time.
Probe: admit 10,000 distinct synthetic consumers, advance time to 1,000,000 seconds, trigger maintenance through acquire. All 10,000 token buckets remain. This is O(U) resident state in lifetime partition cardinality U, including fully replenished buckets. Expire buckets only when replenishment makes removal semantically equivalent to a fresh bucket; apply bounded, incremental maintenance and an explicit cardinality policy. Blind LRU removal of depleted buckets would incorrectly restore capacity.
S2 — P2: Every admission scans all live leases¶
src/serving_limits/core.py:425 invokes _expire_leases on each admission; line 550
reconstructs the entire active lease dictionary. Release also scans it. For N successful
admissions retained within one TTL, total scanning is O(N²), and each operation runs
synchronously under a process lock on the calling async path. Denial retry calculation
adds another global scan and sorting.
Probe: one concurrency rule, capacity N+1, frozen clock 0, TTL 30 seconds, N admissions without release, three fresh-state repetitions per N. Measured seconds:
| N | Repetitions | Median |
|---|---|---|
| 1,000 | 0.06017, 0.06050, 0.06183 | 0.06050 |
| 2,000 | 0.22672, 0.22479, 0.22337 | 0.22479 |
| 4,000 | 0.88248, 0.88418, 0.89026 | 0.88418 |
Doubling N increased time by approximately 3.72× and 3.93×. This local microbenchmark supports the code-level quadratic mechanism, not a production throughput or latency claim. Use an expiry priority queue and bucket-local permit indexes while preserving atomic multi-rule acquisition, partial permit expiry, and idempotent release.
S3 — P2: Personalization expiry uses an unbounded transaction¶
Static evidence: src/recommendations/personalization/state_store.py:346 materializes
all expired Shopper keys, deletes all expired rows, and rebuilds every affected profile
inside one transaction. _rebuild_profile at line 670 additionally materializes the
entire retained history for each Shopper. Time retention does not bound interaction
count. Large backlogs increase memory, transaction duration, locks, and database work.
Recommendation: a bounded, resumable maintenance operation with limited keys/rows per transaction and streamed reduction. Qualification must cover concurrent ingestion, interruption/retry, preservation of profile versions, and eventual backlog clearance. This finding has not yet been reproduced against a large database.
S4 — P2: Holdout evaluation materializes the entire pair graph¶
Static evidence: pipeline/workstore.py:633 fetches all holdout pairs and builds two
set memberships per edge. pipeline/run.py:984 loads that graph during search and
line 420 loads it again for retained evidence. Polars repeats materialization at
polars_workstore.py:404. For E distinct pairs, Python memory is Theta(E), with a
worst-case quadratic graph in Catalog size, outside DuckDB spill limits. Bounding each
interaction group does not bound the accumulated graph. Stream evaluation or introduce
an explicitly qualified evaluation sample; preserve exact metric meanings or disclose
sampling changes. Large-scale memory reproduction remains outstanding.
S5 — P2: Complementary-category fallback repeats equivalent full sorts¶
Static evidence: pipeline/run.py:1548 sorts every Item in paired categories separately
for every anchor, builds full candidate lists, then truncates to 200. For two equally
sized categories, sorting work is Theta(N² log N) although anchors in a category share
the same candidate ordering. Precompute one bounded list per category and reuse the
existing KeyedCandidateProvider pattern. Benchmark with growing category sizes and
verify exact scores, ordering, eligibility and fallback provenance.
S8, additional scalability candidate: pipeline/compatibility.py:52 accumulates every
candidate for one anchor in dictionaries before bounded selection. Per-anchor memory
is Theta(M) and sorting Theta(M log M). Review the supported fan-out contract before
assigning severity or choosing a bounded selection accumulator.
Architectural assessment¶
The repository already has useful patterns: Strategy via backend/provider contracts, Adapter at source/storage/telemetry boundaries, composition roots in API and worker, Factory for work-store selection, and an ordered provider chain for fallback ranking. The reusable admission package preserves its application-independent port.
The demonstrated architectural opportunities are snapshot-bound read consistency, indexed resource lifecycle, and bounded resumable maintenance. These improve invariants behind existing interfaces. Adding generic managers, a service locator, or factories for pure helpers would introduce coupling without addressing these findings.
scripts/validate_architecture.py checks forbidden test/prototype imports,
serving_limits independence, and the telemetry adapter boundary. It does not prove
complete inward dependency direction, acyclic modules, bounded execution, or runtime
consistency. Its green result must not be interpreted as such.
Algorithmic references: RFC 2697 describes finite-capacity token replenishment; the idle-bucket removal recommendation is an inference from that behavior. Python event-loop documentation explains why synchronous CPU work delays an event loop; the observed scan complexity comes from this repository and the local probe.
Coverage and remaining work¶
Direct first-pass review includes serving candidate loaders, admission core/config/ASGI, metadata similarity, fallback ranking, architecture validator, and relevant API setup. Parallel reviewers inspected personalization service/interactions/health/lifecycle; selected storage, state-store, authorization, reranking and worker paths; pipeline reduction/generation and associated tests. Selected-section inspection is not full-file coverage. Static tooling type-checked 78 source files, but that is not manual review.
The architecture reviewer fully read pipeline/content.py, cooccurrence.py,
ingestion.py, groups.py, fallbacks.py, compatibility.py, evaluation.py,
backend.py, workstore_factory.py, and tests/unit/test_content.py. The contract
reviewer fully read serving.py, personalization/service.py, interactions.py,
health.py, and lifecycle.py. Other named large modules received targeted inspection.
Historical review backlog (outside the current planning goal):
- Independently revalidate generation counterexamples against end-to-end Training Runs; assess tied-score top-N behavior against the explicit ranking contract.
- Review all storage publication, lease recovery, migrations, source consistency, retries and cancellation; qualify PostgreSQL behavior on a disposable database.
- Complete ANN residency/search/artifact validation, two-tower training, geography, personalization authorization/tokens/retention and request body bounds.
- Complete simulation, synthetic generation/materialization, observability, CLI scripts, packaging, Docker/Compose and CI review, recording every file's coverage.
- Inspect tests for missing observables and reference-model independence; measure representative training/serving workloads with seeds, sizes, resource limits, repeated timings and correctness signatures. Include peak memory and tail latency.
- Consolidate severity, contract citations, reproductions, pattern recommendations, and any fixes separately authorized by the user; repeat relevant gates after changes.
Later probes confirmed expired-interaction admission and sequence/version regression as C9 and C10. The remediation design covers all 28 confirmed findings and six provisional concerns individually, including verification, dependencies, compatibility, and rollout. Remaining product decisions are identified in that draft; they do not prevent completion of the requested plans.
Commands and handoff¶
Python environment preflight reported Python 3.14.7, project .venv/bin/python, uv.
git status --short: clean baseline; final tree contains two new documentation files, this evidence record and the remediation design. uv-generated SOURCES.txt newline drift was restored to baseline.git log -5 --oneline: established HEAD and recent changes.- Read-only
cat,sed,nl,rg --files,git diffand IDE search calls: inspected skills, guidance, files and evidence; one attemptedcontent.pypath was absent, corrected topipeline/content.py. make test-static: exit 0; Ruff, mypy, documentation and architecture checks passed.make docs-check: exit 0, initially 61 Markdown files and then 62 with the design;git diff --check: exit 0. Separate whitespace checks cover both new files.make test-focused TEST=tests/unit/test_serving_limits.py: 11 passed, exit 0..venv/bin/python .build/review_admission_probe.py: exit 0; retention/bypass counterexamples and timings above. Synthetic descriptors only..venv/bin/python .build/review_for_you_race.py: exit 0; mixed-snapshot assertions.make test: first attempt denied access to the uv cache before tests, exit 2; approved rerun exit 0: static checks passed; 520 unit tests passed, six skipped; 42 contract/delivery passed; 46 integration passed, 22 PostgreSQL cases deselected. Unit suite emitted 18 SQLite datetime-adapter deprecation warnings..venv/bin/python .build/review-architecture-probe.py: reviewer exit 0; metadata, tie ordering and behavioral eligibility counterexamples. Exact inputs are preserved in that ignored scratch script; convert them to observable regressions if fixing.make test-focused TEST=tests/contract/test_source_adapter.py: seven passed.- Five separate focused checks for personalization evaluation, quality,
observability, operations configuration, and the observability adapter: 49 passed.
An initial invocation supplying multiple paths in
TESTfailed collection; the supported single-path invocations passed. - Additional synthetic probes for stale publication, source consistency, retention, expired interactions, ANN lifecycle and metric compatibility, purchase exclusions, API run identifiers, configuration identity, poll intervals, body bounds, evaluation, profile barriers, and two-tower imports: final corrected runs exited 0. Their observations and limitations are recorded with the findings above. Harness setup failures and an ad hoc OpenMP collision were excluded as production findings.
- Remediation coverage audit: 16 C findings, seven S findings, five E findings, and six provisional concerns each have an individual plan and verification entry. Planned regression tests have not been implemented or executed.
Not run: PostgreSQL qualification (no disposable database selected), installed-wheel smoke, full capacity/load benchmarks, and optional backend/deployment qualification. Normal test success does not establish those properties. No migrations, hosted settings, external data, packages, or deployed services were changed.