Unique directed two-tower training pairs¶
Status: implementation and local controlled comparison, 2026-09-30. This addresses M1 in the remediation design. It establishes the pair contract and measures this small synthetic workload; production relevance, resource qualification and broad enablement remain open.
Contract and compatibility¶
The builder selects at most 10,000 unique directed Item-index pairs. Each pair uses its strongest weighted selection score across co-purchases (weight two) and co-views (weight one), including repeated entries within a signal. Reverse directions remain separate. The existing descending weighted-score/anchor/candidate tuple order is preserved, independent of mapping iteration order. Selection weights choose examples; they do not weight the contrastive loss.
A budget-sized map identifies retained pairs. Its heap prunes stale updated entries before using the cutoff and compacts above twice the budget. Rejected or evicted pairs need no full-graph seen set: the retained cutoff never decreases. A later stronger observation can re-enter the selection. One unique directed pair cannot satisfy the minimum of two training examples by repetition.
Training semantics now use torch-two-tower-v2. Version-one artifacts fail closed and require
rebuilding; serving retains its ordinary candidates from the same Recommendation Snapshot. Keep
different learned representation identities separate in evidence comparisons. The architecture,
normalization, contrastive loss, exported NumPy query encoder and metadata representation remain
unchanged. No schema migration is introduced. A clean serving import does not initialize PyTorch.
Controlled temporal comparison¶
Each arm ran in a fresh process on the same native Apple Silicon host and MPS device, Python
3.14.7, PyTorch 2.14.0 and NumPy 2.5.3. The experiment used the unmodified smoke profile:
1,000 Catalog Items, 25,000 views, 10,000 purchases and 90 days. Source seeds were 17, 29 and 41;
the model/projection seed stayed 17. Cutoff was 2026-09-30 00:00 UTC, with a 28-day holdout
starting 2026-09-02. Whole shopping contexts use their latest event to choose the temporal split.
The normal bounded reducer streamed generated inputs into derived-only storage. Training pairs came exclusively from the train partition, using fixed support one, shrinkage ten and 200 candidates per anchor. Both arms used a 10,000-slot budget, 32-dimensional metadata projection, three epochs, batches of 128 and the existing AdamW/in-batch contrastive objective. Native index construction and CPU Torch each used one configured thread. No raw interaction rows were staged or exported. These synthetic sources do not contain merchant or Shopper identities.
The baseline replays _pairs from commit 1dfb34a; the corrected arm uses the current builder.
Derived train-candidate/holdout-truth SHA3-512 identities matched across each pair of arms.
Projected input bytes also matched exactly. Holdout does not select parameters or training pairs.
Recall and NDCG use exact exhaustive eligible Item embedding rankings with self exclusion and
lexical tie order, evaluated separately against view and purchase holdout truth at K10. This
isolates the training change from approximate Faiss retrieval error. It is an Item-relationship
comparison, not evaluation of personalized Shopper requests.
| Source seed | Arm | Selected unique pairs | Available unique pairs | Unique-pair coverage | Mean training loss, epochs 1 / 2 / 3 |
|---|---|---|---|---|---|
| 17 | Baseline | 9,782 | 125,668 | 0.077840 | 4.833005 / 4.822456 / 4.819319 |
| 17 | Unique | 10,000 | 125,668 | 0.079575 | 4.833078 / 4.822711 / 4.819971 |
| 29 | Baseline | 9,773 | 126,978 | 0.076966 | 4.832922 / 4.821362 / 4.818583 |
| 29 | Unique | 10,000 | 126,978 | 0.078754 | 4.832198 / 4.821786 / 4.818302 |
| 41 | Baseline | 9,748 | 128,048 | 0.076128 | 4.832999 / 4.820446 / 4.816383 |
| 41 | Unique | 10,000 | 128,048 | 0.078096 | 4.831877 / 4.820392 / 4.817275 |
Both arms selected 10,000 total examples. Baseline duplicates spent 218–252 slots. Training loss is the arithmetic mean of batch losses per epoch, including the final shorter batch. Different examples are selected, so similar loss values alone do not establish equal relevance.
| Source seed | Arm | View anchors | View Recall@10 | View NDCG@10 | Purchase anchors | Purchase Recall@10 | Purchase NDCG@10 |
|---|---|---|---|---|---|---|---|
| 17 | Baseline | 749 | 0.010368 | 0.117314 | 666 | 0.012614 | 0.032388 |
| 17 | Unique | 749 | 0.010627 | 0.124727 | 666 | 0.010241 | 0.030184 |
| 29 | Baseline | 747 | 0.009435 | 0.103114 | 676 | 0.007691 | 0.029188 |
| 29 | Unique | 747 | 0.009349 | 0.104149 | 676 | 0.008603 | 0.031286 |
| 41 | Baseline | 748 | 0.009365 | 0.102489 | 664 | 0.008641 | 0.021891 |
| 41 | Unique | 748 | 0.010609 | 0.117973 | 664 | 0.009558 | 0.027645 |
The quality changes are mixed: purchase Recall/NDCG decreased for seed 17, view Recall decreased for seed 29, and other measured cells increased. The contract correction does not establish a general relevance improvement. Three synthetic data seeds with one model initialization provide no statistical confidence interval, conversion/uptake claim or production-quality acceptance. Multiple distinct positives for one anchor can still conflict in the unchanged in-batch loss; deduplicating identical pairs does not solve that separate objective question.
| Source seed | Arm | Training | Complete build | Temporary Python build peak | Whole-process peak RSS |
|---|---|---|---|---|---|
| 17 | Baseline | 4.247 s | 5.278 s | 91,095,027 B | 961,658,880 B |
| 17 | Unique | 3.911 s | 5.014 s | 91,072,157 B | 946,552,832 B |
| 29 | Baseline | 3.807 s | 4.836 s | 91,092,905 B | 949,944,320 B |
| 29 | Unique | 3.770 s | 4.795 s | 91,071,737 B | 943,996,928 B |
| 41 | Baseline | 3.589 s | 4.548 s | 91,071,882 B | 960,774,144 B |
| 41 | Unique | 3.583 s | 4.574 s | 91,072,193 B | 933,593,088 B |
Training timing includes loss instrumentation and an MPS completion barrier. Build timing includes projection, model initialization/training, encoding and index serialization. Tracemalloc starts after source reduction and measures temporary Python build allocations, including lazy optimizer initialization; it excludes native/device buffers. Process peak RSS includes imports, source reduction, native state and later exact evaluation, and is not isolated training/device memory. Other local verification work overlapped part of the seed-29 run. Single observations do not establish a speed or memory advantage, a capacity limit, or a serving latency SLO.
Verification and artifacts¶
Two-tower tests prove the cap boundary, reverse directions, input-order invariance, repeated-update heap bound, exhaustive weighted-selection parity, the two-unique-pair training minimum, rejection of v1 artifacts, real model publication/retrieval and framework-free serving imports. The initial cap regression failed six cases and passed two.
The disposable comparison script is .build/m1_two_tower_comparison.py, invoked with
baseline or unique and source seed 17, 29 or 41 in the configured project interpreter.
Its six bounded aggregate JSON records remain under .build/m1-<arm>-<seed>.json; they contain
no raw source rows or shopping-context identifiers. Commands, gate results and remaining work
are recorded in the remediation handoff. The script is local
diagnostic evidence, not a committed benchmark interface or a Qualification Evidence Bundle.