Skip to content

Unique directed two-tower training pairs

Status: implementation and local controlled comparison, 2026-09-30. This addresses M1 in the remediation design. It establishes the pair contract and measures this small synthetic workload; production relevance, resource qualification and broad enablement remain open.

Contract and compatibility

The builder selects at most 10,000 unique directed Item-index pairs. Each pair uses its strongest weighted selection score across co-purchases (weight two) and co-views (weight one), including repeated entries within a signal. Reverse directions remain separate. The existing descending weighted-score/anchor/candidate tuple order is preserved, independent of mapping iteration order. Selection weights choose examples; they do not weight the contrastive loss.

A budget-sized map identifies retained pairs. Its heap prunes stale updated entries before using the cutoff and compacts above twice the budget. Rejected or evicted pairs need no full-graph seen set: the retained cutoff never decreases. A later stronger observation can re-enter the selection. One unique directed pair cannot satisfy the minimum of two training examples by repetition.

Training semantics now use torch-two-tower-v2. Version-one artifacts fail closed and require rebuilding; serving retains its ordinary candidates from the same Recommendation Snapshot. Keep different learned representation identities separate in evidence comparisons. The architecture, normalization, contrastive loss, exported NumPy query encoder and metadata representation remain unchanged. No schema migration is introduced. A clean serving import does not initialize PyTorch.

Controlled temporal comparison

Each arm ran in a fresh process on the same native Apple Silicon host and MPS device, Python 3.14.7, PyTorch 2.14.0 and NumPy 2.5.3. The experiment used the unmodified smoke profile: 1,000 Catalog Items, 25,000 views, 10,000 purchases and 90 days. Source seeds were 17, 29 and 41; the model/projection seed stayed 17. Cutoff was 2026-09-30 00:00 UTC, with a 28-day holdout starting 2026-09-02. Whole shopping contexts use their latest event to choose the temporal split.

The normal bounded reducer streamed generated inputs into derived-only storage. Training pairs came exclusively from the train partition, using fixed support one, shrinkage ten and 200 candidates per anchor. Both arms used a 10,000-slot budget, 32-dimensional metadata projection, three epochs, batches of 128 and the existing AdamW/in-batch contrastive objective. Native index construction and CPU Torch each used one configured thread. No raw interaction rows were staged or exported. These synthetic sources do not contain merchant or Shopper identities.

The baseline replays _pairs from commit 1dfb34a; the corrected arm uses the current builder. Derived train-candidate/holdout-truth SHA3-512 identities matched across each pair of arms. Projected input bytes also matched exactly. Holdout does not select parameters or training pairs. Recall and NDCG use exact exhaustive eligible Item embedding rankings with self exclusion and lexical tie order, evaluated separately against view and purchase holdout truth at K10. This isolates the training change from approximate Faiss retrieval error. It is an Item-relationship comparison, not evaluation of personalized Shopper requests.

Source seed Arm Selected unique pairs Available unique pairs Unique-pair coverage Mean training loss, epochs 1 / 2 / 3
17 Baseline 9,782 125,668 0.077840 4.833005 / 4.822456 / 4.819319
17 Unique 10,000 125,668 0.079575 4.833078 / 4.822711 / 4.819971
29 Baseline 9,773 126,978 0.076966 4.832922 / 4.821362 / 4.818583
29 Unique 10,000 126,978 0.078754 4.832198 / 4.821786 / 4.818302
41 Baseline 9,748 128,048 0.076128 4.832999 / 4.820446 / 4.816383
41 Unique 10,000 128,048 0.078096 4.831877 / 4.820392 / 4.817275

Both arms selected 10,000 total examples. Baseline duplicates spent 218–252 slots. Training loss is the arithmetic mean of batch losses per epoch, including the final shorter batch. Different examples are selected, so similar loss values alone do not establish equal relevance.

Source seed Arm View anchors View Recall@10 View NDCG@10 Purchase anchors Purchase Recall@10 Purchase NDCG@10
17 Baseline 749 0.010368 0.117314 666 0.012614 0.032388
17 Unique 749 0.010627 0.124727 666 0.010241 0.030184
29 Baseline 747 0.009435 0.103114 676 0.007691 0.029188
29 Unique 747 0.009349 0.104149 676 0.008603 0.031286
41 Baseline 748 0.009365 0.102489 664 0.008641 0.021891
41 Unique 748 0.010609 0.117973 664 0.009558 0.027645

The quality changes are mixed: purchase Recall/NDCG decreased for seed 17, view Recall decreased for seed 29, and other measured cells increased. The contract correction does not establish a general relevance improvement. Three synthetic data seeds with one model initialization provide no statistical confidence interval, conversion/uptake claim or production-quality acceptance. Multiple distinct positives for one anchor can still conflict in the unchanged in-batch loss; deduplicating identical pairs does not solve that separate objective question.

Source seed Arm Training Complete build Temporary Python build peak Whole-process peak RSS
17 Baseline 4.247 s 5.278 s 91,095,027 B 961,658,880 B
17 Unique 3.911 s 5.014 s 91,072,157 B 946,552,832 B
29 Baseline 3.807 s 4.836 s 91,092,905 B 949,944,320 B
29 Unique 3.770 s 4.795 s 91,071,737 B 943,996,928 B
41 Baseline 3.589 s 4.548 s 91,071,882 B 960,774,144 B
41 Unique 3.583 s 4.574 s 91,072,193 B 933,593,088 B

Training timing includes loss instrumentation and an MPS completion barrier. Build timing includes projection, model initialization/training, encoding and index serialization. Tracemalloc starts after source reduction and measures temporary Python build allocations, including lazy optimizer initialization; it excludes native/device buffers. Process peak RSS includes imports, source reduction, native state and later exact evaluation, and is not isolated training/device memory. Other local verification work overlapped part of the seed-29 run. Single observations do not establish a speed or memory advantage, a capacity limit, or a serving latency SLO.

Verification and artifacts

Two-tower tests prove the cap boundary, reverse directions, input-order invariance, repeated-update heap bound, exhaustive weighted-selection parity, the two-unique-pair training minimum, rejection of v1 artifacts, real model publication/retrieval and framework-free serving imports. The initial cap regression failed six cases and passed two.

The disposable comparison script is .build/m1_two_tower_comparison.py, invoked with baseline or unique and source seed 17, 29 or 41 in the configured project interpreter. Its six bounded aggregate JSON records remain under .build/m1-<arm>-<seed>.json; they contain no raw source rows or shopping-context identifiers. Commands, gate results and remaining work are recorded in the remediation handoff. The script is local diagnostic evidence, not a committed benchmark interface or a Qualification Evidence Bundle.