GPU-Serving Two-Tower Models for Lightweight Ads Engagement Prediction (opens in new tab)
Pinterest replaced its CPU-served two-tower model for ads lightweight ranking with a GPU-serving architecture based on MMOE and DCN. The more expressive model maintained latency comparable to the CPU baseline while reducing offline CTR loss by 5–10%. Separating standard and shopping ad models produced another 5–10% loss reduction and doubled offline iteration speed, with online improvements in CPC and CTR.
Role of Lightweight Ranking
- Lightweight ranking serves as an intermediate stage in Pinterest’s ads recommendation pipeline.
- It filters a large pool of candidate ads before more complex downstream ranking models process them.
- The two-tower design balances quality and latency:
- The Pin tower generates ad embeddings offline through batch updates.
- The query tower generates real-time user embeddings.
- The prediction score is the sigmoid of the embeddings’ dot product.
MMOE-DCN Model Architecture
- The new system replaces the previous Multi-Task Multi-Domain (MTMD) model.
- It combines:
- Multi-gate Mixture-of-Experts (MMOE) with MLP-based gating.
- Deep & Cross Network (DCN) layers for modeling feature interactions.
- Each expert uses both full-rank and low-rank DCN layers.
- Unlike MTMD, MMOE handles multi-task and multi-domain learning without relying on separate domain-specific modules.
- GPU serving makes it practical to deploy this larger and more computationally demanding model while preserving CPU-baseline latency.
Scenario-Specific Modeling
- Standard and shopping ad scenarios are served as separate models.
- Each model is trained only on data relevant to its scenario.
- This specialization delivered an additional 5–10% reduction in offline loss.
- Separating the models also doubled the speed of offline model iteration.
Training Efficiency Improvements
- Dataloader optimization
- GPU prefetching prepares the next batch while the current batch is processed.
- Additional worker threads take advantage of the 1 TB of CPU memory available on p4d instances.
- Model code optimization
- Operations that previously allocated zero-filled tensors on the CPU were moved to the GPU.
- Fused kernels replaced multiple individual kernels to reduce execution overhead.
- Training configuration
- BF16 precision improved processing speed.
- Larger batch sizes increased GPU memory utilization.
Evaluation Results
- The model uses downstream ranking scores as labels and optimizes KL divergence between those labels and its predictions.
- Evaluation covers both:
- Auction winners—ads ultimately inserted and shown to users.
- Auction candidates—ads passed to downstream ranking.
- Offline loss decreased significantly across all evaluated slices.
- Online experiments showed:
- Lower cost per click (CPC), which is favorable.
- Higher click-through rate (CTR).
GPU-serving a more complex MMOE-DCN two-tower model allowed Pinterest to improve ad engagement prediction without sacrificing serving latency. The results support using GPU infrastructure, scenario-specific models, and targeted training optimizations to scale lightweight ranking systems.