Architectural Optimizations for Boosting the Performance of Recommendation Systems
Open Access
- Author:
- Jain, Rishabh
- Graduate Program:
- Computer Science and Engineering
- Degree:
- Doctor of Philosophy
- Document Type:
- Dissertation
- Date of Defense:
- July 10, 2025
- Committee Members:
- Mahmut Kandemir, Program Head/Chair
Chitaranjan Das, Chair of Committee
Mahmut Kandemir, Dissertation Advisor
Anand Sivasubramaniam, Major Field Member
John Kalamatianos, Special Member
Onur Kayiran, Special Member
Amulya Yadav, Outside Unit & Field Member - Keywords:
- Deep Learning Recommendation Models
DLRM
Inference
CPU and GPU Performance
Software Prefetching and Cache Pinning
Embedding Stage
Heterogeneous Computing
Memory Hierarchy Optimization
Multithreading Performance
AI System Performance
Performance Characterization - Abstract:
- Personalized recommendation has become one of the most pervasive AI applications on the internet, powering user experiences across e-commerce, social media, and entertainment platforms. Deep Learning Recommendation Models (DLRMs) form the algorithmic backbone of these systems, combining dense and sparse features to deliver accurate personalization at scale. These models are deployed massively across datacenters using CPUs and GPUs, accounting for a major share of all AI inference cycles and driving core revenue for hyperscalers like Meta and Google. However, the rapid growth in model complexity, with trillions of parameters and thousands of embedding tables, has amplified the compute, memory, and latency challenges of inference. Since recommendation is a user-interactive application, inference must meet stringent service-level agreement (SLA) requirements, motivating the need for architectures that can deliver high throughput and low latency. The objective of this dissertation is to improve the performance of DNN-based recommendation systems using architectural optimizations across CPUs, GPUs, and heterogeneous CPU–GPU platforms. Specifically, the dissertation targets DLRM inference. On CPUs, our in-depth performance breakdown study across multiple DLRM models and datasets, reveals that the embedding stage dominates end-to-end inference time due to irregular memory accesses. To capture the diversity in access patterns, production traces are categorized into datasets of varying hotness, and their behavior is analyzed through cache profiling and reuse distance modeling. Based on these insights, the dissertation proposes application-specific software prefetching to hide long access latencies and application-aware HyperThreading to overlap computation and memory operations, leading to improved core utilization on server-grade CPUs. As DLRMs continue to scale in table count, pooling operations, and embedding dimension, the work further explores alternative approaches for parallelization of the embedding stage, introducing Table Threading (TT) and Hierarchical Threading (HT) to parallelize individual batches across cores. These techniques result in significant performance benefits. However, even with these approaches, the performance scaling remains imperfect for heterogeneous datasets, where embedding tables vary widely in size, reuse, and pooling operations. To overcome this imbalance, the dissertation develops a load- and memory-level-parallelism (MLP)-aware scheduling framework, which analytically models each table’s workload using reuse-distance–based metrics. This enables intelligent categorization and scheduling of tables across threads. By incorporating table reordering, core grouping, strategic co-location, and lightweight work stealing, the proposed design achieves lower-latency, improved utilization, and enhanced scalability on modern many-core CPUs. On GPUs, this dissertation examines the architectural inefficiencies of DLRM inference under growing compute and memory-bandwidth demands. While GPUs offer massive parallelism and high-bandwidth memory, detailed microarchitectural analysis of NVIDIA A100 and H100 shows that the embedding stage continues to be the primary bottleneck due to low occupancy and long memory stalls. To mitigate these issues, the work proposes compiler-guided optimizations, software prefetching, and L2 cache pinning techniques that improve warp utilization and hide latency through purely software-based means, thus boosting overall embedding stage performance. Building on these insights, the dissertation introduces HetT-HetD, a heterogeneous CPU–GPU execution framework that challenges the one-device-fits-all paradigm by leveraging the inherent device and table diversity in modern DLRMs. By mapping high-reuse, large-capacity tables to CPUs and low-reuse, bandwidth-intensive tables to GPUs, and coordinating them via asynchronous data movement and heuristic-driven scheduling, HetT-HetD leverages both CPU and GPU performantly. Overall, this dissertation advances the architectural understanding and system-level optimization of large-scale DLRM inference, with a central emphasis on achieving low-latency and high-efficiency recommendation serving. All proposed techniques are realized entirely in software and require no modifications to existing hardware or publicly distributed models, making them directly deployable in current datacenter environments. By systematically improving CPU, GPU, and heterogeneous CPU–GPU execution, this work establishes a new performance baseline for future research exploring hardware and software co-designs aimed at further accelerating deep learning–based recommendation systems.
Accessible Version in Progress
We're generating an accessible version of this file to meet ADA Title II requirements. This process may take up to one hour. Please return later to access the accessible copy once it's ready.
You can still download the current version by clicking "OK".
What's happening:
An accessible PDF is being generated using Adobe with AI used to generate alternative text (alt text) for images in the PDF.