Scalability and Large-Scale Systems Engineering
Taming the tail utilization of ads inference at Meta scale

Taming the tail utilization of ads inference at Meta scale

7/10/2024 · Rohith Menon, Bikash Sharma, Deepak Tiwari

What this post added

This post details optimizations to address tail utilization in Meta's ads inference services. Key contributions include tuning load balancing mechanisms with the 'power of two choices' using polling to avoid heavily loaded hosts, and implementing system-level changes. These system changes involved considering memory bandwidth as a resource during replica placement in Shard Manager and resolving expectation mismatches between ServiceRouter and Shard Manager regarding load balancing assumptions. These efforts resulted in a two-thirds reduction in timeout error rates, a 35% increase in work output for the same resources, and a halving of p99 latency.

Read the original post ↗