Hot shards
Detecting and splitting load when one partition owns the storm.
Hot shards
Business constraint
A celebrity user or SKU concentrates traffic on one shard. Average utilization looks fine; one shard melts.
Naive v0
More replicas of the same shard keying. Or ignore per-shard metrics.
Failure drill
p99 latency from one shard. Retries amplify. Rebalancing moves the hotspot without changing the key.
Evolution path
Iteration 1
Per-shard metrics and alerts. Detect skewed keys.
Iteration 2
Split hot keys (salting, subshards) or cache with single-flight in front.
Iteration 3
Consistent hashing with bounded load; planned reshard drills.
Implementation cut
Average CPU is a vanity metric. Practice key salting carefully with read path fan-in.
Numbers
Max/mean shard QPS ratio. Hot key share of traffic. Reshard impact window.
Pattern tags
Self-check
- Why can fleet averages lie?
- What is key salting?
- How does caching help hot keys?
- What does bounded-load hashing change?
- What do you drill before a viral event?
Walkthrough
Recordings will appear here when published. Until then, work the failure drill and lab locally — pause after each iteration and write the metric that moved.