Explore Library
Quiz

Diagnosing a Skewed Hash Join at Scale

Identifying why a distributed hash join stalls on one node despite balanced table sizes.

A distributed hash join between a large 'orders' fact table and a 'customers' dimension runs mostly fast, but one reducer/partition takes 20x longer and nearly OOMs. Row counts per node are roughly equal, statistics are fresh, and both join keys are indexed. What is the MOST likely root cause and best fix?