Flashcard
Shuffles and Joins in Distributed Feature Engineering
How partitioning, shuffles, and join strategies drive the cost of scaling feature engineering in Spark.
Question
Why are shuffles the main bottleneck when scaling feature engineering (groupBy, joins) in Spark?
Click to reveal answer