Lab 11: Deep Spark UI & Memory Tuning
Before we write code, let's understand our Goal: You wrote a query that takes 4 hours to run. Your boss tells you to make it run in 5 minutes. You need to look inside the engine to figure out what is broken.
The Tool: The Spark UI is a visual dashboard that shows exactly what every computer in your cluster is doing at any given millisecond.
1. Diagnosing Data Skew (The Traffic Jam)
Real-World Analogy Mapping: Imagine a 4-lane highway (Your Spark Cluster has 4 CPUs).
- The Ideal Scenario: 400 cars are split perfectly: 100 cars per lane. Traffic flows at 60mph.
- Data Skew: 397 cars try to cram into the far-left lane because they are all going to the same exit. The other 3 lanes are completely empty. The left lane comes to a dead halt.
In Spark, this happens when you GROUP BY a column that has heavily skewed data (e.g., grouping by Country when 99% of your customers live in the "USA"). One computer does 99% of the work while the others sit idle.
How Present Technology Solves It: Look at the Spark UI's Tasks tab. If one Task takes 45 minutes, but the median Task takes 2 seconds, you have Data Skew. To fix it, we "salt" the data (add a random number to the key to force the cars into different lanes), or we use Databricks' built-in Adaptive Query Execution (AQE).
# The Smart Way: Enable AQE Skew Join Optimization
spark.conf.set("spark.sql.adaptive.enabled", "true")
spark.conf.set("spark.sql.adaptive.skewJoin.enabled", "true")
2. Memory Tuning (Out of Memory Errors)
Anatomy Breakdown: When a worker node has 16GB of RAM, Spark does not give you all 16GB.
- Reserved Memory: Spark keeps ~300MB for the system.
- Execution Memory: Used for math (sorting, shuffling, joining).
- Storage Memory: Used for caching data (
df.cache()).
If you cache a massive dataset, you steal memory away from the Execution engine. When the Execution engine runs out of space, it crashes with an Out Of Memory (OOM) error, or it "spills" the math onto the slow hard drive.
# The Long Way: Guessing and getting OOM errors.
# The Smart Way: Tuning the memory fraction
# Tell Spark to dedicate 80% of the RAM to Execution/Storage, and 20% to User Memory.
spark.conf.set("spark.memory.fraction", "0.8")
# If you need to cache data, use DISK to save RAM
from pyspark import StorageLevel
df.persist(StorageLevel.DISK_ONLY)
3. The Garbage Collector (G1GC)
Goal: Clean up old variables that are no longer being used so the RAM doesn't fill up.
If your Spark cluster randomly freezes for 30 seconds and does zero work, the Java Garbage Collector is likely pausing the world to clean up memory. You can fix this by adding a special configuration to your cluster creation script to use the modern "G1" garbage collector.
# Add this to the "Spark Config" section of your Databricks Cluster:
spark.executor.extraJavaOptions -XX:+UseG1GC
← Previous: Lab 10: MLflow & MLOps Expansion | Next: Lab 12: Legacy to Lakehouse Migration Playbook →**