ยท3 min read
๐ PySpark & Big Data Engineering โ Syllabus
pysparkdata-engineeringsyllabus
๐ PySpark & Big Data Engineering โ Syllabus
Course Overview
This module covers the transition from legacy Hadoop architectures to modern, in-memory Apache Spark processing, ending with cutting-edge Lakehouse architectures.
๐ก All notes are designed for beginners using simple analogies, real-world examples, and the Problem โ Solution framework.
โ Chapters
๐ 01 - Big Data & Hadoop Fundamentals
- Characteristics of Big Data (Volume, Velocity, Variety)
- Batch vs Real-Time Processing
- Hadoop Ecosystem & MapReduce limitations
- Data Ingestion with Sqoop
โก 02 - Apache Spark Fundamentals
- Spark Architecture (Driver, Executors, Cluster Manager)
- In-Memory Computing Advantages
- Spark Deployment Modes (Local, Client, Cluster)
๐งฉ 03 - Spark RDDs (Resilient Distributed Datasets)
- What is an RDD? (Immutability, Lineage, Fault Tolerance)
- Transformations (Lazy Evaluation) vs. Actions
- Narrow vs. Wide Dependencies
04 - Advanced RDD Concepts
- RDD Persistence Levels (Cache vs. Persist)
- Partitioning for Parallelization (
repartitionvscoalesce) - Shared Variables (Broadcast, Accumulators)
๐๏ธ 05 - Spark SQL & DataFrames
- The Catalyst Optimizer & Tungsten Engine
- Data Abstractions (DataFrames vs Datasets vs Schema RDDs)
- File Formats (JSON vs Parquet)
- Python UDFs & Performance Hits
๐ 06 - Advanced SQL & Transformations
- Advanced Grouping (ROLLUP, CUBE)
- Common Table Expressions (CTEs) for readable code
- Pivoting Data (Long to Wide)
- Materialized Views
๐ง 07 - Modern Data Architectures (Delta Lake & Iceberg)
- Problems with traditional Data Lakes
- Delta Lake Fundamentals (ACID, Time Travel, Schema Evolution)
- Apache Iceberg features and massive scale
Course Content: