7 min read

    Big Data & Hadoop Fundamentals

    pysparkhadoopbig-data

    ๐Ÿ•ฐ๏ธ The Pre-Big Data Era

    What Existed Previously: Before Big Data, companies stored all their information in traditional Relational Databases (RDBMS) (like MySQL or Oracle) running on single, massive servers.

    Analogy

    Imagine trying to transport 1,000 tons of cargo. What existed before was like trying to build one single, impossibly gigantic truck to carry it all.

    ๐Ÿ›‘ Problems Faced

    • Expensive Scaling: If you needed more storage, you had to buy a bigger, more expensive server (vertical scaling). Eventually, you hit a physical limit.
    • Can't Handle Variety: RDBMS required data to be perfectly structured in tables (rows and columns). They completely failed at storing unstructured data like images, audio, or raw server logs.
    • Processing Bottlenecks: Searching through billions of rows on one machine took hours or days.

    ๐Ÿš€ Enter Big Data & Hadoop

    How Present Technology Solves It: Instead of building one giant truck, what if we used 1,000 normal pickup trucks driving together?

    That is Distributed Computing. Hadoop is the framework that allows hundreds of cheap, normal computers (commodity hardware) to work together as one giant machine (horizontal scaling).

    ๐Ÿ“ˆ What is Big Data? (The 3 Vs)

    • Volume: Massive scale (Terabytes to Petabytes).
    • Velocity: Speed of data generation (e.g., real-time social media streams).
    • Variety: Different types of data (structured SQL, raw text, videos).

    โฑ๏ธ Batch vs Real-Time Processing

    TypeWhat is it?Example
    Batch ProcessingProcessing a massive chunk of data all at once at a scheduled time.Generating monthly payroll reports.
    Real-Time ProcessingProcessing data instantly the moment it arrives.Credit card fraud detection at the swipe.

    ๐Ÿ˜ The Hadoop Ecosystem

    Hadoop solves the Big Data problem using three main tools working together:

    1๏ธโƒฃ HDFS (Hadoop Distributed File System) โ€” The Storage

    Instead of storing a 100GB file on one hard drive, HDFS chops it into tiny blocks and scatters them across 50 different computers. If one computer dies, HDFS automatically has a backup on another.

    2๏ธโƒฃ MapReduce โ€” The Processor

    Analogy

    If you need to count the words in a library of 10,000 books, you don't bring all the books to one person. You send 10,000 people to the books, have them count locally, and then sum the totals.

    Instead of moving massive data to the CPU, MapReduce brings the computation logic directly to the machines where the data is stored.

    3๏ธโƒฃ YARN (Yet Another Resource Negotiator) โ€” The Manager

    It acts as the operating system for the cluster, deciding which computers get to use their RAM and CPU for which tasks.


    ๐Ÿšš Data Ingestion: Sqoop

    What Existed Previously: Companies had decades of important data trapped in old relational databases.

    Problems Faced: Moving Terabytes of data manually into Hadoop is painfully slow and complex.

    How Present Technology Solves It: Apache Sqoop acts as a bridge, bulk-transferring data between traditional databases and Hadoop automatically.

    bash
    # Practical Example: Importing a MySQL table into Hadoop using Sqoop
    sqoop import \
      --connect jdbc:mysql://localhost/bankdb \
      --username admin --table transactions \
      --target-dir /user/hadoop/transactions_data
    

    ๐Ÿ—๏ธ Hadoop Advanced Features (Fault Tolerance)

    Real-World Analogy Mapping: Imagine a massive library where books are stored on different bookshelves (racks). If a bookshelf catches fire, you lose all the books on it. To prevent this, the librarian makes 3 copies of every book (Block Replication) and ensures they are never all placed on the exact same bookshelf (Rack Awareness).

    • Block Replication: HDFS automatically makes multiple copies (default 3) of every data block. If one computer (node) dies, the data is perfectly safe on another.
    • Rack Awareness: HDFS is smart enough to know which physical server racks the computers belong to. It intentionally places the backup copies on different physical racks, so even if an entire rack loses power, the data survives.

    โš™๏ธ Hadoop Deployment Modes

    When you install Hadoop, you can run it in three different modes depending on your goal.

    ModeWhat is it?Use Case
    Standalone ModeRuns everything in a single Java process. HDFS is disabled.Quick debugging and testing logic on your laptop.
    Pseudo-Distributed ModeRuns on one machine, but simulates a real cluster by spinning up separate processes for HDFS, YARN, etc.Advanced local testing and learning the ecosystem.
    Fully Distributed ModeSpreads the processes across hundreds or thousands of physical machines.Production. Real-world Big Data environments.

    ๐Ÿงช Practice Drill

    text
    // Try answering these:
    **Q1.** Why did traditional RDBMS fail to handle the "Variety" aspect of Big Data?
    
    **Q2.** Which component of Hadoop acts as the storage system that chops files into blocks?
    
    **Q3.** If a website needs to recommend a song to a user the moment they click "Like", do you need Batch or Real-Time processing?
    
    ๐Ÿ’ก Click for Solutions

    A1. RDBMS require data to fit perfectly into strict rows and columns. They cannot efficiently store or process unstructured data like images, audio, or raw server logs.

    A2. HDFS (Hadoop Distributed File System).

    A3. Real-Time processing, because the decision must be made instantly based on live user actions.


    โ† Syllabus | Next Topic โ†’ Apache Spark Fundamentals