2 min read

    Lab 08: Orchestration, CI/CD & DABs

    #databricks#workflows#cicd#dabs

    Before we write code, let's understand our Goal: Once your code works perfectly in a notebook, you cannot sit at your computer and click "Run" every morning at 2 AM. We need a system to run it automatically, retry if it fails, and deploy changes safely from development to production.

    The Tools:

    • Databricks Workflows (Jobs): The orchestrator that runs tasks on a schedule.
    • Databricks Asset Bundles (DABs): The deployment tool that turns your jobs into code (YAML) so they can be version-controlled in Git.

    1. Databricks Jobs & Workflows

    Real-World Analogy Mapping: Imagine a Restaurant Kitchen.

    • The Workflow (The Job): The master recipe for cooking a 3-course meal.
    • The Tasks: The individual steps. Task 1 (Make Salad) must happen before Task 2 (Cook Steak).
    • Repair Runs: If the chef burns the steak (Task 2 fails), you don't throw away the salad (Task 1). You use a Repair Run to only re-do the steak, saving time and ingredients (compute costs).

    Key Features of Workflows:

    • Parameters & Dynamic Values: Pass dynamic values into a job (e.g., {{start_date}}), which the notebook retrieves via dbutils.widgets.get("start_date").
    • Retries: Configure a task to retry 3 times upon failure with a 5-minute exponential backoff.
    • Concurrency: Restrict a job to max_concurrent_runs = 1 to ensure two schedules don't overlap and corrupt your Delta tables.

    2. Databricks Asset Bundles (DABs) & CI/CD

    What Existed Previously: Engineers used to manually open the Databricks UI, click "Create Job", and configure the schedules by hand. If someone accidentally deleted the job, it was gone forever. You couldn't track who made changes.

    How Present Technology Solves It: Instead of clicking buttons, engineers define their entire Workflow in a text file (YAML). This is called Infrastructure-as-Code. You push this file to Git, and a CI/CD pipeline (like GitHub Actions) automatically deploys it.

    Anatomy Breakdown of databricks.yml:

    yaml
    bundle:
      name: finance_pipeline # The name of the project
    
    resources:
      jobs:
        daily_etl:
          name: "Daily Finance ETL"
          tasks:
            # Task 1: Ingest Data
            - task_key: "ingest_data"
              notebook_task:
                notebook_path: "./src/ingestion.py"
              job_cluster_key: "default_cluster"
              
            # Task 2: Transform Data (Notice the dependency!)
            - task_key: "transform_data"
              depends_on:
                - task_key: "ingest_data"
              notebook_task:
                notebook_path: "./src/transformation.py"
    

    3. Deployment Commands (CLI)

    Once you write the YAML file, you use the Databricks Command Line Interface (CLI) to deploy it.

    The Smart Way (CLI Workflow):

    1. databricks bundle validate (Checks your YAML for typos or syntax errors before deploying).
    2. databricks bundle deploy -t dev (Zips up your code and deploys it to the Development workspace).
    3. databricks bundle run daily_etl -t dev (Manually triggers the job to test if it works).

    ← Previous: Lab 07: Unity Catalog & Governance SQL | Next: Lab 09: Testing & Monitoring →**