๐ Environment Setup & Workspace Provisioning
Environment Setup & Workspace Provisioning
Before we dive into Data Ingestion and Delta Lake, you need an actual environment to run the code provided in the upcoming chapters and the Hands-On Implementation Labs.
You can view the complete lab roadmap, topics, and learning sequence here:
๐ View the Databricks Hands-On Labs Syllabus
What Existed Previously: Historically, learning Big Data meant spending 3 days installing Java, Scala, Apache Hadoop, and Spark on your local laptop. You had to manually configure hundreds of environment variables, and if your Windows/Mac version didn't perfectly match the Java version, nothing worked.
Problems Faced:
- Installation Hell: Beginners gave up before even writing their first line of code due to configuration errors.
- Resource Limits: Your local laptop only has 16GB of RAM; it crashes when trying to process real big data.
How Present Technology Solves It: Databricks is entirely cloud-native. You do not install anything on your local computer. With a single click, Databricks provisions a complete, pre-configured Spark cluster in the cloud, accessible directly through your web browser.
This chapter will guide you through setting up this "Cold Start" environment from scratch.
1. Option A: Databricks Community Edition (Free & Recommended for Beginners)
If you are a student or just want to practice PySpark and Delta Lake without attaching a credit card, the Databricks Community Edition is the perfect starting point.
What it includes:
- A free micro-cluster (1 Driver node, 15GB memory, 2 Cores).
- Databricks Notebooks environment.
- Basic Delta Lake capabilities.
Limitations:
- No Auto Loader (
cloudFiles). - No Unity Catalog or Databricks SQL.
- Clusters automatically terminate after 2 hours of idle time (and data in temporary DBFS may be wiped).
Setup Steps:
- Navigate to the Databricks Sign Up Page.
- Fill out your details. When prompted to choose a cloud provider, look below the large cloud provider buttons for a smaller link that says "Get started with Community Edition". Click that link.
- Verify your email.
- Log into your free workspace!
2. Option B: Enterprise Cloud Trial (AWS / Azure / GCP)
If you want to practice production-grade features like Unity Catalog, Delta Live Tables (DLT), Serverless SQL, and Auto Loader, you must use a full enterprise workspace.
Databricks offers a 14-day free trial of the platform itself, but you must pay your cloud provider (AWS/Azure) for the actual compute (EC2/VMs) used during the trial.
Azure Databricks Setup:
- Log into your Azure Portal.
- Search for "Azure Databricks" and click Create.
- Select your Subscription and Resource Group.
- Enter a Workspace Name and choose a Region.
- Under Pricing Tier, select Premium (Required for Unity Catalog and Serverless).
- Click Review + Create. Once deployed, click Launch Workspace.
AWS Databricks Setup:
- Navigate to the Databricks Free Trial page and select AWS.
- Follow the Quickstart guide, which provides a CloudFormation Template.
- The CloudFormation template will automatically create the necessary Cross-Account IAM Roles (so the Databricks Control Plane can securely spin up EC2 instances in your AWS account).
3. Creating Your First Compute Cluster
Once you are logged into your Workspace (Community or Enterprise), you need "compute" to run your notebooks.
- On the left sidebar, click Compute, then click Create Compute.
- Cluster Name:
Course_Lab_Cluster. - Databricks Runtime Version: Select the latest LTS (Long Term Support) version (e.g.,
13.3 LTSor14.3 LTS). If you plan to do Machine Learning, select theMLversion. - Node Type:
- Community Edition: You cannot change this.
- Enterprise: Choose a small instance (e.g.,
i3.xlargeon AWS orStandard_DS3_v2on Azure) to save costs.
- Autopilot Options:
- Uncheck "Enable autoscaling" to keep costs predictable.
- CRITICAL: Set "Terminate after X minutes of inactivity" to
15or30minutes to prevent runaway cloud bills.
- Click Create Cluster. (It will take 2-5 minutes to spin up the virtual machines).
4. Importing the Course Labs
To run the PySpark code found in the Implementation Labs chapter, you can copy the code directly into a Databricks Notebook.
- On the left sidebar, click Workspace, then click your username.
- Right-click the whitespace and select Create -> Notebook.
- Name it
Lab_01_PySpark. - Ensure the Default Language is set to Python and attach it to the
Course_Lab_Clusteryou just created. - You can now copy and paste the code from the upcoming chapters directly into the notebook cells and press
Shift + Enterto execute!
โ ๐ The Databricks Data Intelligence Platform | Next Topic โ ๐ฅ Source Systems & The Ingestion Layer