17 min read

    14 - Hands-on Lab - High-Availability Web Cluster

    gcpcloudlabcompute-enginemigload-balancingvpc

    Welcome to Day 14 of Learn GCP in 30 Days! Today is your Week 2 Capstone Hands-on Lab.

    ๐ŸŽฏ

    Today's Goal Today, you will combine everything you learned in Week 2โ€”VPC Networking, Firewall Rules, Compute Engine Templates, Managed Instance Groups (MIGs), Autoscaling, and Cloud Load Balancingโ€”to build a real-world, production-grade High-Availability Web Cluster from scratch!


    ๐Ÿ—๏ธ The Problem: The Fragile Single Server

    Throughout Week 2, we discovered why single-server architectures fail in the real world:

    What Existed Previously:

    Traditional web applications ran on a single physical machine or virtual server with one public IP address.

    Problems Faced:

    • ๐Ÿ’ฅ Single Point of Failure (SPOF): If the server crashes or the data center experiences a power glitch at 3:00 AM, the entire website goes offline.
    • ๐ŸŒ Traffic Congestion & Freezing: When thousands of customers flood the website during a flash sale, the single server runs out of CPU and memory, crashing under the load.
    • ๐Ÿ”ง Zero Maintenance Window: Upgrading code or operating systems meant taking the website offline, losing business and customer trust.

    How Present Technology Solves It:

    By building a High-Availability Web Cluster on Google Cloud:

    1. Custom VPC Isolation: Your servers run inside a secure private network with precise firewall access.
    2. Golden Blueprint (Instance Template): Identical server clones can be launched automatically in seconds.
    3. Self-Healing & Elastic Scaling (MIG): Google Cloud automatically restarts dead servers and scales the number of running servers up or down based on incoming traffic.
    4. Unified Entry Point (Cloud Load Balancer): Millions of worldwide visitors connect to one single public Anycast IP, and traffic is distributed smoothly across all healthy server clones.
    mermaid

    ๐Ÿฌ Real-World Analogy: The International Airport Terminal

    Think of building a high-availability cloud architecture like organizing an International Airport Terminal:

    mermaid
    • ๐Ÿข Custom VPC & Subnet = The Private Airport Building: A secured physical perimeter keeping unauthorized outside traffic out.
    • ๐Ÿ›ก๏ธ Firewall Rules = Security Guards at the Entrance: Checking tickets and only allowing valid passengers through authorized doors (Port 80).
    • ๐Ÿ“‹ Instance Template = Standardized Officer Training Manual: Ensures every counter clerk runs the exact same setup and tools.
    • ๐Ÿ”„ Managed Instance Group = Duty Manager: Ensures at least 2 counters are always staffed. If an officer faints, the manager replaces them immediately.
    • โš–๏ธ Load Balancer = Queue Flow Coordinator: Directs incoming travelers to whichever counter is open and has the shortest line.

    ๐Ÿ—บ๏ธ Capstone Lab Architecture: What We Are Building

    In this lab, you will build and test a 5-component cloud infrastructure:

    mermaid

    ๐Ÿงช Step-by-Step Hands-on Lab Activity

    You can complete this capstone lab using either the Web Console UI (Click-by-Click) or the Cloud Shell CLI (Fast Track)!


    Step 1: Create a Custom VPC Network & Subnet

    First, we create a secure, custom network for our web cluster.

    Option A: Web Console UI

    1. Open console.cloud.google.com.
    2. Press /, type VPC networks, and select it (or navigate via โ˜ฐ โ†’\rightarrow VPC network โ†’\rightarrow VPC networks).
    3. Click Create VPC Network at the top.
    4. Set the following parameters:
      • Name: custom-web-vpc
      • Subnet creation mode: Custom
    5. Under New subnet, fill in:
      • Name: web-subnet-asia
      • Region: asia-south1 (or your nearest region)
      • IPv4 range: 10.0.1.0/24
    6. Leave other settings as default and click Create.

    Option B: Cloud Shell CLI

    bash
    # 1. Create the Custom VPC
    gcloud compute networks create custom-web-vpc --subnet-mode=custom
    
    # 2. Create the Subnet inside asia-south1
    gcloud compute networks subnets create web-subnet-asia \
        --network=custom-web-vpc \
        --region=asia-south1 \
        --range=10.0.1.0/24
    

    Step 2: Configure Firewall Rules (Traffic & Health Checks)

    Our web servers must accept web traffic on Port 80 from the public, and health probe queries from Google's load balancer monitoring system.


    ๐Ÿฉบ Why Allow Google Health Check Probe IPs (35.191.0.0/16 & 130.211.0.0/22)?

    Real-World Analogy Mapping:

    Think of your servers as a Restaurant Kitchen:

    • ๐Ÿฝ๏ธ Public Visitors (0.0.0.0/0) = Hungry customers coming through the front doors for food (web requests).
    • ๐Ÿฉบ Health Checkers (35.191.0.0/16) = The government Food Safety Inspector visiting the kitchen every 5 seconds to verify it is clean and operating.
    • ๐Ÿšช Firewall Rule = The Security Bouncer guarding the kitchen doors.

    If the bouncer only allows "customers" and blocks the "Health Inspector", the inspector can never enter the kitchen! The inspector will write a report: "I couldn't reach the kitchen, so the restaurant must be closed!" The Load Balancer will then stop sending all customers to your healthy servers.

    How Present Technology Solves It:

    Google's Load Balancer health checkers do not ping from the Load Balancer's public IP. Instead, Google uses dedicated, static probe IP ranges:

    • 35.191.0.0/16: Global and regional health checking probes.
    • 130.211.0.0/22: External HTTP(S) load balancer probes.

    By adding these ranges to your firewall rule, Google's monitoring robots can check your VMs and report them as Healthy!


    Option A: Web Console UI

    1. In the VPC network left menu, click Firewall.
    2. Click Create Firewall Rule:
      • Name: allow-web-traffic
      • Network: custom-web-vpc
      • Direction of traffic: Ingress (Inbound)
      • Action on match: Allow
      • Targets: Specified target tags
      • Target tags: web-cluster-node
      • Source IPv4 ranges: 0.0.0.0/0, 35.191.0.0/16, 130.211.0.0/22
      • Protocols and ports: Check Specified protocols and ports, select TCP, and enter 80.
    3. Click Create.

    Option B: Cloud Shell CLI

    bash
    # Create Firewall Rule for HTTP and Google Health Check Probes
    gcloud compute firewall-rules create allow-web-traffic \
        --network=custom-web-vpc \
        --direction=INGRESS \
        --priority=1000 \
        --action=ALLOW \
        --rules=tcp:80 \
        --source-ranges=0.0.0.0/0,35.191.0.0/16,130.211.0.0/22 \
        --target-tags=web-cluster-node
    

    Step 3: Create an Instance Template with a Dynamic Web Page

    We create the golden template that installs Nginx and creates a dynamic web page displaying the specific VM's name and zone when visited.


    ๐Ÿ›Ž๏ธ What is the GCP Metadata Server (http://metadata.google.internal/)?

    Real-World Analogy Mapping:

    Think of booting a VM like checking into a massive Hotel Room:

    • When you enter the room, you don't know the room number or building wing in advance.
    • You pick up the in-room telephone and dial the Hotel Reception Desk (#0): "Which room and wing am I in?"
    • The receptionist answers immediately: "You are in Room 402, Mumbai Branch."

    How Present Technology Solves It:

    When creating an Instance Template, you cannot hardcode VM hostnames or zones because Google will create multiple random clones (web-cluster-mig-4x9z) dynamically.

    Google runs an internal, private Metadata Server beside every VM at http://metadata.google.internal/computeMetadata/v1/ (or http://169.254.169.254/). Newly launched VMs query this internal endpoint to discover their own identity!

    Anatomy Breakdown:

    bash
    ZONE=$(curl -s -H "Metadata-Flavor: Google" http://metadata.google.internal/computeMetadata/v1/instance/zone | awk -F/ '{print $NF}')
    
    • curl -s: Silently fetches data via HTTP without printing progress bars.
    • -H "Metadata-Flavor: Google": A mandatory security badge header required by GCP to prevent malicious website scripts from stealing VM data.
    • http://metadata.google.internal/.../instance/zone: Returns the VM's zone string (e.g. projects/12345/zones/asia-south1-a).
    • awk -F/ '{print $NF}': Grabs only the last piece after the slash (asia-south1-a).

    Option A: Web Console UI

    1. Navigate via โ˜ฐ โ†’\rightarrow Compute Engine โ†’\rightarrow Instance templates.
    2. Click Create instance template.
    3. Configure the template:
      • Name: web-cluster-template
      • Machine type: e2-micro (Cost-effective for testing)
    4. Scroll down to Advanced options โ†’\rightarrow expand Networking:
      • Network: custom-web-vpc
      • Subnetwork: web-subnet-asia
      • Network tags: web-cluster-node
    5. Expand Management โ†’\rightarrow under Startup script, paste the following script:
      bash
      #!/bin/bash
      apt-get update
      apt-get install -y nginx
      HOSTNAME=$(hostname)
      ZONE=$(curl -s -H "Metadata-Flavor: Google" http://metadata.google.internal/computeMetadata/v1/instance/zone | awk -F/ '{print $NF}')
      cat <<EOF > /var/www/html/index.html
      <!DOCTYPE html>
      <html>
      <head>
        <title>GCP High-Availability Web Cluster</title>
        <style>
          body { font-family: 'Segoe UI', Tahoma, Geneva, Verdana, sans-serif; background: #0f172a; color: #f8fafc; text-align: center; padding: 50px; }
          .card { background: #1e293b; border-radius: 12px; padding: 30px; display: inline-block; box-shadow: 0 10px 25px rgba(0,0,0,0.5); }
          h1 { color: #38bdf8; }
          .badge { background: #0284c7; color: white; padding: 6px 14px; border-radius: 20px; font-weight: bold; }
        </style>
      </head>
      <body>
        <div class="card">
          <h1>๐Ÿš€ Week 2 Capstone Web Cluster</h1>
          <p>Successfully served by backend VM:</p>
          <h2><span class="badge">$HOSTNAME</span></h2>
          <p>Hosted in Google Cloud Zone: <strong>$ZONE</strong></p>
        </div>
      </body>
      </html>
      EOF
      systemctl restart nginx
      
    6. Click Create.

    Option B: Cloud Shell CLI

    ๐Ÿ’ก

    How Code Files are Created in this Lab (2 Ways)

    • โšก Fast 1-Click Way (Recommended): Simply copy & paste the cat << 'EOF' ... EOF command block below directly into your Cloud Shell terminal. It creates and saves startup.sh automatically in 1 second!
    • ๐Ÿ“ Manual Way (If you want to edit code):
      1. Open editor: nano startup.sh
      2. Paste code: Ctrl+V (or right-click โ†’\rightarrow Paste)
      3. Save: Ctrl+O then press Enter
      4. Exit: Ctrl+X
      5. (Or click the graphical Open Editor ๐Ÿ“ button in Cloud Shell).
    bash
    # 1. Create the startup script file
    cat << 'EOF' > startup.sh
    #!/bin/bash
    apt-get update
    apt-get install -y nginx
    HOSTNAME=$(hostname)
    ZONE=$(curl -s -H "Metadata-Flavor: Google" http://metadata.google.internal/computeMetadata/v1/instance/zone | awk -F/ '{print $NF}')
    cat << 'PAGE' > /var/www/html/index.html
    <!DOCTYPE html>
    <html>
    <head>
      <title>GCP High-Availability Web Cluster</title>
      <style>
        body { font-family: sans-serif; background: #0f172a; color: #f8fafc; text-align: center; padding: 50px; }
        .card { background: #1e293b; border-radius: 12px; padding: 30px; display: inline-block; box-shadow: 0 10px 25px rgba(0,0,0,0.5); }
        h1 { color: #38bdf8; }
        .badge { background: #0284c7; color: white; padding: 6px 14px; border-radius: 20px; font-weight: bold; }
      </style>
    </head>
    <body>
      <div class="card">
        <h1>๐Ÿš€ Week 2 Capstone Web Cluster</h1>
        <p>Successfully served by backend VM:</p>
        <h2><span class="badge">HOST_PLACEHOLDER</span></h2>
        <p>Hosted in Google Cloud Zone: <strong>ZONE_PLACEHOLDER</strong></p>
      </div>
    </body>
    </html>
    PAGE
    sed -i "s/HOST_PLACEHOLDER/$HOSTNAME/g" /var/www/html/index.html
    sed -i "s/ZONE_PLACEHOLDER/$ZONE/g" /var/www/html/index.html
    systemctl restart nginx
    EOF
    
    # 2. Create Instance Template using --metadata-from-file
    gcloud compute instance-templates create web-cluster-template \
        --region=asia-south1 \
        --network=custom-web-vpc \
        --subnet=web-subnet-asia \
        --machine-type=e2-micro \
        --tags=web-cluster-node \
        --metadata-from-file=startup-script=startup.sh
    

    Step 4: Create a Managed Instance Group (MIG) with Autoscaling

    Now, launch a Managed Instance Group that maintains a minimum of 2 VM clones across the zone and scales up to 4 if traffic increases.

    Option A: Web Console UI

    1. Navigate via โ˜ฐ โ†’\rightarrow Compute Engine โ†’\rightarrow Instance groups.
    2. Click Create instance group.
    3. Fill in the group configuration:
      • Name: web-cluster-mig
      • Instance template: Select web-cluster-template
      • Location: Single zone โ†’\rightarrow asia-south1-a (or your chosen zone)
    4. Under Autoscaling:
      • Autoscaling mode: On: permit Scale in and Scale out
      • Minimum number of instances: 2
      • Maximum number of instances: 4
      • Autoscaling signal: CPU utilization (Target: 60%)
    5. Click Create.

    Option B: Cloud Shell CLI

    bash
    # 1. Create Managed Instance Group
    gcloud compute instance-groups managed create web-cluster-mig \
        --zone=asia-south1-a \
        --template=web-cluster-template \
        --size=2
    
    # 2. Configure Autoscaling (Min: 2, Max: 4, Target CPU: 60%)
    gcloud compute instance-groups managed set-autoscaling web-cluster-mig \
        --zone=asia-south1-a \
        --min-num-replicas=2 \
        --max-num-replicas=4 \
        --target-cpu-utilization=0.60 \
        --cool-down-period=60
    

    Step 5: Deploy an External HTTP Load Balancer

    Now, place a Global External Application Load Balancer in front of web-cluster-mig.

    Option A: Web Console UI

    1. Navigate via โ˜ฐ โ†’\rightarrow Network services (or Networking) โ†’\rightarrow Load balancing.
    2. Click Create Load Balancer.
    3. Select Application Load Balancer (HTTP/S) โ†’\rightarrow Click Next.
    4. Choose Public facing (external) โ†’\rightarrow Best for global workloads โ†’\rightarrow Global external application load balancer โ†’\rightarrow Click Configure.
    5. Name: web-cluster-lb
    6. Frontend configuration:
      • Protocol: HTTP | Port: 80 | IP Version: IPv4 โ†’\rightarrow Click Done.
    7. Backend configuration:
      • Click Backend services & backend buckets โ†’\rightarrow Create a backend service.
      • Name: web-backend-svc
      • Backend type: Instance group
      • Under Backends, choose web-cluster-mig (Port: 80).
      • Under Health check, select Create a health check:
        • Name: web-http-health-check
        • Protocol: HTTP | Port: 80
        • Click Save.
      • Click Create.
    8. Click Review and create โ†’\rightarrow Click Create!

    Option B: Cloud Shell CLI

    bash
    # 1. Create Health Check
    gcloud compute health-checks create http web-http-health-check \
        --port=80 \
        --check-interval=5s \
        --healthy-threshold=2 \
        --unhealthy-threshold=2
    
    # 2. Create Global Backend Service
    gcloud compute backend-services create web-backend-svc \
        --protocol=HTTP \
        --health-checks=web-http-health-check \
        --global
    
    # 3. Add the MIG to the Backend Service
    gcloud compute backend-services add-backend web-backend-svc \
        --instance-group=web-cluster-mig \
        --instance-group-zone=asia-south1-a \
        --global
    
    # 4. Create URL Map, HTTP Proxy, and Global Forwarding Rule (Frontend)
    gcloud compute url-maps create web-cluster-url-map \
        --default-service=web-backend-svc
    
    gcloud compute target-http-proxies create web-cluster-http-proxy \
        --url-map=web-cluster-url-map
    
    gcloud compute forwarding-rules create web-cluster-frontend-rule \
        --global \
        --target-http-proxy=web-cluster-http-proxy \
        --ports=80
    

    Step 6: Live Verification & Testing Failover

    Let's test our live cluster to see traffic balancing and fault tolerance in action!

    1. Retrieve your Load Balancer's public IP address:

      bash
      gcloud compute forwarding-rules describe web-cluster-frontend-rule \
          --global \
          --format="value(IPAddress)"
      
    2. Open a web browser and visit http://YOUR_LOAD_BALANCER_IP. (Note: Global Load Balancers take about 2โ€“3 minutes to propagate DNS and warm up backend routes).

    3. Verify Load Balancing:

      • Refresh the page several times (or open it in an Incognito / Private window).
      • You will observe the displayed VM Hostname badge change between your 2 different VM instances (web-cluster-mig-xxxx and web-cluster-mig-yyyy)!
    4. Verify Auto-Healing & Fault Tolerance (Chaos Test):

      • Go to Compute Engine โ†’\rightarrow VM Instances.
      • Manually click on one of the running VMs in web-cluster-mig and click DELETE.
      • Refresh your browser: The Load Balancer instantly routes all traffic to the remaining healthy VM with zero downtime.
      • Within 60 seconds, check the VM Instances page again: The MIG automatically detects that an instance is missing and provisions a brand-new VM clone to keep the cluster at healthy capacity!

    ๐Ÿงน Step 7: Clean Up All Resources (Credit Safety Guarantee)

    To guarantee your trial account incurs $0.00 in continued charges, tear down all created lab resources in reverse order:

    Option A: Web Console UI

    1. Load Balancing: Go to Load balancing โ†’\rightarrow Select web-cluster-lb โ†’\rightarrow Click Delete (select all attached backend services, health checks, and forwarding rules).
    2. Instance Groups: Go to Compute Engine โ†’\rightarrow Instance groups โ†’\rightarrow Select web-cluster-mig โ†’\rightarrow Click Delete.
    3. Instance Templates: Go to Instance templates โ†’\rightarrow Select web-cluster-template โ†’\rightarrow Click Delete.
    4. Firewall Rules: Go to VPC network โ†’\rightarrow Firewall โ†’\rightarrow Delete allow-web-traffic.
    5. VPC & Subnet: Go to VPC networks โ†’\rightarrow Select custom-web-vpc โ†’\rightarrow Click Delete VPC network.

    Option B: Cloud Shell CLI (One-Click Cleanup)

    bash
    # 1. Delete Load Balancer Components
    gcloud compute forwarding-rules delete web-cluster-frontend-rule --global --quiet
    gcloud compute target-http-proxies delete web-cluster-http-proxy --quiet
    gcloud compute url-maps delete web-cluster-url-map --quiet
    gcloud compute backend-services delete web-backend-svc --global --quiet
    gcloud compute health-checks delete web-http-health-check --quiet
    
    # 2. Delete Managed Instance Group & Template
    gcloud compute instance-groups managed delete web-cluster-mig --zone=asia-south1-a --quiet
    gcloud compute instance-templates delete web-cluster-template --quiet
    
    # 3. Delete Firewall, Custom VPC Network & Temp Files
    gcloud compute firewall-rules delete allow-web-traffic --quiet
    gcloud compute networks subnets delete web-subnet-asia --region=asia-south1 --quiet
    gcloud compute networks delete custom-web-vpc --quiet
    rm -f startup.sh
    

    Signal vs. Noise: Key Concepts & Noise Filter

    ๐Ÿง 

    Good to Know (Key Concepts)

    • End-to-End Cluster Workflow: Custom VPC โ†’\rightarrow Firewall Rules โ†’\rightarrow Instance Template โ†’\rightarrow Managed Instance Group โ†’\rightarrow Cloud Load Balancer.
    • Health Check Probe IPs: Always allow 35.191.0.0/16 and 130.211.0.0/22 in firewall rules when using Google Cloud Load Balancers.
    • Metadata Server: http://metadata.google.internal/computeMetadata/v1/ allows startup scripts on VMs to dynamically query their own hostname, project ID, and zone.
    • Auto-Healing vs Autoscaling: Auto-healing replaces broken/unhealthy servers; autoscaling changes the total server count based on CPU or network load.
    โ„น๏ธ

    Noise Filter (Don't Memorize)

    • Do NOT memorize low-level Linux networking kernel parameters (sysctl) or manual IP routing tables.
    • Do NOT stress about the exact internal IP ranges used by Google's health-checker probes; you can always search the official GCP documentation for "load balancer health check probe IP ranges".

    Common Doubts & Interview Traps

    Q1: Why did my Load Balancer return a 502 Bad Gateway or Server Error during the first 2 minutes after creation?

    • Answer: Global Cloud Load Balancers take about 2 to 3 minutes to synchronize routing rules across Google's edge locations worldwide and complete their first set of health check probes against your backend instances. Once the health checks pass, traffic flows smoothly.

    Q2: What happens if I forget to allow Google's Health Check IP ranges (35.191.0.0/16 and 130.211.0.0/22) in my firewall rules?

    • Answer: The Load Balancer's health checker will be blocked from reaching your VMs. Even if your Nginx servers are running perfectly, the Load Balancer will mark every VM as Unhealthy and refuse to send any traffic to them!

    Q3: Can an Instance Template be edited once created?

    • Answer: No! Instance Templates in GCP are immutable (cannot be modified after creation). To change machine types or startup scripts, you create a new template (or copy the existing one) and perform a rolling update on the MIG.

    Daily Practice Drill & Self-Check

    Test your understanding of today's capstone lab:

    text
    // Try answering these:
    1. What are the two IP ranges that must be allowed through your firewall for Google Cloud Load Balancer health checks to succeed?
    2. If one VM in your Managed Instance Group is accidentally deleted by a team member, what component automatically detects the loss and creates a replacement?
    3. Which component of the load balancer holds the single public IP address that users type into their browsers?
    
    ๐Ÿ’ก Click for Solutions
    1. 35.191.0.0/16 and 130.211.0.0/22!
    2. The Managed Instance Group (MIG)! It compares the current instance count against the target size and launches a new VM from the Instance Template.
    3. The Frontend (Forwarding Rule)!

    ๐ŸŽ‰ Huge Congratulations! You have completed Day 14 and finished Week 2: Core Compute & Networking!

    Over the past 7 days, you have mastered:

    • VPC Networks, Custom Subnets, and CIDR ranges (Day 08)
    • Firewall Rules and Port Management (Day 09)
    • Compute Engine Virtual Machines & SSH access (Day 10)
    • Machine Types, Disk Configurations, and Spot VM savings (Day 11)
    • Managed Instance Groups, Auto-Healing, and Autoscaling (Day 12)
    • Global and Internal Load Balancing (Day 13)
    • Built a complete, production-ready High-Availability Web Cluster (Day 14)!

    Take some time to celebrate your progress. Tomorrow, we start Week 3: Serverless & Containerized Applications with Day 15: Containers 101 and Artifact Registry!


    โ† 13 - Global and Internal Load Balancing | Next Topic โ†’ 15 - Containers 101 and Artifact Registry