29 - Vertex AI Studio and Gemini API
Welcome to Day 29 of Learn GCP in 30 Days!
Yesterday, you mastered Cloud Security and CI/CD pipelines with Secret Manager and Cloud Build. Today, we step into the most transformative domain in modern cloud engineering: Generative AI with Google Cloud Vertex AI & Gemini.
Today's Goal Today, you will understand what Foundation Models actually are and why they revolutionized software development, learn how to control model outputs using Temperature and System Instructions, test prompts visually inside Vertex AI Studio, build a Python application that sends images and text to the Gemini API, and verify all parameters with our $0.00 credit safety guarantee!
ποΈ What is a Foundation Model? (The Architectural Shift)
To understand modern AI, we first need to understand the shift from Narrow AI to Foundation Models:
π§± Why is it called a "Foundation" Model?
Think of the concrete foundation of a building:
- A construction crew pours a single, solid concrete slab. On top of that same foundation, you can build a residential home, a hospital, a retail store, or a school.
- In AI, Google trains one massive, versatile base model (Gemini) on broad, internet-scale datasets (text, code, books, audio, images, and video).
- This single base model serves as the "Foundation" upon which developers can build thousands of different specialized applicationsβfrom medical assistants to code generatorsβsimply by giving it instructions, without ever training a machine learning model from scratch.
π Core AI Concepts Demystified
Here is the exact technical meaning of the core terms you will encounter in enterprise AI:
| Term | Exact Meaning | Why It Matters |
|---|---|---|
| Foundation Model | A massive, general-purpose base model pre-trained on broad multimodal data, designed to be adapted to a wide range of downstream tasks. | Eliminates the need to collect millions of training samples and build custom neural networks from scratch. |
| Prompt | The specific input text, question, or data payload sent to the model to guide its response. | The quality and clarity of your prompt directly determines the accuracy of the output. |
| System Instructions | Pre-conditioning context that sets the model's operational persona, behavioral constraints, and response formatting rules before it processes user input. | Enforces strict business guardrails (e.g. "You are a banking compliance assistant. Output only JSON. Never guess."). |
| Temperature | A mathematical sampling parameter () that controls the degree of randomness in token selection. | Low temperature () forces factual, deterministic output. High temperature () encourages creative, diverse phrasing. |
| Multimodality | The architectural ability of an AI model to simultaneously ingest, process, and reason across multiple data formats (text, code, images, audio, video) in a single unified model. | You can feed diagrams, video clips, or audio recordings directly to Gemini alongside text without needing separate conversion tools. |
| Vertex AI Studio | Google Cloud's interactive web console workspace for testing prompts, tuning model parameters, evaluating responses, and exporting production SDK code. | Allows engineers and architects to prototype and validate AI behavior before writing backend code. |
π‘οΈ Deep Dive: Understanding Temperature (Randomness Control)
When an AI generates a response, it predicts the next word (token) based on mathematical probability. Temperature controls how the model samples from these probabilities:
π The Core Problem: Consumer AI vs. Enterprise Vertex AI
If developers can already access consumer chatbots on the web, why do companies build on Google Cloud Vertex AI?
What Existed Previously:
When generative AI first emerged, enterprise security teams blocked employee access:
- Pasting proprietary financial records or private patient healthcare data into public consumer tools created severe compliance and data leak risks.
- Applications needed reliable programmatic APIs with guaranteed response formats, not unpredictable web interfaces.
How Google Cloud Vertex AI Solves It:
- Enterprise Data Isolation: Google explicitly guarantees that your prompts, embeddings, and uploaded files are 100% private to your project and never used to train base foundation models.
- Native Cloud Storage Integration: You can point Gemini directly to large PDFs, audio recordings, or high-resolution images stored in your private Cloud Storage (GCS) buckets via
gs://...URIs without uploading files over the public internet. - Structured Outputs: Gemini supports strict JSON Schema output enforcement, ensuring that the model's reply can be immediately ingested by downstream database tables and microservices.
β‘ The Modern Gemini Model Family
Google provides specialized foundation models engineered for different speed, cost, and reasoning requirements:
| Model Name | Key Strengths | Best Use Case | Context Window |
|---|---|---|---|
gemini-2.5-flash | Latest frontier multimodal speed, real-time vision, structured outputs | Modern AI web apps, live agents, vision analysis | 1,000,000 tokens |
gemini-1.5-flash | High-throughput efficiency, ultra-low cost ($0.075 / 1M input tokens) | High-volume classification, chatbots, document extraction | 1,000,000 tokens |
gemini-1.5-pro | Complex logical reasoning, cross-document synthesis, complex coding | Deep legal/financial analysis, massive codebase refactoring | 2,000,000 tokens |
π οΈ Step-by-Step Hands-On Lab: Build 2 Real-World AI Apps
In this hands-on lab, we will build two enterprise-grade AI applications using Google's latest Gemini models:
- App 1: Multimodal Vision Analyzer β Gemini "looks" at a cloud architecture diagram image stored in GCS and evaluates its design.
- App 2: Structured JSON Data Extractor β Gemini extracts raw text into a strict, validated JSON object ready for a database!
Step 1: Visual Prototyping in Vertex AI Studio
- In the Google Cloud Console search bar, search for Vertex AI Studio and select it.
- Click Generate with Gemini Freeform.
- In the Model dropdown, select
gemini-2.0-flash(orgemini-1.5-flash). - In the System Instructions field on the left configuration panel, paste:
text
You are an expert Google Cloud Solutions Architect. Provide structured technical assessments with clear section headers. - Set the Temperature slider to
0.2(for focused, factual reasoning). - In the user prompt box, enter:
text
Compare Google Cloud Run vs Google Kubernetes Engine (GKE) in 3 key technical criteria. - Click Submit Observe how Gemini produces a disciplined, well-structured comparison adhering to your system rules.
- Click the Get Code button at the top-right toolbar to view the automatically generated Python SDK boilerplate!
Step 2: Enable Vertex AI API & Set Up Environment
Open Google Cloud Shell and run:
# 1. Export project environment variables
export PROJECT_ID=$(gcloud config get-value project)
export REGION="us-central1"
# 2. Enable Vertex AI and Cloud Storage APIs
gcloud services enable \
aiplatform.googleapis.com \
storage.googleapis.com
π Under the Hood:
aiplatform.googleapis.com: Activates the Google Cloud Vertex AI API endpoint for programmatic inference requests.
Step 3: Stage a Sample Diagram in Cloud Storage
Let's create a Cloud Storage bucket and store an architecture diagram image for Gemini to analyze:
# 1. Create a dedicated storage bucket
export BUCKET_NAME="ai-demo-${PROJECT_ID}"
gcloud storage buckets create gs://${BUCKET_NAME} --location=${REGION}
# 2. Download a sample Google Cloud architecture diagram
curl -s -o sample_diagram.png https://cloud.google.com/static/images/architecture/serverless-three-tier-web-app.png
# 3. Copy the image into your private bucket
gcloud storage cp sample_diagram.png gs://${BUCKET_NAME}/sample_diagram.png
Step 4: Author App 1 β Multimodal Vision Analysis (vision_analyzer.py)
Let's create our application workspace and author the multimodal vision script using the latest Gemini 2.5 Flash model:
mkdir -p ~/vertex-ai-app
cd ~/vertex-ai-app
Terminal File Creation Options Choose either the 1-Click command or the manual editor:
β‘ Option A: Fast 1-Click Way (Copy & Paste):
cat << 'EOF' > vision_analyzer.py
import os
import sys
import json
import urllib.request
import subprocess
# 1. Initialize Project & Access Tokens
PROJECT_ID = subprocess.check_output("gcloud config get-value project", shell=True).decode().strip()
ACCESS_TOKEN = subprocess.check_output("gcloud auth print-access-token", shell=True).decode().strip()
BUCKET_NAME = os.environ.get("BUCKET_NAME", f"ai-demo-{PROJECT_ID}")
# 2. Determine target image from GCS
if len(sys.argv) > 1:
filename = sys.argv[1]
image_uri = f"gs://{BUCKET_NAME}/{filename}" if not filename.startswith("gs://") else filename
else:
output = subprocess.check_output(f"gcloud storage ls gs://{BUCKET_NAME}/", shell=True).decode()
images = [line.strip() for line in output.splitlines() if any(line.lower().endswith(ext) for ext in ['.jpg', '.jpeg', '.png', '.webp'])]
user_images = [img for img in images if 'sample_diagram' not in img]
image_uri = user_images[-1] if user_images else images[-1]
mime_type = "image/jpeg" if image_uri.lower().endswith(('.jpg', '.jpeg')) else "image/png"
print(f"\nπΈ Selected Image: {image_uri}")
print("π€ Model: Gemini 2.5 Flash")
print("π Analyzing person, scene, and architecture...\n")
# 3. Prompt for Person, Scene & Technical Analysis
prompt = """
Analyze this image in detail:
1. If there is a person/people: Identify if it is a well-known public figure, celebrity, or historical personality. Describe their appearance, clothing, posture, and facial expression.
2. If this is a diagram: Identify the Google Cloud services and request flow depicted.
3. Describe the overall setting, background, and 3 key notable visual details.
"""
payload = {
"contents": [
{
"role": "user",
"parts": [
{
"fileData": {
"mimeType": mime_type,
"fileUri": image_uri
}
},
{
"text": prompt
}
]
}
],
"generationConfig": {
"temperature": 0.2,
"maxOutputTokens": 1024
}
}
# 4. Call Vertex AI Endpoint directly with Gemini 2.5 Flash
url = f"https://us-central1-aiplatform.googleapis.com/v1/projects/{PROJECT_ID}/locations/us-central1/publishers/google/models/gemini-2.5-flash:generateContent"
req = urllib.request.Request(
url,
data=json.dumps(payload).encode("utf-8"),
headers={
"Authorization": f"Bearer {ACCESS_TOKEN}",
"Content-Type": "application/json"
},
method="POST"
)
with urllib.request.urlopen(req) as resp:
data = json.loads(resp.read().decode("utf-8"))
text = data["candidates"][0]["content"]["parts"][0]["text"]
print("=" * 60)
print("π€ GEMINI 2.5 FLASH ANALYSIS RESULT:")
print("=" * 60)
print(text)
print("=" * 60 + "\n")
EOF
π Option B: Manual Way with Nano:
nano vision_analyzer.py
(Paste the code above, press Ctrl+O Enter to save, then Ctrl+X to exit).
Step 5: Author App 2 β Structured JSON Extraction (json_extractor.py)
Now let's build an app that forces Gemini 2.5 Flash to return 100% valid JSON:
β‘ Fast 1-Click Way (Copy & Paste):
cat << 'EOF' > json_extractor.py
import os
import json
import urllib.request
import subprocess
PROJECT_ID = subprocess.check_output("gcloud config get-value project", shell=True).decode().strip()
ACCESS_TOKEN = subprocess.check_output("gcloud auth print-access-token", shell=True).decode().strip()
raw_review = """
Customer Feedback (2026-08-31):
Hi, I ordered 3 pairs of CloudRunner Shoes (Order #94812) on Friday.
The delivery was lightning fast (2 days!), but the shoe size was slightly smaller than expected.
Rating: 4 out of 5 stars. Customer sentiment: positive but requested an exchange.
"""
prompt = f"""
You are a pure data extraction engine. Extract the following information into a strict JSON object:
- order_id (integer)
- product_name (string)
- quantity (integer)
- rating (integer out of 5)
- sentiment (POSITIVE, NEUTRAL, or NEGATIVE)
- action_required (string)
Return ONLY pure valid JSON, without backticks or markdown formatting.
Text to analyze:
{raw_review}
"""
payload = {
"contents": [{"role": "user", "parts": [{"text": prompt}]}],
"generationConfig": {
"temperature": 0.0,
"responseMimeType": "application/json"
}
}
url = f"https://us-central1-aiplatform.googleapis.com/v1/projects/{PROJECT_ID}/locations/us-central1/publishers/google/models/gemini-2.5-flash:generateContent"
req = urllib.request.Request(
url,
data=json.dumps(payload).encode("utf-8"),
headers={
"Authorization": f"Bearer {ACCESS_TOKEN}",
"Content-Type": "application/json"
},
method="POST"
)
with urllib.request.urlopen(req) as resp:
data = json.loads(resp.read().decode("utf-8"))
text = data["candidates"][0]["content"]["parts"][0]["text"]
# Parse and format verified JSON dictionary
parsed_json = json.loads(text)
print("\n--- π Extracted Database-Ready JSON ---")
print(json.dumps(parsed_json, indent=2))
print("---------------------------------------\n")
EOF
Step 6: Execute Both Applications!
Run both Python applications in Cloud Shell:
export PROJECT_ID=$(gcloud config get-value project)
export BUCKET_NAME="ai-demo-${PROJECT_ID}"
# 1. Run Multimodal Image Analysis (with Gemini 2.5 Flash)
python3 vision_analyzer.py
# 2. Run Structured JSON Extraction (with Gemini 2.5 Flash)
python3 json_extractor.py
π What You Will See:
vision_analyzer.py: Gemini 2.5 Flash reads the image from Cloud Storage and performs immediate visual / person / architecture recognition.json_extractor.py: Gemini 2.5 Flash extracts customer review data into a clean, verified JSON dictionary ready for Firestore or BigQuery!
π‘οΈ Step 7: Credit Safety & Resource Teardown ($0.00 Guarantee)
Gemini API calls are billed per token (Gemini 1.5 Flash costs less than $0.0001 for this test). To keep your Google Cloud account completely clean:
Clean Up Demo Resources
# Delete Cloud Storage bucket and image
gcloud storage rm --recursive gs://${BUCKET_NAME}
# Delete local workspace files
rm -rf ~/vertex-ai-app sample_diagram.png
π§ Daily Practice Drill & Self-Check
Test your understanding of enterprise AI concepts:
// Try answering these:
1. What is the defining characteristic of a "Foundation Model" compared to traditional task-specific AI?A) It can only run on quantum computersB) It is a single massive general-purpose base model pre-trained on broad data that can be adapted to hundreds of different tasks via promptingC) It only works for processing numbers in spreadsheetsD) It requires manually labeling 1 million training images for every new question
2. Why would an engineer set the temperature parameter to 0.0 or 0.1 when using Gemini in a software application?A) To cool down the physical server CPUsB) To enforce strict, deterministic, and factual outputs with minimal randomness (ideal for code generation, data extraction, and math)C) To make the model generate fictional storiesD) To delete the model cache
3. What is the primary data privacy benefit of using Google Cloud Vertex AI over free consumer chatbot tools?A) Vertex AI has a dark mode themeB) Vertex AI guarantees enterprise data privacy: your prompts, customer data, and documents are never used to train Google's public foundation modelsC) Consumer chatbots only work during business hoursD) Vertex AI disables all user passwords
4. What does "Multimodality" mean in the Gemini model architecture?A) The model requires 3 separate graphics cards to bootB) The model natively ingests, understands, and reasons across multiple data formats (text, code, images, audio, video) in a single unified systemC) The model only understands English and SpanishD) The model can only generate JPEG images
π‘ Click for Solutions
- B (Massive general-purpose base model) β Foundation models serve as a reusable core base that can perform translation, summarization, coding, and analysis without task-specific retraining.
- B (Deterministic & factual output) β Low temperature minimizes token sampling randomness, ensuring repeatable, accurate technical outputs.
- B (Enterprise privacy guarantee) β Vertex AI enforces strict enterprise data governance, ensuring private data remains completely isolated and protected by GCP IAM.
- B (Native Multimodal reasoning) β Gemini was built from the ground up to process text, audio, video, code, and images simultaneously within a single context window.
π Day 29 Cheat Sheet Summary
Tomorrow is our grand celebration: Day 30: Final Capstone Project & Course Graduation! We will bring together everything you learned across 30 days into a complete production system! ππ
β 28 - Secret Manager & Cloud Build CI-CD | Next Topic β 30 - Capstone Project - Production Deployment