2 min read

    Lab 13: Generative AI and Vector Search

    #databricks#ai#rag#vector-search

    Before we write code, let's understand our Goal: You want to build a ChatGPT-style bot that can answer questions about your company's private HR Handbook. Since ChatGPT has never read your private handbook, if you ask it a question, it will hallucinate and make up an answer.

    The Tool: Retrieval-Augmented Generation (RAG) combined with a Vector Database.


    1. Vector Search (The Open Book Test)

    Real-World Analogy Mapping: Imagine asking a student a highly specific history question.

    • The Long Way (Standard LLM): The student tries to guess the answer from memory. They get it wrong. (Hallucination).
    • The Smart Way (RAG): You give the student an open textbook, a highlighter, and 5 minutes to find the paragraph. They read the exact paragraph and give you a 100% accurate answer.

    A Vector Database mathematically groups sentences by "meaning." When a user asks "Maternity leave?", the Vector Database instantly finds the paragraph about "Parental Time Off" (even though the keywords don't match perfectly) and hands that paragraph to the AI.

    python
    from databricks.vector_search.client import VectorSearchClient
    
    vsc = VectorSearchClient()
    
    # The Smart Way: Create a Vector Index directly from a Delta Table
    # Databricks automatically converts the text into mathematical embeddings using a built-in AI model!
    vsc.create_delta_sync_index(
        endpoint_name="enterprise_vector_endpoint",
        source_table_name="enterprise_data.hr.employee_handbooks",
        index_name="enterprise_data.hr.handbook_index",
        pipeline_type="TRIGGERED",
        primary_key="document_id",
        embedding_source_column="content",
        embedding_model_endpoint_name="databricks-bge-large-en" 
    )
    

    2. LangChain (Building the AI Agent)

    Now we use Python's LangChain library to connect our Vector Database (the textbook) to a Large Language Model (the student).

    python
    from langchain_community.chat_models import ChatDatabricks
    from langchain_community.vectorstores import DatabricksVectorSearch
    from langchain.chains import RetrievalQA
    
    # 1. The Student (A massive AI model hosted securely inside Databricks)
    llm = ChatDatabricks(endpoint="databricks-dbrx-instruct")
    
    # 2. The Textbook (Our Vector Index)
    vector_store = DatabricksVectorSearch(
        endpoint="enterprise_vector_endpoint",
        index_name="enterprise_data.hr.handbook_index"
    )
    # Tell the database to retrieve the top 3 most relevant paragraphs
    retriever = vector_store.as_retriever(search_kwargs={"k": 3})
    
    # 3. Connect them together into an Agent!
    rag_chain = RetrievalQA.from_chain_type(
        llm=llm,
        chain_type="stuff",
        retriever=retriever
    )
    
    # 4. Ask the Agent a question. It will read the 3 paragraphs and summarize the answer.
    answer = rag_chain.run("How many weeks of maternity leave do I get?")
    print(answer)
    

    ← Previous: Lab 12: Legacy to Lakehouse Migration Playbook | Next: Lab 14: Infrastructure as Code with Terraform →**