Building a RAG Pipeline | Kamiwaza Docs
Documentation for Kamiwaza 0.7.1
Version: 0.7.1
Overview
Retrieval-Augmented Generation (RAG) combines the power of large language models with your own documents and data to provide accurate, contextual responses. This guide walks you through building a complete RAG pipeline using Kamiwaza's core services.
What You'll Build
By the end of this guide, you'll have:
- Document ingestion system that processes various file formats
- Embedding pipeline that converts text to vector representations
- Vector search system for finding relevant context
- LLM integration that generates responses using retrieved context
- Web interface for querying your documents
Prerequisites
Before starting, ensure you have:
- Kamiwaza installed and running
- At least 16GB of available RAM
- Sample documents (markdown format) to process
- Basic familiarity with Python (for SDK examples)
Architecture Overview
A RAG pipeline consists of four main components:
- Documents
- Ingestion & Chunking
- Embedding Generation
- Vector Storage
- Similarity Search
- LLM Generation
- Response
Step 1: Deploy Required Models
First, we'll deploy an embedding model for vectorizing text and a language model for generating responses.
Deploy an Embedding Model
The embedding model will be automatically loaded when you create an embedder - no manual deployment needed:
from kamiwaza_client import KamiwazaClient
client = KamiwazaClient(base_url="http://localhost:7777/api/")
embedder = client.embedding.get_embedder(
model="BAAI/bge-base-en-v1.5",
provider_type="huggingface_embedding"
)
print("✅ Embedding model ready for use")
Deploy a Language Model
Deploy a language model using Kamiwaza for response generation:
from kamiwaza_client import KamiwazaClient
client = KamiwazaClient(base_url="http://localhost:7777/api/")
model_repo = "Qwen/Qwen3-0.6B-GGUF"
models = client.models.search_models(model_repo, exact=True)
client.models.initiate_model_download(model_repo)
client.models.wait_for_download(model_repo)
deployment_id = client.serving.deploy_model(repo_id=model_repo)
openai_client = client.openai.get_client(repo_id=model_repo)
Check Deployment Status
# List active deployments to verify
deployments = client.serving.list_active_deployments()
for deployment in deployments:
print(f"✅ {deployment.m_name} is {deployment.status}")
Step 2: Document Ingestion Pipeline
Now we'll create a pipeline to process documents, chunk them, and generate embeddings.
Document Processing Script
import os
from pathlib import Path
from typing import List, Dict
from kamiwaza_client import KamiwazaClient
class RAGPipeline:
def __init__(self, base_url="http://localhost:7777/api/"):
self.client = KamiwazaClient(base_url=base_url)
self.embedder = self.client.embedding.get_embedder(
model="BAAI/bge-base-en-v1.5",
provider_type="huggingface_embedding"
)
def add_documents_to_catalog(self, filepaths: List[str]) -> List:
datasets = []
for filepath in filepaths:
dataset = self.client.catalog.create_dataset(
dataset_name=filepath,
platform="file",
environment="PROD",
description=f"RAG document: {Path(filepath).name}"
)
if dataset.urn:
datasets.append(dataset)
return datasets
def process_document(self, file_path: str):
doc_path = Path(file_path)
with open(doc_path, 'r', encoding='utf-8') as f:
content = f.read()
chunks = self.embedder.chunk_text(
text=content,
max_length=1024,
overlap=102
)
embeddings = self.embedder.embed_chunks(chunks)
# Handle metadata saving omitted for brevity
return len(chunks)
# Usage
pipeline = RAGPipeline()
DOCUMENT_PATHS = ["./docs/intro.md", "./docs/models/overview.md"]
datasets = pipeline.add_documents_to_catalog(DOCUMENT_PATHS)
for doc_path in DOCUMENT_PATHS:
chunks = pipeline.process_document(doc_path)
Step 3: Implement Retrieval and Generation
Now we'll create the query interface that retrieves relevant documents and generates responses.
from typing import List, Dict
from kamiwaza_client import KamiwazaClient
class RAGQuery:
def __init__(self, base_url="http://localhost:7777/api/"):
self.client = KamiwazaClient(base_url=base_url)
self.embedder = self.client.embedding.get_embedder(
model="BAAI/bge-base-en-v1.5"
)
def semantic_search(self, query: str, limit: int = 5) -> List[Dict]:
query_embedding = self.embedder.create_embedding(query).embedding
results = self.client.vectordb.search(
query_vector=query_embedding,
limit=limit
)
return results
def generate_response(self, query: str, context: str) -> str:
prompt = f"'''
Context:
{context}
Question: {query}
Answer:"""
response = self.openai_client.chat.completions.create(
messages=[{"role": "user", "content": prompt}]
)
return response.choices[0].message.content
def query(self, user_question: str, limit: int = 5) -> Dict:
search_results = self.semantic_search(user_question, limit=limit)
context = self.format_context(search_results)
response = self.generate_response(user_question, context)
return {'question': user_question, 'answer': response}
Step 4: Example Queries
rag = RAGQuery()
query_response = rag.query("What is one cool thing about Kamiwaza?")
print(query_response['answer'])
Step 5: Production Considerations
Resource Management
def monitor_system_health():
client = KamiwazaClient(base_url="http://localhost:7777/api/")
deployments = client.serving.list_active_deployments()
for deployment in deployments:
print(f" - {deployment.m_name}: {deployment.status}")
Clean up
def cleanup_rag_system(chat_model_repo="Qwen/Qwen3-0.6B-GGUF"):
client = KamiwazaClient(base_url="http://localhost:7777/api/")
success = client.serving.stop_deployment(repo_id=chat_model_repo)
Best Practices
- Chunk Size: Keep chunks between 200-800 tokens for optimal retrieval
- Overlap: Add 50-100 token overlap between chunks to preserve context
Troubleshooting
Common Issues
Poor Retrieval Quality
- Check embedding model performance on your domain
Next Steps
Now that you have a working RAG pipeline with Kamiwaza-deployed models:
- Try Different Models: Experiment with larger models for better response quality
Key Benefits of This SDK-Based Approach
- Simplified Integration: No need for manual HTTP requests - the SDK handles all API communication
- Production Ready: SDK handles error cases, retries, and connection management