Deploy Private AI on
Your Own Infrastructure
Take control of your data and reduce latency by deploying open-source Large Language Models (LLMs) like LLaMA 3 and Mistral on rented VPS environments. Build robust Retrieval-Augmented Generation (RAG) pipelines using high-performance vector databases.
100% Data Privacy
Your prompts and proprietary documents never leave your server. Full compliance with data regulations.
High Performance
Optimized inference engines utilizing CPU/GPU for maximum tokens/second throughput.
Vector Native
Seamless integration with Chroma and Qdrant for semantic search and document retrieval.
Hardware Requirements
Choosing the right VPS hardware is critical for reasonable generation speeds. Models are heavy on RAM and memory bandwidth. Here is what you need depending on your target model:
| Model Size | Quantization | Min. VRAM/RAM | Recommended Server |
|---|---|---|---|
| 7B - 8B (Mistral, LLaMA 3) | 4-bit (Q4_K_M) | ~6 GB | 4 vCPU, 8GB RAM (CPU only) OR Nvidia T4 |
| 13B - 14B (Qwen) | 4-bit (Q4_K_M) | ~10 GB | 8 vCPU, 16GB RAM OR Nvidia A10G |
VPS Deployment
Deploying a robust AI inference node requires the right tooling. We recommend Ollama for managing weights, quantization, and executing models efficiently via a REST API on your server architecture.
Installing the Inference Engine
# Install Ollama on Linux/Ubuntu VPS curl -fsSL https://ollama.com/install.sh | sh # Start the background service to keep API alive systemctl start ollama systemctl enable ollama
Running LLaMA & Mistral
Once the engine is running, pull open-source models provided by Meta or Mistral AI. The system handles quantization and hardware offloading automatically.
# Pull and run Mistral 7B ollama run mistral # Pull Meta's LLaMA 3 (8B parameters) ollama run llama3
API Endpoints
REST API Generation Request
The local inference server exposes an API gateway on port 11434. You can interact with it via standard JSON payloads, making it easy to plug into web apps.
curl -X POST http://localhost:11434/api/generate -d '{ "model": "mistral", "prompt": "Explain the core concept of Vector Databases.", "stream": false }'
Vector DBs (RAG)
To build a Retrieval-Augmented Generation (RAG) system, you need to store document embeddings. We utilize ChromaDB running as a separate dockerized service. Python clients then bridge the Vector DB semantic search with the local LLM generation.
import chromadb import requests # 1. Connect to local ChromaDB instance client = chromadb.HttpClient(host='localhost', port=8000) collection = client.get_collection(name="company_docs") # 2. Perform semantic search (retrieve context) results = collection.query( query_texts=["How do I reset my password?"], n_results=1 ) context = results['documents'][0][0] # 3. Pass Context to Local LLM API payload = { "model": "mistral", "prompt": f"Context: {context}\n\nQuestion: How to reset password?", "stream": False } response = requests.post("http://localhost:11434/api/generate", json=payload)
Simulation Lab
Local Instance ActiveInteract with the simulated local model below. This demonstrates how a self-hosted inference node processes background requests and returns typed JSON responses to the frontend client. Try asking about "Hardware", "RAG", or "Deploy".
Local Inference Terminal
Model: mistral-7b-instruct-v0.2.Q4_K_M.gguf
System initialized. Model loaded into VRAM. Awaiting inference requests...