Local LLM Deployment with Ollama and vLLM is a hands-on course in running open-weight models on your own machines, from choosing models, quantisation and GPU sizing to serving them through Ollama and vLLM with an OpenAI-compatible API and Open WebUI. It suits IT, DevOps and developer teams whose data must stay in-house, and you leave with a working prototype of an internal LLM service.
Many organisations want to use LLMs on internal documents, customer data or source code, but the rules say that data must not leave for an outside provider. The answer is to run open-weight models on the organisation's own machines or servers. Once a team actually tries it, though, the IT side faces hard questions: which model to pick, how many GB of GPU memory are needed, how far to quantise before quality drops, whether to use Ollama or vLLM, and how to open the service to the whole organisation safely.
This course takes learners through setting up an in-house LLM service from scratch to real use. It starts with choosing models and reading licence terms, understanding quantisation and the GGUF format, and calculating the VRAM, RAM and context length a workload needs. Learners then go deep on Ollama, covering Modelfiles, concurrency settings and the OpenAI-compatible API with structured output and tool calling, before moving
to vLLM on a GPU to serve many users at once. They tune memory, measure performance with benchmarks, add a chat front end with Open WebUI, and put security and monitoring in place. The course closes with a capstone in which each learner designs and deploys an in-house LLM system of their own. (2 days, 6 hours per day, 12 hours in total, Intermediate level.)
What you’ll gain
Choose open-weight models that fit the task, the Thai language, the hardware and the licence terms
Understand quantisation and pick a precision level that balances quality against resources
Calculate the VRAM and RAM needed from model size, context length and concurrent users
Install and configure Ollama with Modelfiles and environment variables for real workloads
Call models through an OpenAI-compatible API with structured output, tool calling and embeddings
Run and tune vLLM on a GPU to handle a large number of concurrent requests
Build an in-house LLM service with a chat interface, authentication and monitoring
Who this course is for
IT infrastructure teams and system engineers asked to set up an LLM for internal use
DevOps and platform engineers who run AI services on the organisation's servers or cloud
Developers who need an LLM API for internal apps that keeps data inside the organisation
AI engineers comparing the cost and performance of cloud APIs with self-hosting
Organisations whose confidentiality or PDPA requirements limit the use of cloud AI
Prerequisites
Basic Linux command-line and Docker skills
Basic Python and experience calling REST APIs
Some experience with ChatGPT or another LLM and a basic idea of prompting
A laptop with at least 16 GB of RAM that can run Docker, and a Google account for Google Colab
Curriculum
Course Details
This course runs for 2 days, 6 hours per day (12 hours in total, 09:00-16:00), as lectures with a lab in every topic. Intermediate level. Day one covers model selection, resource sizing and Ollama, with labs on the learner's own laptop using small models. Day two covers vLLM on a GPU, Open WebUI, security and monitoring, using Google Colab or a cloud GPU rented by the learner, whose cost is the learner's responsibility. The course focuses on deploying and serving models, and pairs well with the On-Premise AI Chatbots with LangChain and LangGraph course, which focuses on building chatbot applications on top of in-house models. Learners take home a lab guide, docker-compose files, a GPU sizing sheet and benchmark scripts.
Day 1Choosing Models, Sizing Resources and Ollama in Depth
Section 1: Why Run LLMs Yourself, and the Tool Landscape
Why organisations self-host: confidentiality, PDPA, cost and offline operation
The trade-offs compared with cloud models: quality, operational effort and hardware cost
An overview of inference engines (Ollama, llama.cpp and vLLM) and a reference architecture for an in-house LLM service
Lab: install Ollama, pull a first model and chat from the command line
Section 2: Choosing the Right Open-Weight Model
Common model families such as Qwen, Llama and Gemma, and Thai models such as Typhoon
Base, instruct and reasoning models, and when to use each
Reading model cards and licence terms before commercial use
Testing Thai ability with questions drawn from the organisation's real work
Lab: compare the Thai answers of three models on the same task
Section 3: Lab: Quantisation and the GGUF Format
What quantisation is and why large models can run on small machines
Common GGUF levels such as Q4_K_M, Q5_K_M and Q8_0
GPU formats such as AWQ, GPTQ and FP8 used with vLLM
How quantisation affects answer quality, speed and memory
Lab: run the same model at two precision levels and compare the results
Section 4: Workshop: Sizing GPUs and Memory
Estimating weight memory from parameter count and precision
What the KV cache is and how it grows with context length and concurrent users
Running on CPU, one GPU, several GPUs or Apple Silicon, and how they differ
Choosing GPUs and cloud instances that make sense for the workload
Workshop: build a sizing sheet for your own organisation's use case
Section 5: Lab: Ollama and a Standard API
Build an organisation model with a Modelfile: system prompt, template and parameters
Import GGUF files from Hugging Face and set context length, keep-alive and concurrent requests
Run Ollama in Docker with a GPU and expose it to other machines on the network
Call it through the OpenAI-compatible API by changing only the base URL of existing code
Structured output, tool calling and embeddings with in-house models
Lab: a Python script that extracts data from documents without sending data off the machine
Day 2vLLM and Serving LLMs to the Whole Organisation
Section 6: Lab: vLLM on a GPU and Tuning
Why vLLM handles many concurrent users: PagedAttention and continuous batching
Install it on Google Colab or a cloud GPU, and run it with the official Docker image
Start an OpenAI-compatible server with vllm serve, with an API key, tool calling and structured output
Balance max model length and GPU memory utilisation
Use AWQ, GPTQ or FP8 models and enable prefix caching to get more from the GPU
Tensor parallelism across GPUs, and serving several LoRA adapters on one base model
Section 7: Lab: Benchmarking and Choosing Ollama or vLLM
The metrics that matter: time to first token, tokens per second and total throughput
Load testing with the vLLM benchmark tools and a concurrent request script
Reading results as users increase, and finding where the system starts to slow down
Criteria for choosing Ollama or vLLM based on team size and hardware
Lab: measure the same model on both systems and turn the results into a recommendation
Section 8: Lab: An Organisation Chat Interface with Open WebUI
Install Open WebUI with Docker Compose and connect it to Ollama and vLLM
Manage users, roles and access to each model
Upload documents for basic question answering with an internal model
Set system prompts and ready-made model presets for each department
The Open WebUI licence terms to know before using it in an organisation
Section 9: Security and Operations
Close unneeded ports and put a reverse proxy with TLS and API keys in front
Segment the network, restrict callers and keep logs in line with data policy
Monitoring with vLLM metrics, Prometheus and Grafana
Update models and inference engines safely without disrupting users
Lab: build a dashboard for GPU usage and response times
Section 10: Capstone: An In-House LLM Service
Take a brief from a sample organisation with data and budget constraints
Choose the model, quantisation level and inference engine, with reasons
Assemble the system with Docker Compose: inference, Open WebUI and reverse proxy
Test usage and load, then present a plan for scaling
Wrap up with a checklist before opening the service to the whole organisation
Schedule & training options
For individuals — public rounds
No public rounds are open right now. Join the waiting list and we will contact you first when the next round opens, or ask us on LINE. Or call 02-570-8449 or 088-807-9770
Who is Local LLM Deployment with Ollama and vLLM for, and what background is needed?
Built for IT infrastructure teams and system engineers asked to set up an LLM for internal use · DevOps and platform engineers who run AI services on the organisation's servers or cloud · Developers who need an LLM API for internal apps that keeps data inside the organisation Background you should have: Basic Linux command-line and Docker skills · Basic Python and experience calling REST APIs Not sure the fit is right? Talk to our team on LINE @itgenius or call 02-570-8449.
How much does Local LLM Deployment with Ollama and vLLM cost and how long does it run?
THB 7,900 (currently THB 7,110 on promotion). The course runs 12 hours. The price excludes 7% VAT (for payment in a company's name). Pay by bank transfer to the company account, confirm it on our payment page, and we can issue the receipt or tax invoice in your company's name.
Do I get a certificate?
Yes. Everyone who completes the course receives a Certificate of Completion from IT Genius Institute. Each certificate carries its own number, and anyone holding that number can verify it online on our certificate page, so you can add it to your portfolio or pass it to HR as evidence of training.
Where does the training take place, and is there an online option?
You can attend onsite at IT Genius Institute or arrange to join online, and we also run it as a private in-house session for your team. Ask about dates and venues on LINE @itgenius or call 02-570-8449.
What if I fall behind or miss a session — can I retake it?
Yes. You may retake the same course free of charge in a later round, under the institute's conditions. Tell our team which course and round you attended, and we will check it and offer you the rounds that still have seats. Ask us on LINE @itgenius or call 02-570-8449.
How do I enrol, or request a quotation for my company?
Enrol online with the registration form on this page. You can register several attendees at once and enter your tax ID and billing address for the tax invoice. Or request a company quotation straight from the quote button. For anything else call 02-570-8449 or reach us on LINE @itgenius.
Good prompts and RAG solve a lot of problems, but some tasks need more. You may want a model that always answers in your organisation's format, uses domain terms correctly, classifies Thai documents accurately, or lets a small in-house model replace a large…
General AI models are strong at language but do not know your organization's internal information - manuals, policies, contracts, customer data or accumulated know-how. When asked something specific to the organization, AI often answers incorrectly…
12 hours2 days
8,910THB9,900 THB
View details
Local LLM Deployment with Ollama and vLLM12 hrs · 2 days