AI · AIC-85

Local LLM Deployment with Ollama and vLLM

Local LLM Deployment with Ollama and vLLM is a hands-on course in running open-weight models on your own machines, from choosing models, quantisation and GPU sizing to serving them through Ollama and vLLM with an OpenAI-compatible API and Open WebUI. It suits IT, DevOps and developer teams whose data must stay in-house, and you leave with a working prototype of an internal LLM service.

Updated
From 7,110 THB / person 7,900 −10% excl. VAT 7% · group rates available
PDFDownload the course outline
  • Duration12 hours · 2 days
  • FormatOnsite / live online
  • Next roundOn request
  • CertificateIncluded

Course overview

Many organisations want to use LLMs on internal documents, customer data or source code, but the rules say that data must not leave for an outside provider. The answer is to run open-weight models on the organisation's own machines or servers. Once a team actually tries it, though, the IT side faces hard questions: which model to pick, how many GB of GPU memory are needed, how far to quantise before quality drops, whether to use Ollama or vLLM, and how to open the service to the whole organisation safely.

This course takes learners through setting up an in-house LLM service from scratch to real use. It starts with choosing models and reading licence terms, understanding quantisation and the GGUF format, and calculating the VRAM, RAM and context length a workload needs. Learners then go deep on Ollama, covering Modelfiles, concurrency settings and the OpenAI-compatible API with structured output and tool calling, before moving

to vLLM on a GPU to serve many users at once. They tune memory, measure performance with benchmarks, add a chat front end with Open WebUI, and put security and monitoring in place. The course closes with a capstone in which each learner designs and deploys an in-house LLM system of their own. (2 days, 6 hours per day, 12 hours in total, Intermediate level.)

What you’ll gain

  • Choose open-weight models that fit the task, the Thai language, the hardware and the licence terms
  • Understand quantisation and pick a precision level that balances quality against resources
  • Calculate the VRAM and RAM needed from model size, context length and concurrent users
  • Install and configure Ollama with Modelfiles and environment variables for real workloads
  • Call models through an OpenAI-compatible API with structured output, tool calling and embeddings
  • Run and tune vLLM on a GPU to handle a large number of concurrent requests
  • Build an in-house LLM service with a chat interface, authentication and monitoring

Who this course is for

  • IT infrastructure teams and system engineers asked to set up an LLM for internal use
  • DevOps and platform engineers who run AI services on the organisation's servers or cloud
  • Developers who need an LLM API for internal apps that keeps data inside the organisation
  • AI engineers comparing the cost and performance of cloud APIs with self-hosting
  • Organisations whose confidentiality or PDPA requirements limit the use of cloud AI

Prerequisites

  • Basic Linux command-line and Docker skills
  • Basic Python and experience calling REST APIs
  • Some experience with ChatGPT or another LLM and a basic idea of prompting
  • A laptop with at least 16 GB of RAM that can run Docker, and a Google account for Google Colab

Curriculum

Course Details

This course runs for 2 days, 6 hours per day (12 hours in total, 09:00-16:00), as lectures with a lab in every topic. Intermediate level. Day one covers model selection, resource sizing and Ollama, with labs on the learner's own laptop using small models. Day two covers vLLM on a GPU, Open WebUI, security and monitoring, using Google Colab or a cloud GPU rented by the learner, whose cost is the learner's responsibility. The course focuses on deploying and serving models, and pairs well with the On-Premise AI Chatbots with LangChain and LangGraph course, which focuses on building chatbot applications on top of in-house models. Learners take home a lab guide, docker-compose files, a GPU sizing sheet and benchmark scripts.

Day 1 Choosing Models, Sizing Resources and Ollama in Depth

Section 1: Why Run LLMs Yourself, and the Tool Landscape

  • Why organisations self-host: confidentiality, PDPA, cost and offline operation
  • The trade-offs compared with cloud models: quality, operational effort and hardware cost
  • An overview of inference engines (Ollama, llama.cpp and vLLM) and a reference architecture for an in-house LLM service
  • Lab: install Ollama, pull a first model and chat from the command line

Section 2: Choosing the Right Open-Weight Model

  • Common model families such as Qwen, Llama and Gemma, and Thai models such as Typhoon
  • Base, instruct and reasoning models, and when to use each
  • Reading model cards and licence terms before commercial use
  • Testing Thai ability with questions drawn from the organisation's real work
  • Lab: compare the Thai answers of three models on the same task

Section 3: Lab: Quantisation and the GGUF Format

  • What quantisation is and why large models can run on small machines
  • Common GGUF levels such as Q4_K_M, Q5_K_M and Q8_0
  • GPU formats such as AWQ, GPTQ and FP8 used with vLLM
  • How quantisation affects answer quality, speed and memory
  • Lab: run the same model at two precision levels and compare the results

Section 4: Workshop: Sizing GPUs and Memory

  • Estimating weight memory from parameter count and precision
  • What the KV cache is and how it grows with context length and concurrent users
  • Running on CPU, one GPU, several GPUs or Apple Silicon, and how they differ
  • Choosing GPUs and cloud instances that make sense for the workload
  • Workshop: build a sizing sheet for your own organisation's use case

Section 5: Lab: Ollama and a Standard API

  • Build an organisation model with a Modelfile: system prompt, template and parameters
  • Import GGUF files from Hugging Face and set context length, keep-alive and concurrent requests
  • Run Ollama in Docker with a GPU and expose it to other machines on the network
  • Call it through the OpenAI-compatible API by changing only the base URL of existing code
  • Structured output, tool calling and embeddings with in-house models
  • Lab: a Python script that extracts data from documents without sending data off the machine
Day 2 vLLM and Serving LLMs to the Whole Organisation

Section 6: Lab: vLLM on a GPU and Tuning

  • Why vLLM handles many concurrent users: PagedAttention and continuous batching
  • Install it on Google Colab or a cloud GPU, and run it with the official Docker image
  • Start an OpenAI-compatible server with vllm serve, with an API key, tool calling and structured output
  • Balance max model length and GPU memory utilisation
  • Use AWQ, GPTQ or FP8 models and enable prefix caching to get more from the GPU
  • Tensor parallelism across GPUs, and serving several LoRA adapters on one base model

Section 7: Lab: Benchmarking and Choosing Ollama or vLLM

  • The metrics that matter: time to first token, tokens per second and total throughput
  • Load testing with the vLLM benchmark tools and a concurrent request script
  • Reading results as users increase, and finding where the system starts to slow down
  • Criteria for choosing Ollama or vLLM based on team size and hardware
  • Lab: measure the same model on both systems and turn the results into a recommendation

Section 8: Lab: An Organisation Chat Interface with Open WebUI

  • Install Open WebUI with Docker Compose and connect it to Ollama and vLLM
  • Manage users, roles and access to each model
  • Upload documents for basic question answering with an internal model
  • Set system prompts and ready-made model presets for each department
  • The Open WebUI licence terms to know before using it in an organisation

Section 9: Security and Operations

  • Close unneeded ports and put a reverse proxy with TLS and API keys in front
  • Segment the network, restrict callers and keep logs in line with data policy
  • Monitoring with vLLM metrics, Prometheus and Grafana
  • Update models and inference engines safely without disrupting users
  • Lab: build a dashboard for GPU usage and response times

Section 10: Capstone: An In-House LLM Service

  • Take a brief from a sample organisation with data and budget constraints
  • Choose the model, quantisation level and inference engine, with reasons
  • Assemble the system with Docker Compose: inference, Open WebUI and reverse proxy
  • Test usage and load, then present a plan for scaling
  • Wrap up with a checklist before opening the service to the whole organisation

Schedule & training options

For individuals — public rounds

No public rounds are open right now. Join the waiting list and we will contact you first when the next round opens, or ask us on LINE. Or call 02-570-8449 or 088-807-9770

For organisations — in-house / private

  • Tailor the content to your team’s tools and projects
  • Your dates, at your office or live online
  • Quotation with tax ID for procurement
Corporate training quote

Instructors

Frequently asked questions

Who is Local LLM Deployment with Ollama and vLLM for, and what background is needed?

Built for IT infrastructure teams and system engineers asked to set up an LLM for internal use · DevOps and platform engineers who run AI services on the organisation's servers or cloud · Developers who need an LLM API for internal apps that keeps data inside the organisation Background you should have: Basic Linux command-line and Docker skills · Basic Python and experience calling REST APIs Not sure the fit is right? Talk to our team on LINE @itgenius or call 02-570-8449.

How much does Local LLM Deployment with Ollama and vLLM cost and how long does it run?

THB 7,900 (currently THB 7,110 on promotion). The course runs 12 hours. The price excludes 7% VAT (for payment in a company's name). Pay by bank transfer to the company account, confirm it on our payment page, and we can issue the receipt or tax invoice in your company's name.

Do I get a certificate?

Yes. Everyone who completes the course receives a Certificate of Completion from IT Genius Institute. Each certificate carries its own number, and anyone holding that number can verify it online on our certificate page, so you can add it to your portfolio or pass it to HR as evidence of training.

Where does the training take place, and is there an online option?

You can attend onsite at IT Genius Institute or arrange to join online, and we also run it as a private in-house session for your team. Ask about dates and venues on LINE @itgenius or call 02-570-8449.

What if I fall behind or miss a session — can I retake it?

Yes. You may retake the same course free of charge in a later round, under the institute's conditions. Tell our team which course and round you attended, and we will check it and offer you the rounds that still have seats. Ask us on LINE @itgenius or call 02-570-8449.

How do I enrol, or request a quotation for my company?

Enrol online with the registration form on this page. You can register several attendees at once and enter your tax ID and billing address for the tax invoice. Or request a company quotation straight from the quote button. For anything else call 02-570-8449 or reach us on LINE @itgenius.