Section 1: Workshop: Prompting, RAG or Fine-Tuning
- What fine-tuning can and cannot fix, such as adding knowledge that changes often
- Tasks that benefit from fine-tuning: answer format, classification, extraction and domain terms
- The full cost of fine-tuning: data, GPUs, time and looking after the model afterwards
- Replacing a large model with a small fine-tuned one to cut cost and run in-house
- Workshop: assess your own organisation's use case with a decision checklist
Section 2: How LLMs Are Trained, and the Principles of LoRA and QLoRA
- Pre-training, supervised fine-tuning and preference tuning, and how they differ
- Full fine-tuning compared with parameter-efficient fine-tuning
- How LoRA works and what rank, alpha and target modules mean
- QLoRA: loading a model in 4-bit to train on a GPU with limited memory
- Estimating the GPU memory needed for a given model size and training method
Section 3: Lab: Setting Up a Fine-Tuning Environment
- Choosing between Google Colab, a rented cloud GPU or a local GPU
- Install Unsloth, Transformers, PEFT, TRL and bitsandbytes with compatible versions
- Working with the Hugging Face Hub: tokens, downloading models and private repositories
- An overview of Unsloth Studio for no-code fine-tuning
- Lab: run inference on the base model and record the results as a baseline
Section 4: Designing a Training Dataset
- Instruction and conversation formats, and each model's chat template
- Quality over quantity: diversity, consistency and examples at the edges of the task
- Generating synthetic data with an LLM and having people review it
- Points to watch with Thai data: word segmentation, token counts and consistent language
- PDPA and removing personal data before training on organisational data
Section 5: Lab: Getting Data Ready for Training
- Clean, deduplicate and filter out low-quality examples
- Split into train, validation and test sets without leakage between them
- Convert data to the chat template and check the token length of each example
- Save the dataset to the Hugging Face Hub as a private dataset
- Lab: prepare a Thai dataset for classification and question answering for a sample business