Programming · PRC-37

NLP and Thai Text Analytics with LLMs

NLP and Thai Text Analytics with LLMs is a hands-on course in analysing Thai reviews and social media text with Python, from word segmentation with PyThaiNLP and sentiment models to topic discovery, LLM classification and structured extraction. It suits data analysts, marketing teams and customer experience teams, and you leave with a Voice of Customer report built from sample data.

Updated
From 6,750 THB / person 7,500 −10% excl. VAT 7% · group rates available
PDFDownload the course outline
  • Duration12 hours · 2 days
  • FormatOnsite / live online
  • Next roundOn request
  • CertificateIncluded

Course overview

Customer text is one of the most valuable and least used sources of data: product reviews, social media comments, chats with admins and open-ended survey answers. Thai text is harder to analyse than English because words are not separated by spaces, and it is full of slang, misspellings, emoji and repeated characters such as 555. Many organisations still read messages one by one or only count star ratings, so they cannot see what customers are actually happy or unhappy about.

This course takes learners through Thai text analysis with Python, from the fundamentals to using LLMs. They start by segmenting and cleaning text with PyThaiNLP, summarise key terms with TF-IDF, build a classic sentiment model with scikit-learn, and use embeddings and Thai language models from Hugging Face. They then discover the topics customers talk about with topic modelling, use LLMs to classify text, analyse sentiment by aspect and extract data as JSON, while measuring accuracy and controlling cost. The course ends by turning the results into an insight report and a review analysis project from start to finish. (2 days, 6 hours per day, 12 hours in total, Intermediate level.)

What you’ll gain

  • Understand the challenges of Thai NLP and choose an analysis method that fits the problem
  • Segment, clean and normalise Thai text with PyThaiNLP
  • Summarise key terms and compare groups of text with TF-IDF
  • Build sentiment classifiers with scikit-learn and Thai language models from Hugging Face
  • Discover the topics customers talk about with embeddings and topic modelling
  • Use LLMs to classify text, analyse sentiment by aspect and extract data as JSON
  • Measure accuracy and compare cost between classic models and LLMs
  • Turn analysis results into an insight report the business can act on

Who this course is for

  • Data analysts and data scientists who need to analyse Thai text
  • Marketing and social listening teams who track customer opinion
  • Customer experience and call centre teams with large volumes of complaints or chats
  • Python developers who want to build automatic text classification
  • Researchers and students working with Thai text

Prerequisites

  • Basic Python and basic table handling with pandas
  • A basic grasp of machine learning concepts helps but is not required
  • A Google account for Google Colab
  • An API key for at least one LLM, such as Gemini, OpenAI or Claude

Curriculum

Course Details

This course runs for 2 days, 6 hours per day (12 hours in total, 09:00-16:00), as lectures with labs on Google Colab using open Thai review and comment datasets and text from a fictional company. Intermediate level. It builds on Python for Machine Learning, focusing specifically on Thai text and on using LLMs for text analysis. Learners use their own or their company's Google account and LLM API key, and API usage may be charged. Before analysing real customer text, learners should always check the source platform's terms and PDPA. Learners take home notebooks for every lab, reusable Thai text cleaning functions, prompt templates and a Voice of Customer report template.

Day 1 Thai NLP Fundamentals with PyThaiNLP

Section 1: Thai Text Analytics Overview

  • Text analysis in business: reviews, social media, chats and open-ended surveys
  • Why Thai is hard: no spaces between words, slang, misspellings and emoji
  • Choosing an approach: rules, machine learning, language models or LLMs
  • Lab: set up Google Colab and load a sample Thai review dataset

Section 2: Lab: Word Segmentation and Processing with PyThaiNLP

  • Segment words with different engines and compare the results
  • Split sentences, tag parts of speech and normalise text
  • Add business terms, such as product and branch names, to the dictionary
  • Lab: segment the whole review set and count the most frequent words

Section 3: Cleaning Social Media Text

  • Handle URLs, hashtags, mentions and emoji
  • Collapse repeated characters and turn 555 into something meaningful
  • Thai stopwords and when not to remove them
  • Lab: build a reusable text cleaning function

Section 4: Lab: Key Terms and Comparison with TF-IDF

  • Bag-of-words, n-grams and TF-IDF with Thai text
  • Find the terms that stand out in positive versus negative reviews
  • Thai word clouds that render the font correctly
  • Lab: compare key terms by branch or by product

Section 5: Lab: A Classic Sentiment Model

  • Prepare labels and split data into train and test sets
  • TF-IDF with logistic regression in scikit-learn
  • Measure with precision, recall, F1 and a confusion matrix
  • Review misclassified text to find the cause
  • Lab: build a positive, neutral and negative review classifier

Section 6: Embeddings and Thai Models on Hugging Face

  • Embeddings and transformers explained simply
  • Thai models such as WangchanBERTa and multilingual embedding models
  • Use Hugging Face pipelines to classify text without training your own model
  • Lab: find reviews with similar meaning using embeddings
Day 2 Deeper Text Analysis with LLMs

Section 7: Lab: Topic Discovery with Topic Modelling

  • Group text with embeddings and clustering
  • Use BERTopic with a Thai word segmenter
  • Have an LLM name and summarise each topic in Thai
  • Lab: find what customers talked about most each month

Section 8: LLMs for Classification and Sentiment

  • Zero-shot and few-shot classification with prompts
  • Aspect-based sentiment: separate feelings about price, service and product
  • Handle sarcasm and messages that cover several issues
  • Lab: classify complaints into categories defined by the business

Section 9: Lab: Extracting Data as JSON with LLMs

  • Structured output that follows a defined schema
  • Extract product, branch, issue and urgency from text
  • Detect and mask personal data before sending text to an LLM
  • Process large volumes in batches and handle errors
  • Lab: turn customer chats into an analysis-ready table

Section 10: Measuring Accuracy and Controlling Cost

  • Build a human-labelled reference set
  • Compare classic models, Thai language models and LLMs on the same metrics
  • Work out cost per thousand messages and cut it by filtering first
  • Lab: a comparison table of accuracy, speed and cost

Section 11: From Insight to Business Report

  • Track sentiment and topic trends over time
  • Have an LLM summarise key issues with real example quotes
  • Export results to CSV or Google Sheets for a dashboard
  • Lab: a one-page Voice of Customer report for management

Section 12: Workshop: Capstone End-to-End Review Analysis

  • Choose a sample dataset or your own text with no personal data
  • Run the full flow from cleaning and classification to sentiment and topics
  • Measure accuracy and summarise the cost of the chosen approach
  • Present the insights and review them together

Schedule & training options

For individuals — public rounds

No public rounds are open right now. Join the waiting list and we will contact you first when the next round opens, or ask us on LINE. Or call 02-570-8449 or 088-807-9770

For organisations — in-house / private

  • Tailor the content to your team’s tools and projects
  • Your dates, at your office or live online
  • Quotation with tax ID for procurement
Corporate training quote

Instructors

Frequently asked questions

Who is NLP and Thai Text Analytics with LLMs for, and what background is needed?

Built for Data analysts and data scientists who need to analyse Thai text · Marketing and social listening teams who track customer opinion · Customer experience and call centre teams with large volumes of complaints or chats Background you should have: Basic Python and basic table handling with pandas · A basic grasp of machine learning concepts helps but is not required Not sure the fit is right? Talk to our team on LINE @itgenius or call 02-570-8449.

How much does NLP and Thai Text Analytics with LLMs cost and how long does it run?

THB 7,500 (currently THB 6,750 on promotion). The course runs 12 hours. The price excludes 7% VAT (for payment in a company's name). Pay by bank transfer to the company account, confirm it on our payment page, and we can issue the receipt or tax invoice in your company's name.

Do I get a certificate?

Yes. Everyone who completes the course receives a Certificate of Completion from IT Genius Institute. Each certificate carries its own number, and anyone holding that number can verify it online on our certificate page, so you can add it to your portfolio or pass it to HR as evidence of training.

Where does the training take place, and is there an online option?

You can attend onsite at IT Genius Institute or arrange to join online, and we also run it as a private in-house session for your team. Ask about dates and venues on LINE @itgenius or call 02-570-8449.

What if I fall behind or miss a session — can I retake it?

Yes. You may retake the same course free of charge in a later round, under the institute's conditions. Tell our team which course and round you attended, and we will check it and offer you the rounds that still have seats. Ask us on LINE @itgenius or call 02-570-8449.

How do I enrol, or request a quotation for my company?

Enrol online with the registration form on this page. You can register several attendees at once and enter your tax ID and billing address for the tax invoice. Or request a company quotation straight from the quote button. For anything else call 02-570-8449 or reach us on LINE @itgenius.