AI · AIC-74

AI Voice, Audio and Music for Content

AI Voice, Audio and Music for Content is a hands-on course in producing audio with AI, from scripts written for listening and Thai voice-over in ElevenLabs and Gemini TTS to responsible voice cloning, podcasts from documents in NotebookLM, music with Suno and mixing in CapCut. It suits content creators, marketing teams and producers of corporate training materials.

Updated
From 3,510 THB / person 3,900 −10% excl. VAT 7% · group rates available
PDFDownload the course outline
  • Duration6 hours · 1 days
  • FormatOnsite / live online
  • Next roundOn request
  • CertificateIncluded

Course overview

Sound is a key part of today's content: voice-over in short videos, company podcasts, announcements, training materials and music for ads. Hiring voice talent, booking a studio or licensing music every time costs both time and money. AI audio tools can produce Thai voice-over, dialogue and background music in minutes, but the results only sound natural when the script is written for listening, pronunciation is under control and the team understands the rights involved in voices and music.

This course covers audio production with AI from start to finish. Learners write scripts for the ear, create Thai voice-over and direct emotion with ElevenLabs, control tone and two-speaker dialogue with Gemini TTS, clone their own voice properly, turn documents into podcasts with NotebookLM, and compose background music and jingles with Suno. Everything is then cleaned up, mixed and captioned in CapCut, with guidance on copyright, consent for using a person's voice and commercial-use terms. The day closes with a workshop that produces an audio package for a real campaign. (1 day, 6 hours, Beginner level.)

What you’ll gain

  • Choose the right AI audio tool for voice-over, dialogue, podcasts and background music
  • Write scripts for listening and fix pronunciation of Thai, numbers and English words
  • Create Thai voice-over and adjust emotion and pacing with ElevenLabs
  • Direct tone with natural-language instructions and create two-voice dialogue with Gemini TTS
  • Clone your own voice with consent and identity verification as each tool requires
  • Create podcasts from documents with NotebookLM and background music with Suno
  • Clean up audio, mix speech with music and add captions in CapCut
  • Check copyright, voice rights and commercial-use terms before publishing

Who this course is for

  • Content creators, marketing teams and social media admins who make videos and audio clips
  • HR, training and corporate communications teams who produce training materials and internal announcements
  • Business owners who want to make their own podcasts, audio ads or brand jingles
  • Teachers and learning content producers who need Thai narration quickly

Prerequisites

  • Basic computer skills, including handling audio or video files
  • Some experience with an AI chat tool such as Gemini or ChatGPT is helpful
  • No music, voice acting or audio editing background required
  • Learners should bring headphones and may bring their own non-confidential scripts or documents for the workshop

Curriculum

Course Details

A 1-day, 6-hour hands-on course (09:00-16:00) at Beginner level for people who produce voice-over, podcasts or music for content. It uses ElevenLabs, Gemini TTS, NotebookLM, Suno and CapCut through the learner's own or company accounts; voice cloning and commercial music rights require paid plans. The course focuses on AI voice and music work and pairs well with AI Video Editor and AI Content Creator for Business. Learners take home a script-for-listening template, a voice and music prompt set, the audio spot, jingle and short podcast made in the workshop, and a rights checklist for audio work.

Day 1: Creating Voice-over, Podcasts and Music with AI

Section 1 AI Audio Overview and Choosing Tools
  • Text-to-speech, voice cloning, voice changers, dubbing, music and sound effects
  • The strengths of ElevenLabs, Gemini TTS, NotebookLM, Suno and CapCut
  • Which jobs suit AI voices and which still call for a human voice actor
  • Thai accuracy, audio quality and credit cost
Section 2 Lab: Writing Scripts for the Ear
  • Write for listeners: short sentences, rhythm and room to breathe
  • Where AI often misreads: numbers, dates, abbreviations, names and English mixed into Thai
  • Add pronunciation hints and split lines by speaker
  • Use an AI chat tool to fit a script to a target number of seconds
  • Lab: write a 30-second ad script and a 60-second product introduction
Section 3 Lab: Thai Voice-over with ElevenLabs
  • Choose a voice from the Voice Library that suits the brand and audience
  • Generate Thai voice-over with the latest model and compare results
  • Use audio tags to direct emotion, such as excited, whispering or laughing
  • Fix mispronunciations, adjust speed and generate several takes to choose from
  • Lab: voice the script from Section 2 in two styles
Section 4 Lab: Directing Tone and Dialogue with Gemini TTS
  • Generate speech from text in Google AI Studio
  • Direct style, tone, pacing and emotion with natural-language instructions
  • Create two-speaker dialogue, such as host and guest or staff and customer
  • Compare Thai results with ElevenLabs to choose the right tool for each job
  • Lab: produce a two-voice dialogue for a how-to clip
Section 5 Lab: Responsible Voice Cloning
  • How instant and professional voice cloning differ
  • Clone only your own voice, or a voice you have written consent to use
  • Identity verification and prohibited uses under each provider's terms
  • Record a clean source sample for a better-quality clone
  • Lab: clone your own voice and have it read a Thai script
Section 6 Lab: Podcasts from Documents with NotebookLM
  • Add documents, reports or websites as sources in a notebook
  • Generate a Thai Audio Overview and choose its format and length
  • Set the focus and audience so the conversation stays on target
  • Check the content against the source documents before publishing
  • Lab: create a podcast that summarises a company policy or handbook
Section 7 Lab: Background Music and Jingles with Suno
  • Write music prompts: genre, mood, tempo, instruments and length
  • Instrumental tracks for clips and songs with lyrics for jingles
  • Trying Thai lyrics and the limits of their pronunciation
  • Create sound effects for clips with ElevenLabs
  • Lab: make a 15-second brand jingle and an instrumental track for an ad
Section 8 Lab: Cleaning and Mixing Audio in CapCut
  • Reduce noise and make speech clearer
  • Duck music under speech and place sound effects with care
  • Keep loudness consistent for social media and podcasts
  • Generate Thai captions automatically and correct them
  • Lab: mix voice-over, music and sound effects into a sample clip
Section 9 Copyright, Voice Rights and Commercial Terms
  • AI voices and music are not automatically free of copyright, and your rights depend on the plan you use
  • Never imitate real people, artists or copyrighted songs
  • Voices are personal data: consent and storage under PDPA
  • Telling listeners when AI voices are used, and platform policies on AI content
  • A checklist for audio before publishing
Section 10 Workshop: An Audio Package for a Campaign
  • Workshop: take a campaign brief, write the scripts and choose voices
  • Produce a 30-second audio spot, a jingle and a short podcast episode
  • Mix and export files ready for each channel
  • Review against the rights checklist, then listen together and share feedback

Schedule & training options

For individuals — public rounds

No public rounds are open right now. Join the waiting list and we will contact you first when the next round opens, or ask us on LINE. Or call 02-570-8449 or 088-807-9770

For organisations — in-house / private

  • Tailor the content to your team’s tools and projects
  • Your dates, at your office or live online
  • Quotation with tax ID for procurement
Corporate training quote

Instructors

Frequently asked questions

How do you write a script so that an AI voice reads Thai correctly?

Start with short sentences that are easy to follow by ear and leave natural pauses for breath. Then add pronunciation guides where AI voices often slip, such as numbers, dates, abbreviations, proper names and English words mixed into Thai, split the script by speaker, and listen to the result to fix any words that still sound wrong. Asking an AI chat tool to trim the script to a target number of seconds also saves a lot of rework.

What is voice cloning, and how do you use it legitimately?

Voice cloning means an AI learns the character of a person's voice from sample recordings and then speaks new text in that voice. Instant cloning works from a short sample, while professional cloning uses more source audio for a closer match. Clone only your own voice or one you have written consent for, complete the provider's identity verification, and never imitate real people or artists, because a voice counts as personal data under PDPA.

How do you write a Suno prompt for a jingle or music to accompany a video?

A good music prompt states the genre, mood, tempo, instruments and length clearly, for example upbeat pop with acoustic guitar for a short advert. For background music under a voice-over, choose an instrumental so it does not compete with speech. Jingles with lyrics usually need several attempts because Thai lyrics still have pronunciation limits, and before commercial use you must check that your plan grants those rights.

Who is the AI Voice, Audio and Music for Content course for, and what should I know beforehand?

It suits content creators, marketing teams and social media admins, HR and training teams producing learning materials, business owners who want their own podcasts or audio ads, and teachers who need Thai narration. It is a beginner course: basic computer skills and handling audio or video files are enough, with no music, voice acting or audio-editing experience needed, though some experience with an AI chat tool helps you move faster.

What will I be able to do after the AI Voice, Audio and Music for Content course?

You will be able to write scripts for listening and produce Thai voice-overs in ElevenLabs, create two-voice dialogues with Gemini TTS, turn documents into podcasts with NotebookLM, compose jingles and background music with Suno, and mix speech, music and captions in CapCut, checking rights before publishing. Participants receive a Certificate of Completion with its own number that can be verified online.

What do I need to bring to the AI Voice, Audio and Music for Content course?

You use ElevenLabs, Gemini TTS, NotebookLM, Suno and CapCut with your own or your organisation's accounts; voice cloning and commercial music rights require paid plans. Bring headphones, and you are welcome to bring your own non-confidential scripts or documents to use in the workshop.