Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Foundations of Speech Synthesis and Voice Cloning
- A comprehensive overview of text-to-speech (TTS) and neural voice synthesis
- Distinguishing voice cloning from general speech generation: exploring use cases and limitations
- Review of leading models including Tacotron, WaveNet, FastSpeech, and VITS
Utilising Commercial Platforms
- Practical application of ElevenLabs and Resemble AI
- Techniques for voice creation, duplication, and fine-tuning
- Managing API access and optimising text-to-speech workflows
Development Using Open-Source Tools
- Setup and configuration of Coqui TTS
- Training custom voice models and managing associated datasets
- Producing speech with granular control over pitch, speed, and emotional tone
Data Preparation and Voice Dataset Administration
- Strategies for collecting and refining voice samples
- Processes for segmenting, labelling, and aligning audio transcripts
- Ensuring ethical data sourcing and obtaining proper voice consent
Application Integration Strategies
- Embedding TTS capabilities into websites and digital applications
- Developing IVR systems and interactive conversational bots
- Generating synthetic dialogue for video productions and gaming environments
Quality Assurance and Realism Assessment
- Applying MOS (Mean Opinion Score) and intelligibility testing standards
- Mastering the control of expressiveness and prosody
- Analysing and comparing latency, audio fidelity, and overall realism
Ethical, Legal, and Governance Frameworks
- Mitigating deepfake risks through responsible usage protocols
- Navigating consent, attribution, and copyright implications
- Adhering to relevant regulations and establishing organisational policies
Conclusion and Future Directions
Requirements
- Solid grasp of machine learning fundamentals
- Proficiency with various audio file formats and editing software
- Competence in basic Python programming
Target Audience
- AI developers and engineers keen on advancing their speech synthesis capabilities
- Content creators and media technologists exploring the potential of voice generation
- R&D teams focused on developing personalised or dynamic audio systems