Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours (2 days)
Course Outline
Foundations of Speech Synthesis and Voice Cloning
- Introduction to text-to-speech (TTS) and neural voice synthesis concepts
- Distinguishing voice cloning from speech generation: exploring use cases and limitations
- Key architectural models: Tacotron, WaveNet, FastSpeech, and VITS
Utilizing Commercial Platforms
- Working with ElevenLabs and Resemble AI
- Processes for voice creation, replication, and refinement
- Managing API access and text-to-speech workflows
Development with Open-Source Tools
- Setup and configuration of Coqui TTS
- Training bespoke voices and managing associated datasets
- Generating speech with precise control over pitch, speed, and emotional tone
Data Preparation and Voice Dataset Stewardship
- Methods for collecting and refining voice samples
- Techniques for segmenting, labeling, and aligning transcripts
- Ethical sourcing practices and securing voice consent
Application Integration Strategies
- Embedding TTS capabilities into websites and software applications
- Designing IVR systems and interactive bot experiences
- Producing synthetic dialogue for video content and gaming environments
Assessing Quality and Realism
- Applying MOS (Mean Opinion Score) and intelligibility testing standards
- Regulating expressiveness and prosodic features
- Evaluating performance across latency, fidelity, and naturalness
Ethical, Legal, and Governance Frameworks
- Mitigating deepfake risks through responsible usage practices
- Addressing consent, attribution, and copyright considerations
- Navigating relevant regulations and organizational policies
Recap and Future Directions
Requirements
- Foundational knowledge of machine learning principles
- Proficiency with audio file formats and editing software
- Competence in basic Python programming
Target Audience
- AI developers and engineers with an interest in speech synthesis technologies
- Content creators and media technologists investigating voice generation tools
- R&D teams developing personalized or dynamic audio systems