Cast an AI voice the same way you'd cast a voice actor. From text-to-speech to AI dubbing, create content effortlessly with TTS that understands emotion and context.
Vision
The AI we create becomes
the most natural interface for our lives.
Products
Cast an AI voice the same way you'd cast a voice actor. From text-to-speech to AI dubbing, create content effortlessly with TTS that understands emotion and context.
Get content automation, apps, and conversational services up and running in five minutes, with no complex set up needed.
Design AI agent personas tailored to any channel or service. Bring your own IP to life and connect with customers in real time through Live Chat.
Technology & Research
Technology built on years of R&D and dozens of published papers.
Our research has been presented at the world's leading AI conferences,
including NeurIPS, Interspeech, and ICASSP.
Typecast SSFM (Speech Synthesis Foundation Model) is built on billions of parameters and over one million hours of speech data. It encodes speech into tokens, and a transformer learns them the way it learns language. The result is a model that does more than read a sentence. It understands how people speak.
Add context, emotion, and style to your script as a prompt, and our in-house audio-visual foundation model generates the voice, facial expressions, and gestures together as a single video. Every element of expression can be controlled precisely through parameters and prompts.
We develop the full stack of listening, thinking, and speaking ourselves, and connect it as one seamless system. The result is an AI agent that responds without delay, converses as naturally as a person, and carries out real tasks.
PixSwap: High-Resolution Face Swapping for Effective Reflection of Identity via Pixel-Level Supervision with Synthetic Paired Dataset
DRAFT: Dense Retrieval Augmented Few-shot Topic classifier Framework
Cross-speaker Emotion Transfer by Manipulating Speech Style Latents
Text-driven Emotional Style Control and Cross-speaker Style Transfer in Neural TTS
One-Shot Face Reenactment on Megapixels
EdiTTS: Score-based Editing for Controllable Text-to-Speech
MLP Singer: Towards Rapid Parallel Korean Singing Voice Synthesis
Diff-TTS: A Denoising Diffusion Model for Text-to-Speech
Large-scale Speaker Retrieval on Random Speaker Variability Subspace
Learning Pronunciation from a Foreign Language in Speech Synthesis Networks
Robust and Fine-grained Prosody Control of End-to-End Speech Synthesis
Voice Imitating Text-to-Speech Neural Networks
Emotional End-to-End Neural Speech Synthesizer