U2-TTS-Design
Design your exclusive voice with text
Describe the voice style you want in text to generate an exclusive voice
U2-TTS-Design: Design your exclusive voice with text
U2-TTS-Design does not require uploading real-person samples or picking from a fixed voice library. Simply describe the voice you want in natural language (e.g. gender, age, pitch, speed, emotion, accent, texture, character personality, use case), and it can generate a reusable exclusive voice from scratch.
Core Advantages
Zero reference audio, text-only voice creation
No real-person samples, no recording sessions, no professional post-production. Customize a new voice entirely through text descriptions — easy even for beginners.
Full-dimensional voice control
Freely customize gender, age, pitch, speed, emotional tone, dialect accent, character personality, and fine-grained parameters to match your vision.
Generate once, reuse forever
Each generated voice has a unique ID that can be saved long-term and invoked anytime across multiple scripts, businesses, and scenarios with consistent voice identity.
Original and compliant, avoiding copyright risk
Focuses on creating entirely new voices without replicating real people or public figures, avoiding voice infringement and portrait-right disputes for safer enterprise use.
Fast iteration, simplified workflow cost
If the voice is unsatisfactory, simply revise the text description and regenerate — saving time, labor, and communication costs compared with dubbing, sampling, and repeated tuning.
Technical Highlights
More freedom in voice design
Not limited to preset system voices or user audio samples — create brand, character, broadcast, narrative, and companion voices directly through natural language.
Finer control granularity
Describe timbre, emotion, rhythm, pitch, accent, tone, character personality, and scenario use in one go — covering multi-dimensional control from voice to performance.
Stronger expression consistency
Created voices can be stored as reusable voice assets, maintaining a stable voice identity across multiple contents, calls, and business lines.
Better suited for compliant original voices
Voice Design emphasizes describing voice traits rather than imitating specific people — ideal for enterprises building original brand voices and character voice libraries.
Synergy with voice cloning capabilities
First design an ideal base voice style through text, then connect with platform voice cloning for personalized refinement — ideal for virtual hosts, AI assistants, and game NPCs in long-term content production.
Application Scenarios
Personalized audiobooks and text listening
Let AI read to you in your favorite character voice.
Emotional companionship
Whether daily chat and comfort or late-night emotional sharing, get voice responses with authenticity and intimacy.
Personalized navigation and travel companion
Set your own voice or a favorite character voice as your assistant and navigation voice — making every trip more familiar and fun.
Fun dubbing for short videos and vlogs
Easily add quirky, emotional, or professional narration through voice transformation or custom voices — even playing multiple roles alone.
Capabilities
Custom voice creationInput voice feature descriptions plus preview text to generate an exclusive voice ID and preview audio.
Preview and iterative optimizationAdjust text descriptions based on preview results and regenerate voices suited to your scenario.
Voice reuse for synthesisUse the generated voice ID for speech synthesis to keep voice identity consistent.
Voice asset managementQuery lists, view details, and delete unused voices to build your exclusive voice library.
Compliant and safe generationGuides objective voice feature descriptions and restricts real-person or celebrity replication requests for compliant commercial use.
Flexible pricing, tailored solutions, and private deployment
U2-TTS-Design
Design your exclusive voice with text
Describe the voice style you want in text to generate an exclusive voice
U2-TTS-Design: Design your exclusive voice with text
U2-TTS-Design does not require uploading real-person samples or picking from a fixed voice library. Simply describe the voice you want in natural language (e.g. gender, age, pitch, speed, emotion, accent, texture, character personality, use case), and it can generate a reusable exclusive voice from scratch.
Core Advantages
Zero reference audio, text-only voice creation
No real-person samples, no recording sessions, no professional post-production. Customize a new voice entirely through text descriptions — easy even for beginners.
Full-dimensional voice control
Freely customize gender, age, pitch, speed, emotional tone, dialect accent, character personality, and fine-grained parameters to match your vision.
Generate once, reuse forever
Each generated voice has a unique ID that can be saved long-term and invoked anytime across multiple scripts, businesses, and scenarios with consistent voice identity.
Original and compliant, avoiding copyright risk
Focuses on creating entirely new voices without replicating real people or public figures, avoiding voice infringement and portrait-right disputes for safer enterprise use.
Fast iteration, simplified workflow cost
If the voice is unsatisfactory, simply revise the text description and regenerate — saving time, labor, and communication costs compared with dubbing, sampling, and repeated tuning.
Technical Highlights
More freedom in voice design
Not limited to preset system voices or user audio samples — create brand, character, broadcast, narrative, and companion voices directly through natural language.
Finer control granularity
Describe timbre, emotion, rhythm, pitch, accent, tone, character personality, and scenario use in one go — covering multi-dimensional control from voice to performance.
Stronger expression consistency
Created voices can be stored as reusable voice assets, maintaining a stable voice identity across multiple contents, calls, and business lines.
Better suited for compliant original voices
Voice Design emphasizes describing voice traits rather than imitating specific people — ideal for enterprises building original brand voices and character voice libraries.
Synergy with voice cloning capabilities
First design an ideal base voice style through text, then connect with platform voice cloning for personalized refinement — ideal for virtual hosts, AI assistants, and game NPCs in long-term content production.
Application Scenarios
Personalized audiobooks and text listening
Let AI read to you in your favorite character voice.
Emotional companionship
Whether daily chat and comfort or late-night emotional sharing, get voice responses with authenticity and intimacy.
Personalized navigation and travel companion
Set your own voice or a favorite character voice as your assistant and navigation voice — making every trip more familiar and fun.
Fun dubbing for short videos and vlogs
Easily add quirky, emotional, or professional narration through voice transformation or custom voices — even playing multiple roles alone.
Capabilities
- Custom voice creationInput voice feature descriptions plus preview text to generate an exclusive voice ID and preview audio.
- Preview and iterative optimizationAdjust text descriptions based on preview results and regenerate voices suited to your scenario.
- Voice reuse for synthesisUse the generated voice ID for speech synthesis to keep voice identity consistent.
- Voice asset managementQuery lists, view details, and delete unused voices to build your exclusive voice library.
- Compliant and safe generationGuides objective voice feature descriptions and restricts real-person or celebrity replication requests for compliant commercial use.





