Unisound U2

Native Agent model — faster, leaner, stronger

Redefining speed, cost, and reliable execution with dense intelligence

U1-InsureMed

Records, reports, insurance questions: ask when in doubt

From reports to insurance questions, it explains complex issues clearly

U2-ASR

Beyond hearing, truly understanding expression

Recognizes seven Chinese dialects, mixed speech and industry terms

U2-TTS

Meaning through voice, warmth in every expression

More than reading text aloud, it interprets tone, emotion, and detail

U2-TTS-Clone

Few-shot high fidelity, fast voice cloning

Voice timbre cloning and emotion transfer for more human expression

U1-OCR

Recognize IDs, parse documents, extract data

Intelligent ID and document parsing, one-click key data extraction

U1-OCR-Med

Medical docs, smart layouts, precise extraction

Medical document parsing for classification, archiving, and extraction

Token Plan

5 leading AI models, one key to access

Releases

View the latest model release updates

Docs

API reference and developer guides

Unisound Token Hub Model Matrix

Multimodal coverage across text, audio, and vision for all scenarios

Unisound U2

Text

Supports autonomous planning, tool calling, multi-agent collaboration, long-horizon task orchestration, and self-correction, enabling end-to-end execution from goal understanding to result delivery while balancing efficient reasoning with low token consumption.

View details

U1-InsureMed

Text

Healthcare and insurance Q&A over medical records, labs, checkups, and policy documents, summarizing, interpreting, and surfacing key points. Multi-turn dialogue keeps context for follow-up questions.

View details

U2-ASR

Speech

Recognition accuracy in complex noise and dialect scenarios is the first in the industry to exceed 90%, with multilingual and full-system dialect coverage, structured long-audio transcription, and fast, reliable deployment.

View details

U2-TTS

Speech

Breakthroughs in semantic understanding and emotion for natural, expressive speech

View details

U2-TTS-Clone

Speech

Clone voice timbre in seconds from a single-sentence sample, with emotion transfer and cross-lingual Chinese-English synthesis, enabling rapid accumulation of exclusive brand/character voice assets.

View details

U1-OCR

Vision

Integrates visual and semantic parsing capabilities with comprehensive document processing. Compatible with multiple languages, handwriting, and various non-standard documents, meeting the needs of both personal document management and enterprise document digitalization.

View details

U1-OCR-Med

Vision

Tailor-made for medical scenarios, adapting to various medical documents such as medical records, prescriptions, inspection reports, and billing receipts. Efficiently handles complex medical expressions and special cases like handwriting and seals, empowering smart healthcare businesses.

View details

Build with Leading Partners

Deep industry expertise, delivering value to 500+ enterprises

Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo
Partner logo

Flexible pricing, tailored solutions, and private deployment