Pay-As-You-Go
Model pricing may be adjusted from time to time based on business conditions, operating costs, and strategies. Please pay attention to the latest rules.
Text
Unisound U2
Possesses full-stack software engineering capabilities, capable of completing development-related tasks such as project delivery, code debugging, and security testing. It is also equipped with a comprehensive agent tool ecosystem, supporting various tool and skill calls, and adapting to diverse implementation scenarios.
| Model Name | Input (CNY / 1M tokens) | Cache Hit (CNY / 1M tokens) | Output (CNY / 1M tokens) |
|---|---|---|---|
| Unisound U2 | 1 | 0.02 | 2 |
U2-Med
A professional-grade large model built for the medical, medical insurance, and pharmaceutical industries. It has deep mastery of three core knowledge systems—clinical diagnosis and treatment, medical insurance policy, and pharmaceutical regulations—and provides end-to-end AI capabilities from content parsing and intelligent auditing to assisted decision-making.
| Model Name | Input (CNY / 1M tokens) | Cache Hit (CNY / 1M tokens) | Output (CNY / 1M tokens) |
|---|---|---|---|
| U2-Med | 2.5 | 0.5 | 10 |
Speech
U2-ASR
A professional audio file transcription large model that supports multi-scenario and multi-lingual audio-to-text. It can quickly and accurately transcribe audio content such as meeting recordings, classroom lectures, and customer service recordings into text, helping developers realize the digitization and intelligent analysis of voice content.
Recharge Now Resource packages are more affordable, as low as 0.3 CNY/hour
| Model Name | Input (CNY/hour) |
|---|---|
| U2-ASR | 0.4 |
U2-TTS
Provides natural and fluent text-to-speech capabilities, supporting multi-timbre and multi-emotion speech generation. It can be used in scenarios such as intelligent voice broadcasting and content reading, providing applications with human-like and high-quality voice output experiences.
Recharge Now Resource packages are more affordable, as low as 0.3 CNY/10k chars
| Model Name | Input (CNY/10k chars) |
|---|---|
| U2-TTS | 0.4 |
U2-TTS-Clone
Supports voice cloning and personalized speech generation, restoring target voice features with just one voice sample. Allows developers to easily build personalized voice interaction experiences.
Recharge Now Resource packages are more affordable, as low as 3.38 CNY/10k chars
| Model Name | Input (CNY/10k chars) |
|---|---|
| U2-TTS-Clone | 4.5 |
U2-TTS-Design
Simply describe the target voice in natural language—for example, gender, age, pitch, speech rate, emotion, accent, timbre texture, character personality, and use case—and the model generates a new reusable voice, making it easy to build personal and enterprise-exclusive voice asset libraries.
Recharge Now Resource packages are more affordable, as low as 3.38 CNY/10k chars
| Model Name | Input (CNY/10k chars) |
|---|---|
| U2-TTS-Design | 4.5 |
Vision
U1-OCR
Supports intelligent parsing of complex documents such as PDFs, scanned copies, and images, automatically extracting content and key information such as text, tables, and signatures, outputting standardized structured data, adapting to multiple scenarios, and assisting office automation and intelligent data processing.
Recharge Now Resource packages are more affordable, as low as 0.68 CNY/1M tokens
| Model Name | Output (CNY/1M tokens) |
|---|---|
| U1-OCR | 0.9 |
U1-OCR-Med
Possesses medical document classification and professional information extraction capabilities, accurately adapting to various medical document scenarios such as medical records, examination reports, prescriptions, and billing receipts, efficiently solving industry pain points such as illegible handwriting, terminology abbreviations, seal occlusion, and complex layouts.
Recharge Now Resource packages are more affordable, as low as 0.23 CNY/hour
| Model Name | Output (CNY/1M tokens) |
|---|---|
| U1-OCR-Med | 0.3 |
