GeneralFoundational model capabilities for multi-industry, multi-scenario use
TextUnisound U2
Native agent LLM for task execution with top-tier reasoning, coding, and long-text analysis.
SpeechU2-ASR
Transcribe audio to text quickly and accurately for multiple scenarios.
SpeechU2-TTS
Text-to-speech with natural, expressive speech in multiple voices.
SpeechU2-TTS-Clone
All-in-one TTS and voice cloning: clone in seconds with high fidelity.
SpeechU2-TTS-Design
Freely tune voice style and emotion to produce custom voices quickly and meet diverse speech synthesis needs.
VisionU1-OCR
Four core capabilities: document classification, ID recognition, layout restoration, and key information extraction.
MedicalSpecialized model capabilities for healthcare scenarios
TextU2-Med
Professional LLM for medical, insurance, and pharma scenarios, covering clinical, auditing, risk control, and health management.
SpeechU2-ASR-MedComing Soon
Medical speech recognition model
SpeechU2-TTS-MedComing Soon
Medical speech synthesis model
VisionU1-OCR-Med
Medical OCR combining document classification, layout parsing, and professional information extraction.
VisionU2-RadiMedComing Soon
Medical imaging model
U2 SkillHub
AI Agent skill store with one-click install, a wide selection of curated skills, and custom Skill uploads.
U2 Agent
Native Agent LLM intelligent assistant with chat and expert Agent dual modes, suited for office, finance, and other scenarios.
U2Claw
Desktop AI Agent lobster tool that aggregates multi-domain AI experts.
  • Home
Text
  • Text generation
Speech
  • One-shot ASR
  • Real-time ASR
  • Speech to Text (Async)
  • Speech Synthesis (Async)
  • Speech Synthesis (Sync)
  • Fast Voice Clone
  • Voice Design
Vision
  • Document Parsing (Async)
  • OCR Extract
Resource Management
  • Voice List
  • File List

Welcome to Unisound Maas Platform

One-stop debugging and management of Unisound Open Platform APIs

Text Generation

Debug the text generation API with multi-turn conversations

One-shot ASR

Short utterance recognition with streaming partial results, up to 25s per session

Real-time ASR

Long-form streaming transcription with low latency

Speech to Text (Async)

Debug ASR API, transcribe long audio files to text

Speech Synthesis (Async)

Text to speech, generate long audio asynchronously and download

Speech Synthesis (Sync)

Real-time WebSocket text-to-speech for short text audition and download

Fast Voice Clone

Quickly clone and generate highly similar personalized voices with few audio samples

Voice Design

Customize voices with text, real-time audition, one-click fine-tuning

Document Parsing (Async)

Debug OCR API, extract document content, support multi-format output

OCR Extract

Debug image information extraction API, identify and extract image content

Voice List

Manage and view all available voices, including system presets and custom clones

File Management

Manage your input/output files, voice clones and sample audios

Quick Start

Select the module to debug from the left menu
Configure API parameters and send the request
View the response to quickly validate API capabilities