Unisound U2 Native Agent Model
Token Plan Launch
Model Type
Vendor
Use Case

Unisound U2

Updated 2026-06-24

Text GenVibe CodingAgent

Unisound's flagship U2 model with high intelligence per token and native agent capabilities. Compact yet capable, lower inference cost; autonomously plans, executes, and refines tasks end to end.

Input: 1 ¥/M tokensOutput: 2 ¥/M tokens

U1-InsureMed

Updated 2026-06-24

Text GenMedical Insurance

Industry model for medical insurance. Parses records, reports, policies, and coverage docs with plain-language summaries, Q&A, and deep analysis.

Input: 2.5 ¥/M tokensOutput: 10 ¥/M tokens

U2-TTS

Updated 2026-06-24

Speech Synth.AIGC

Human-like TTS with style, emotion, and multi-language/dialect support. Async long-text synthesis and flexible audio output for production scale.

Input: 0.4 ¥/10k chars

U2-TTS-Clone

Updated 2026-06-24

Voice CloneAIGC

Clone voices from minimal samples with natural, emotional speech. Persistent cloned voices for reusable brand assets.

Input: 4.5 ¥/10k chars

U2-ASR

Updated 2026-06-24

Speech Recog.

Universal ASR with context understanding and custom vocabulary. Multi-dialect, bilingual, with one-shot and real-time transcription.

Input: 0.4 ¥/hour

U1-OCR

Updated 2026-06-24

Visual Und.

Intelligent document parsing beyond OCR—classification, layout restoration, and key info extraction. Converts unstructured docs to structured data.

Output: 0.9 ¥/M tokens

U1-OCR-Med

Updated 2026-06-24

Visual Und.Medical OCR

Medical document intelligence for records, reports, prescriptions, and bills. Handles handwriting, abbreviations, stamps, and complex layouts.

Output: 0.3 ¥/M tokens

DeepSeek-V4-Flash

Updated 2026-06-24

Text GenVibe CodingAgent

Efficient MoE (284B total, 13B active) with million-token context. Fast, low-latency, cost-effective for chat, content, RAG, and batch copywriting.

Input: 1 ¥/M tokensOutput: 2 ¥/M tokens

DeepSeek-V4-Pro

Updated 2026-06-24

Text GenVibe CodingAgent

Flagship MoE (1.6T total, 49B active) with million-token context. Top math, reasoning, code, and long-text analysis for research and advanced agents.

Input: 12 ¥/M tokensOutput: 24 ¥/M tokens

Kimi-K2.5

Updated 2026-06-24

Text GenVibe CodingAgent

Multimodal agent model unifying vision-language understanding, instant/thinking modes, and chat/agent paradigms.

Input: 4 ¥/M tokensOutput: 21 ¥/M tokens

Kimi-K2.6

Updated 2026-06-24

Text GenVibe CodingAgent

Stronger long-horizon coding with better instruction following and self-correction. Text, image, and video input; thinking and non-thinking modes.

Input: 6.5 ¥/M tokensOutput: 27 ¥/M tokens

GLM-5.2

Updated 2026-06-24

Text GenVibe Coding

Zhipu flagship for long-horizon tasks with 1M lossless context. Full-stack engineering from design to deployment for complex projects.

Input: 8 ¥/M tokensOutput: 28 ¥/M tokens

GLM-5.1

Updated 2026-06-24

Text GenVibe CodingAgent

Long-horizon model (744B params, 200K context, 128K output). Strong reasoning, long-text, and code generation for interactive and enterprise use.

Input: 6 ¥/M tokensOutput: 24 ¥/M tokens

GLM-5

Updated 2026-06-24

Text GenVibe CodingAgent

Next-gen model for coding and agents—open-source SOTA in complex engineering and long tasks, nearing Claude Opus in real coding.

Input: 4 ¥/M tokensOutput: 18 ¥/M tokens

MiniMax-M2.5

Updated 2026-06-24

Text GenVibe CodingAgent

Production-grade agent model built for real-world productivity—faster, stronger, and more cost-effective.

Input: 2.1 ¥/M tokensOutput: 8.4 ¥/M tokens

Qwen3.7-Plus

Updated 2026-06-24

Text GenVibe CodingAgent

High-value Plus model with upgraded vision-language, coding, tools, and agent workflows. Scene perception, GUI ops, and visual-reference code gen.

Input: 2 ¥/M tokensOutput: 8 ¥/M tokens

Qwen3.7-Max-2026-06-08

Updated 2026-06-24

Text GenVibe CodingAgent

Largest Max model with broad, deep agent skills—coding, productivity, long autonomous runs, plus vision understanding.

Input: 12 ¥/M tokensOutput: 36 ¥/M tokens

Qwen3.6-Plus

Updated 2026-06-24

Text GenVibe CodingAgent

Native vision-language Plus rivaling frontier models. Major gains over 3.5 in agentic coding, frontend, OCR, and object localization.

Input: 2 ¥/M tokensOutput: 12 ¥/M tokens

Qwen3.6-Flash

Updated 2026-06-24

Text GenVibe CodingAgent

Fast vision-language model with major gains over 3.5-Flash. Stronger agentic coding, math, code reasoning, and spatial detection.

Input: 1.2 ¥/M tokensOutput: 7.2 ¥/M tokens

Qwen3.6-35B-A3B

Updated 2026-06-24

Text GenVibe CodingAgent

35B-A3B vision-language MoE with linear attention and sparse experts. Better efficiency; gains in agentic coding, reasoning, and spatial detection.

Input: 1.8 ¥/M tokensOutput: 10.8 ¥/M tokens