Hy3 preview

Chat Tencent Cloud
API Docs

Description

Hunyuan Hy3 preview is designed for Agent workloads, adopting 295B/21B activated MoE architecture. Within the same model, it provides three modes: no_think (ultra-fast response), think_low (fast thinking), think_high (deep reasoning), adapting to different latency and depth needs from high-frequency interaction to complex engineering tasks. Approaching current best levels on code benchmarks like SWE-bench Verified, 256K context supports cross-file code refactoring and long document analysis. Suitable for developers who need reliable task completion while being sensitive to inference costs.

Specifications

Type
Chat
Vendor
Tencent Cloud
Model ID
hy3-preview
Context Window
256K
Max Input
192K
Max Output
128K
Requests / Minute
60
Tokens / Minute
1000000

Pricing

Token<16k

Input + Output
Input Base Price
$0.0012 / 1000 tokens
Output Base Price
$0.004 / 1000 tokens

Token<16k Cache Hitin

Input Only
Input Base Price
$0.0004 / 1000 tokens

16k<=Token<32k

Input + Output
Input Base Price
$0.0016 / 1000 tokens
Output Base Price
$0.0064 / 1000 tokens

Token>=32k

Input + Output
Input Base Price
$0.002 / 1000 tokens
Output Base Price
$0.008 / 1000 tokens

16k<=Token<32k Cache Hitin

Input Only
Input Base Price
$0.0006 / 1000 tokens

Token>=32k Cache Hitin

Input Only
Input Base Price
$0.0008 / 1000 tokens