deepseek-v4-flash

Chat Deepseek
API Docs

Description

High-efficiency lightweight MoE model, total parameters 284B, activated 13B, natively supporting million-token ultra-long context. Fast reasoning speed, low latency, low calling cost, balanced comprehensive capability, focusing on high-concurrency lightweight tasks, suitable for daily dialogue, content creation, basic RAG, batch copywriting and other essential scenarios.

Specifications

Type
Chat
Vendor
Deepseek
Model ID
deepseek-v4-flash
Context Window
1M
Max Input
1M
Max Output
384K
Requests / Minute
15000
Tokens / Minute
1200000

Pricing

Default

Input + Output
Input Base Price
$0.001 / 1000 tokens
Output Base Price
$0.002 / 1000 tokens

Cache Hitin

Input Only
Input Base Price
$2.0E-5 / 1000 tokens