deepseek-v4-flash
Description
High-efficiency lightweight MoE model, total parameters 284B, activated 13B, natively supporting million-token ultra-long context. Fast reasoning speed, low latency, low calling cost, balanced comprehensive capability, focusing on high-concurrency lightweight tasks, suitable for daily dialogue, content creation, basic RAG, batch copywriting and other essential scenarios.
Specifications
- Type
- Chat
- Vendor
- Deepseek
- Model ID
-
deepseek-v4-flash - Context Window
- 1M
- Max Input
- 1M
- Max Output
- 384K
- Requests / Minute
- 15000
- Tokens / Minute
- 1200000
Pricing
Default
Input + Output- Input Base Price
- $0.001 / 1000 tokens
- Output Base Price
- $0.002 / 1000 tokens
Cache Hitin
Input Only- Input Base Price
- $2.0E-5 / 1000 tokens