DeepSeek, a Chinese AI development company, announced the new "DeepSeek-V4.1-Flash" model on September 10 (local time). The company stated that by adopting a new architecture and training methods, the model outperforms in some areas the company's own flagship model, "DeepSeek-V4-Pro," as well as OpenAI's "GPT-5.6 Sol" and Anthropic's "Claude Opus 5."
DeepSeek-V4.1-Flash is a multimodal model employing MoE (Mixture-of-Experts), which increases processing efficiency by activating only portions of the model depending on the input content, with a total parameter count of 552 billion (552B). The company claims that it can reduce input costs by adopting a new architecture that improves cost efficiency during data input and reduces the amount of "KV cache," which retains and reuses previous computation results.
Along with the announcement of this new model, DeepSeek applied a reduction in API pricing effective the same day. Under the new pricing structure, off-peak input costs are $0.003 for cache hits (previously $0.007) and $0.15 for cache misses (previously $0.22), while output is $0.6 (previously $0.66). During peak hours, double these rates apply.
Source:
- 「DeepSeek-V4.1-Flash」発表 値上げから一転「より低コストでフラッグシップ超え」うたう (ITmedia AI+, 2026-09-10)
- Hugging Face Model Page