DeepSeek: DeepSeek V4.1 Flash API

DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with 552B backbone parameters that natively processes images and text at up to 1M token context. Its Causal Encoder-Decoder architecture activates only 8B parameters during prefill and 16B during decode, cutting the KV cache footprint to roughly a quarter of DeepSeek-V4-Flash for cost-efficient agentic workloads.

Frequently Asked Questions

What is the context window of DeepSeek V4.1 Flash?

DeepSeek V4.1 Flash supports a context window of up to 1,040,000 tokens.

Does DeepSeek V4.1 Flash support function calling?

Yes. DeepSeek V4.1 Flash supports tool / function calling.

Does DeepSeek V4.1 Flash support reasoning?

Yes. DeepSeek V4.1 Flash is a reasoning-capable model.