toolcall.
ModelsSep 10, 2026, 14:28 UTC

DeepSeek pushes AI model prices lower with V4.1 Flash

The new API model supports a 1M-token context, tool calls and OpenAI- or Anthropic-compatible access, with cache-hit input priced from 0.003 USD per 1M tokens.

DeepSeek has made V4.1 Flash its new low-cost API model, sharpening the price pressure around frontier-style AI services. The company now lists deepseek-flash as DeepSeek-V4.1-Flash in its API docs, with OpenAI-format and Anthropic-format endpoints, tool calls, Responses API support and vision support.

The pricing is the headline. DeepSeek lists off-peak cache-hit input at 0.003 USD per 1M tokens, peak cache-hit input at 0.006 USD, cache-miss input from 0.15 USD, and output from 0.6 USD per 1M tokens. The model also offers a 1M-token context window and a maximum output length of 384K tokens.

DeepSeek says V4.1 Flash has surpassed V4 Pro in performance, cost, speed and total time in its own testing, and that V4 Pro traffic will be routed to V4.1 Flash until a future Pro successor is ready. Bloomberg framed the launch as a fresh blow to rivals including OpenAI, Anthropic and Z AI, because it makes capable API access cheaper for builders who can switch providers.

The usual caveat applies: listed prices do not prove equivalent quality for every task. But for agent, coding and high-volume API workloads, DeepSeek is again pushing the market toward lower inference costs.

Sources

Mentioned

APIdeepseekmodelspricing