Skip to content

DeepSeek V4.1 Flash

Cheapest route in the catalog. Use it for autocomplete, short edits and latency-sensitive product features.

DeepSeekFastCodingavailable

Model ID

Pass this as the model field on every request.

modeldeepseek-v4.1-flash
Provider
DeepSeek
Context window
256K tokens
Max output
64K tokens
Access
All plans
Best for
Fast code and chat completion

Endpoints

Both request shapes are accepted for this model.

  • /v1/chat/completionsOpenAI chat completions format
    available
  • /v1/messagesAnthropic message format
    available
Base URLhttps://api.prbycode.com/v1

At a glance

The numbers most teams compare before switching a route.

Context256K tokens
Input$0.20 / 1M
Output$0.80 / 1M
SpeedFast
StreamingSupported

Compare with

Other models in the catalog, side by side.

ModelContextInput / 1MOutput / 1M
Claude Sonnet 51M$3.00$15.00Compare
GLM 5.3 Flash256K$0.25$1.10Compare
GLM 5.2256K$0.60$2.20Compare
DeepSeek V4 Pro256K$0.90$3.40Compare
  • Claude Sonnet 51M context · $3.00 in / $15.00 out
  • GLM 5.3 Flash256K context · $0.25 in / $1.10 out
  • GLM 5.2256K context · $0.60 in / $2.20 out
  • DeepSeek V4 Pro256K context · $0.90 in / $3.40 out