Unlock exclusive introductory pricing for the newly launched Gemini 3.5 Flash.
D

DeepSeek V4 Flash

์ž…๋ ฅ:$0.12/M
์ถœ๋ ฅ:$0.24/M
์ถœ์‹œ์ผ:Apr 23, 2026

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance.

์ƒˆ๋กœ์šด
์ธ๊ธฐ
์ƒ์—…์  ์‚ฌ์šฉ

DeepSeek V4 Flash์˜ Playground

DeepSeek V4 Flash์˜ Playground๋ฅผ ํƒ์ƒ‰ํ•˜์„ธ์š” โ€” ๋ชจ๋ธ์„ ํ…Œ์ŠคํŠธํ•˜๊ณ  ์‹ค์‹œ๊ฐ„์œผ๋กœ ์ฟผ๋ฆฌ๋ฅผ ์‹คํ–‰ํ•˜๋Š” ๋Œ€ํ™”ํ˜• ํ™˜๊ฒฝ์ž…๋‹ˆ๋‹ค. ํ”„๋กฌํ”„ํŠธ๋ฅผ ์‹œ๋„ํ•˜๊ณ , ๋งค๊ฐœ๋ณ€์ˆ˜๋ฅผ ์กฐ์ •ํ•˜๋ฉฐ, ์ฆ‰์‹œ ๋ฐ˜๋ณตํ•˜์—ฌ ๊ฐœ๋ฐœ์„ ๊ฐ€์†ํ™”ํ•˜๊ณ  ์‚ฌ์šฉ ์‚ฌ๋ก€๋ฅผ ๊ฒ€์ฆํ•˜์„ธ์š”.

Technical specifications of DeepSeek-V4-Flash

ItemDetails
ModelDeepSeek-V4-Flash
ProviderDeepSeek
FamilyDeepSeek-V4 preview series
ArchitectureMixture-of-Experts (MoE)
Total parameters284B
Activated parameters13B
Context length1,000,000 tokens
PrecisionFP4 + FP8 mixed
Reasoning modesNon-think, Think, Think Max
Release statusPreview model
LicenseMIT License

What is DeepSeek-V4-Flash?

DeepSeek-V4-Flash is DeepSeekโ€™s efficiency-focused preview model in the V4 series. It is built as a Mixture-of-Experts language model with a relatively small active footprint for its size, which helps it stay responsive while still supporting a very large 1M-token context window.

Main features of DeepSeek-V4-Flash

  • Million-token context: The model supports a 1,000,000-token context window, which makes it suitable for very long documents, large codebases, and multi-step agent sessions.
  • Efficiency-first MoE design: It uses 284B total parameters but only 13B activated parameters per request, a setup aimed at faster and more efficient inference.
  • Three reasoning modes: Non-think, Think, and Think Max let you trade speed for deeper reasoning when the task gets harder.
  • Strong long-context architecture: DeepSeek says the V4 series combines Compressed Sparse Attention and Heavily Compressed Attention to improve long-context efficiency.
  • Competitive coding and agent behavior: The model card reports strong results on coding and agentic benchmarks, including HumanEval, SWE Verified, Terminal Bench 2.0, and BrowseComp.
  • Open weights and local deployment: The release includes model weights, local inference guidance, and an MIT License, which makes self-hosting and experimentation practical.

Benchmark performance of DeepSeek-V4-Flash

Selected results from the official model card show that DeepSeek-V4-Flash improves over DeepSeek-V3.2-Base on several core benchmarks:

BenchmarkDeepSeek-V3.2-BaseDeepSeek-V4-Flash-BaseDeepSeek-V4-Pro-Base
AGIEval (EM)80.182.683.1
MMLU (EM)87.888.790.1
MMLU-Pro (EM)65.568.373.5
HumanEval (Pass@1)62.869.576.8
LongBench-V2 (EM)40.244.751.5

In the reasoning-and-agent table, the Flash variant also posts solid results on terminal and software tasks, with Flash Max reaching 56.9 on Terminal Bench 2.0 and 79.0 on SWE Verified, while still trailing the larger Pro model on the hardest knowledge-heavy and agentic tasks.

DeepSeek-V4-Flash vs DeepSeek-V4-Pro vs DeepSeek-V3.2

ModelBest fitTradeoff
DeepSeek-V4-FlashFast, long-context work, coding assistants, and high-throughput agent flowsSlightly behind Pro on pure knowledge and the most complex agentic tasks
DeepSeek-V4-ProHighest-capability tasks, deeper reasoning, and harder agent workflowsHeavier and less efficiency-oriented than Flash
DeepSeek-V3.2Older baseline for comparison and migration planningLower benchmark performance than V4-Flash on the official tables

Typical use cases for DeepSeek-V4-Flash

  1. Long-document analysis for contracts, research packs, support knowledge bases, and internal wikis.
  2. Coding assistants that need to inspect big repos, follow instructions across many files, and keep context alive.
  3. Agent workflows where the model needs to reason, call tools, and iterate without losing the thread.
  4. Enterprise chat systems that benefit from a very large context window and low-friction deployment.
  5. Prototype local deployments for teams that want to evaluate DeepSeek-V4 behavior before production hardening.

How to access and use Deepseek v4 Flash API

Step 1: Sign Up for API Key

Log in toย cometapi.com. If you are not our user yet, please register first. Sign into yourย CometAPI console. Get the access credential API key of the interface. Click โ€œAdd Tokenโ€ at the API token in the personal center, get the token key: sk-xxxxx and submit.

Step 2: Send Requests toย deepseek v4 flash API

Select the โ€œdeepseek-v4-flashโ€ endpoint to send the API request and set the request body. The request method and request body are obtained from our website API doc. Our website also provides Apifox test for your convenience. Replace <YOUR_API_KEY> with your actual CometAPI key from your account.ย Where to call it:ย ย Anthropic Messagesย format andย Chatย format.

Insert your question or request into the content fieldโ€”this is what the model will respond to . Process the API response to get the generated answer.

Step 3: Retrieve and Verify Results

Process the API response to get the generated answer. After processing, the API responds with the task status and output data.Enable features such as streaming, prompt caching, or long-context handling via standard parameters.

์ž์ฃผ ๋ฌป๋Š” ์งˆ๋ฌธ

DeepSeek V4 Flash ๊ฐ€๊ฒฉ

[๋ชจ๋ธ๋ช…]์˜ ๊ฒฝ์Ÿ๋ ฅ ์žˆ๋Š” ๊ฐ€๊ฒฉ์„ ์‚ดํŽด๋ณด์„ธ์š”. ๋‹ค์–‘ํ•œ ์˜ˆ์‚ฐ๊ณผ ์‚ฌ์šฉ ์š”๊ตฌ์— ๋งž๊ฒŒ ์„ค๊ณ„๋˜์—ˆ์Šต๋‹ˆ๋‹ค. ์œ ์—ฐํ•œ ์š”๊ธˆ์ œ๋กœ ์‚ฌ์šฉํ•œ ๋งŒํผ๋งŒ ์ง€๋ถˆํ•˜๋ฏ€๋กœ ์š”๊ตฌ์‚ฌํ•ญ์ด ์ฆ๊ฐ€ํ•จ์— ๋”ฐ๋ผ ์‰ฝ๊ฒŒ ํ™•์žฅํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค. [๋ชจ๋ธ๋ช…]์ด ๋น„์šฉ์„ ๊ด€๋ฆฌ ๊ฐ€๋Šฅํ•œ ์ˆ˜์ค€์œผ๋กœ ์œ ์ง€ํ•˜๋ฉด์„œ ํ”„๋กœ์ ํŠธ๋ฅผ ์–ด๋–ป๊ฒŒ ํ–ฅ์ƒ์‹œํ‚ฌ ์ˆ˜ ์žˆ๋Š”์ง€ ์•Œ์•„๋ณด์„ธ์š”.

Comet Price (USD / M Tokens)Official Price (USD / M Tokens)Discount
์ž…๋ ฅ:$0.12/M
์ถœ๋ ฅ:$0.24/M
์ž…๋ ฅ:$0.15/M
์ถœ๋ ฅ:$0.3/M
-20%

DeepSeek V4 Flash์˜ ์ƒ˜ํ”Œ ์ฝ”๋“œ ๋ฐ API

[๋ชจ๋ธ ์ด๋ฆ„]์˜ ํฌ๊ด„์ ์ธ ์ƒ˜ํ”Œ ์ฝ”๋“œ์™€ API ๋ฆฌ์†Œ์Šค์— ์•ก์„ธ์Šคํ•˜์—ฌ ํ†ตํ•ฉ ํ”„๋กœ์„ธ์Šค๋ฅผ ๊ฐ„์†Œํ™”ํ•˜์„ธ์š”. ์ž์„ธํ•œ ๋ฌธ์„œ๋Š” ๋‹จ๊ณ„๋ณ„ ๊ฐ€์ด๋“œ๋ฅผ ์ œ๊ณตํ•˜์—ฌ ํ”„๋กœ์ ํŠธ์—์„œ [๋ชจ๋ธ ์ด๋ฆ„]์˜ ๋ชจ๋“  ์ž ์žฌ๋ ฅ์„ ํ™œ์šฉํ•  ์ˆ˜ ์žˆ๋„๋ก ๋•์Šต๋‹ˆ๋‹ค.

curl https://api.cometapi.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $COMETAPI_KEY" \
  -d '{
    "model": "deepseek-v4-flash",
    "messages": [
      {
        "role": "system",
        "content": "You are a helpful assistant."
      },
      {
        "role": "user",
        "content": "Hello!"
      }
    ],
    "thinking": {
      "type": "enabled"
    },
    "reasoning_effort": "high",
    "stream": false
  }'

cURL Code Example

curl https://api.cometapi.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $COMETAPI_KEY" \
  -d '{
    "model": "deepseek-v4-flash",
    "messages": [
      {
        "role": "system",
        "content": "You are a helpful assistant."
      },
      {
        "role": "user",
        "content": "Hello!"
      }
    ],
    "thinking": {
      "type": "enabled"
    },
    "reasoning_effort": "high",
    "stream": false
  }'

Python Code Example

from openai import OpenAI
import os

# Get your CometAPI key from https://www.cometapi.com/console/token, and paste it here
COMETAPI_KEY = os.environ.get("COMETAPI_KEY") or "<YOUR_COMETAPI_KEY>"
BASE_URL = "https://api.cometapi.com/v1"

client = OpenAI(base_url=BASE_URL, api_key=COMETAPI_KEY)

completion = client.chat.completions.create(
    model="deepseek-v4-flash",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Hello!"},
    ],
    stream=False,
    extra_body={
        "thinking": {"type": "enabled"},
        "reasoning_effort": "high",
    },
)

print(completion.choices[0].message.content)

JavaScript Code Example

import OpenAI from "openai";

// Get your CometAPI key from https://www.cometapi.com/console/token, and paste it here
const api_key = process.env.COMETAPI_KEY || "<YOUR_COMETAPI_KEY>";
const base_url = "https://api.cometapi.com/v1";

const client = new OpenAI({
  apiKey: api_key,
  baseURL: base_url,
});

const completion = await client.chat.completions.create({
  model: "deepseek-v4-flash",
  messages: [
    { role: "system", content: "You are a helpful assistant." },
    { role: "user", content: "Hello!" },
  ],
  thinking: { type: "enabled" },
  reasoning_effort: "high",
  stream: false,
});

console.log(completion.choices[0].message.content);

Uptime

์ง€๋‚œ 30์ผ๊ฐ„์˜ ์š”์ฒญ ์„ฑ๊ณต๋ฅ ๋กœ, ๊ฐ ๋ชจ๋ธ ์ œ๊ณต์ž์˜ ์‹ ๋ขฐ์„ฑ์„ ๋ฐ˜์˜ํ•ฉ๋‹ˆ๋‹ค. CometAPI๋Š” ์—ฐ๊ฒฐ๋œ ๋ชจ๋“  ์ œ๊ณต์ž๋ฅผ ์‹ค์‹œ๊ฐ„์œผ๋กœ 24์‹œ๊ฐ„ ๋ชจ๋‹ˆํ„ฐ๋งํ•ฉ๋‹ˆ๋‹ค.

RespondLIVE
2866msAvg. Response
UptimeLIVE
100.0%Avg. Uptime