BytePlus: DeepSeek V4 Flash (GA)
byteplus/deepseek-v4-flash-ga-260731
Try DeepSeek V4 Flash (GA) for free in your browser, and add it to your app for free with Puter.js AI API.
Try it free Add to your app// npm install @heyputer/puter.js
import { puter } from '@heyputer/puter.js';
puter.ai.chat("Explain quantum computing in simple terms", {
model: "byteplus/deepseek-v4-flash-ga-260731"
}).then(response => {
document.body.innerHTML = response.message.content;
});
<html>
<body>
<script src="https://js.puter.com/v2/"></script>
<script>
puter.ai.chat("Explain quantum computing in simple terms", {
model: "byteplus/deepseek-v4-flash-ga-260731"
}).then(response => {
document.body.innerHTML = response.message.content;
});
</script>
</body>
</html>
Model Card
DeepSeek V4 Flash (GA) is the July 31, 2026 general-availability release of DeepSeek V4 Flash, a Mixture-of-Experts chat model with 284 billion total parameters and 13 billion activated per token. It is served here through BytePlus's ModelArk platform with a 1 million token context window and adjustable reasoning effort.
This checkpoint was re-post-trained for coding agents and tool use. DeepSeek reports 82.7 on Terminal Bench 2.1 (up from 61.8 for the April preview), 76.7 on Cybergym, 70.3 on Toolathlon, and 54.4 on DeepSWE, ahead of the larger V4 Pro preview on these agentic tasks. These figures are vendor-reported.
It supports tool calling and up to 384,000 output tokens, which suits coding agents, multi-step tool workflows, and long-context codebase or document analysis where cost per token matters.
Context Window 1M
tokens
Max Output 384K
tokens
Input Cost $0.44
per million tokens
Output Cost $1.32
per million tokens
Input text
modalities
Tool Use Yes
Release Date Jul 31, 2026
Try DeepSeek V4 Flash (GA) for free
Try DeepSeek V4 Flash (GA) instantly in your browser.
This playground uses the Puter.js AI API — no API keys or setup required.
More AI Models From BytePlus
Dola Seedream 5.0 Flash
Dola Seedream 5.0 Flash is ByteDance's fast, lower-cost image generation and editing model in the Seedream 5.0 family, offered here through BytePlus's ModelArk API under the "Dola" branding. It supports text-to-image and reference-guided editing with up to 10 input images, local edits targeted with point and bounding-box coordinates in the prompt, transparent-background output, and decomposition of an image into separate transparent layers. Output is 1K to 2K. BytePlus positions it as "faster image creation, stronger layouts, lower cost at scale." A third-party launch-day test reported 15 to 20 seconds per text-to-image request. It is the cheapest model in the Seedream 5.0 line, and Dola Seedream 5.0 Pro remains the better choice for photorealism and fine material detail. It suits high-volume work such as ad concepts, thumbnails, social visuals, and rapid design iteration.
ChatDeepSeek V4.1 Flash
DeepSeek V4.1 Flash is a Mixture-of-Experts chat model from DeepSeek, made available here through BytePlus's ModelArk platform. It succeeds DeepSeek V4 Flash, keeps a 1 million token context window, and accepts text and image input. In DeepSeek's own tests it scored 90.6 on Terminal-Bench 2.1, ahead of Claude Opus 5 (89.1) and GPT-5.6 Sol (88.8), and 74.2 on DeepSWE v1.1 versus 74.0 for Claude Opus 5. DeepSeek also reports it outperforming the larger V4 Pro. It supports tool calling and up to 384,000 output tokens, which suits coding agents, tool-calling pipelines, and long-context workloads at $0.30 per million input tokens and $1.20 per million output tokens.
ChatGLM-5.3 Flash
GLM-5.3 Flash is a multimodal Mixture-of-Experts model from Zhipu AI (Z.ai), available here through BytePlus's hosted API. It has 320 billion total parameters with 18 billion active per token, accepts text and image input, and supports a 1 million token context window. Z.ai reports 84.3 on Terminal-Bench 2.1 and 63.4 on DeepSWE v1.1 (GLM-5.2 scored 46.2). On Terminal-Bench 2.1 that is close to Claude Opus 4.8 at 85.0. Artificial Analysis gives it 57 on its Intelligence Index, three points behind the larger GLM-5.3. It is the lower-cost model in the GLM-5.3 family and supports tool calling. It suits coding agents and multi-step automation where per-token price matters.
Frequently Asked Questions
You can access DeepSeek V4 Flash (GA) by BytePlus through Puter.js AI API. Include the library in your web app or Node.js project and start making calls with just a few lines of JavaScript — no backend and no configuration required. You can also use it with Python or cURL via Puter's OpenAI-compatible API.
DeepSeek V4 Flash (GA) is free to try with a Puter account. Every account includes a free AI allowance, and you can chat with it in the playground on this page. You can upgrade your account anytime for a larger allowance.
DeepSeek V4 Flash (GA) is free to integrate using the Puter.js AI API. With the User-Pays Model, you can add AI to your app for $0, since users cover their own AI usage through their Puter account.
| Price per 1M tokens | |
|---|---|
| Input | $0.44 |
| Output | $1.32 |
DeepSeek V4 Flash (GA) was created by BytePlus and released on Jul 31, 2026.
DeepSeek V4 Flash (GA) supports a context window of 1M tokens. For reference, that is roughly equivalent to 2,048 pages of text.
DeepSeek V4 Flash (GA) can generate up to 384K tokens in a single response.
DeepSeek V4 Flash (GA) accepts the following input types: text. It produces: text.
Yes, DeepSeek V4 Flash (GA) supports tool use (function calling), allowing it to interact with external tools, APIs, and data sources as part of its response flow.
Yes — the DeepSeek V4 Flash (GA) API works with any JavaScript framework, Node.js, or plain HTML through Puter.js. Just include the library and start building. See the documentation for more details.
Add DeepSeek V4 Flash (GA) to your app for free
Developers can integrate DeepSeek V4 Flash (GA) for free using the Puter.js AI API.
With the User-Pays Model, each user covers their own AI usage instead of the developer.