Research

Should You Switch to DeepSeek V4 Flash 0731? Model, Cost, and Self-Hosting Comparison — August 2026

As of August 1, 2026, public evidence yields no production winner: test deepseek-v4-flash first for low-cost text coding and record its 0731 fingerprint, Gemini 3.6 Flash first for multimodal work, and Claude Sonnet 5 for fixed-snapshot or ZDR needs; the MIT 0731 weights are public, while local adoption starts after a four-accelerator TCO trial and owned tasks pass with the incumbent ready for rollback.

On a warm hand-drawn investigation desk, three model work trays connect through a blue-steel switchboard to a trial gate and rollback lever, with a one-card bench, multi-card nodes, and a multi-rack machine room behind them
In this article

Research basisAs of 2026-08-01 14:00 CST, this report reviewed the public evaluations, API identity, pricing, context, tool interfaces, and same-day MIT weight release for DeepSeek V4 Flash 0731; pinned Hugging Face revision 7872f01b1d1fe23eabc4c98b48bffcef5a386062 and summed the exact bytes of 48 safetensors from the blobs manifest; calculated a billing example with 200K uncached input and 20K output for DeepSeek plus eight current managed comparators; and recorded official or framework-validated hardware topologies, artifact sizes, licenses, and public GPU rent for five local routes. SOSEC did not run cross-model quality, throughput, latency, energy, or security tests on one common task set, so public evidence determines POC order only; each production decision must come from the reader's own repositories, tools, and acceptance criteria.

SourceDeepSeek API documentation, update notice, official Hugging Face weights, and fixed repository manifest / official model and pricing material from OpenAI, Anthropic, Google, Z.ai, Moonshot AI, Alibaba Cloud, xAI, MiniMax, and Qwen / official vLLM, SGLang, and Runpod deployment and pricing material / SOSEC identity, price, weight-byte, and break-even review on 2026-08-01

1 The decision now: test 0731 first for low-cost coding and wait for owned results before production

Verdict: DeepSeek V4 Flash 0731 is the first August 2026 POC for a low-cost, text-only coding agent. Send requests to deepseek-v4-flash, leave thinking enabled, begin with high effort, and record model, system_fingerprint, cache use, reasoning and output tokens, finish reason, latency, and retries. Public evidence yields no production winner. DeepSeek has not yet released the harness used for its public code-agent results, two test sets are internal, and a current same-harness comparison across vendors is absent. Keep the current production backend available and admit traffic only when owned repository tasks meet prespecified quality, repair, latency, and cost gates.

Start multimodal work with Gemini 3.6 Flash, whose published input surface covers text, images, video, audio, and PDF while its managed tools include search, code execution, file search, and functions. Use Claude Sonnet 5 as the production comparator when a fixed identity or an eligible Zero Data Retention arrangement matters. Existing GPT, Grok, GLM, Kimi, MiniMax, and Qwen routes remain in place and enter the same owned evaluation when their current job or procurement boundary makes them relevant.

JobFirst POC or current defaultObject to freezeProduction admission and exit
Low-cost, text-only coding agentdeepseek-v4-flash, currently version 0731Prompt, tool schema, task commit, effort, maximum turns, and timeout; record returned model and fingerprintCanary after owned success and repair gates pass; route back on fingerprint change, failure increase, or p95 breach
One agent must consume images, audio, video, or PDFGemini 3.6 FlashStable model ID, media fixtures, tool permissions, search, and code-execution switchesValidate extraction, citations, and tool side effects by modality; retain the old route for any regressing modality
Fixed snapshot, governed data handling, or ZDR comparisonClaude Sonnet 5claude-sonnet-5 snapshot, organizational data settings, region, logs, and contract conditionsVendor eligibility and owned evaluation both pass; remeasure after tokenizer or contract changes
Data must stay in an owned boundary or serving must be customizedOfficial MIT-licensed 0731 weights; a small team uses Qwen3.6-35B-A3B first as a single-card controlWeight revision, container digest, topology, KV configuration, parsers, sampling, routing, and capacityThe same task-quality gate passes and successful jobs per node-hour cover total cost; return to a managed API when capacity, reliability, or governance evidence expires

The map turns a vague search for the “strongest model” into four purchasable jobs. Price and published agent results earn 0731 the first low-cost coding trial. Modality, fixed identity, data location, and operational ownership route the other jobs elsewhere. Each job has one first test; a backup enters when the first candidate lacks a required capability, hits an exit condition, or loses on the common owned task set.

1.1 An engineering-ready API, local-node, acceptance, and rollback ticket

Managed API trial

Entry point
Use Chat Completions or Responses at https://api.deepseek.com and request deepseek-v4-flash. Thinking is enabled by default. Begin at reasoning_effort=high; send a job to max only after its long-reasoning budget is approved.
Price clock
At this report's cutoff, cache-hit input, uncached input, and output cost $0.0028, $0.14, and $0.28 per million tokens. The pricing page also announces a future peak period at twice the current rates, with its effective date still pending. Refresh prices on the production-admission day, put the live rate into the budget gate, and withdraw the candidate on a breach.
Request identity
Freeze the system prompt, tools, temperature, top_p, maximum output, timeout, maximum turns, and task commit. Use a pseudonymous internal user_id with no personal information and retain its internal audit mapping.
Response receipt
Store the returned model, system_fingerprint, finish reason, reasoning/output tokens, cache hit/miss, first-token and total latency, tool-argument validation, retries, and final tests. The documented request ID moves; the response identity reveals backend changes.
Failure handling
A length finish enters context or output-budget repair. A content_filter finish enters human review or an approved alternate route. An insufficient_system_resource finish gets bounded backoff retries and then the old backend. Local schema and authorization checks run before any tool side effect.

Local 0731 trial

Frozen artifact
Use official revision 7872f01b…. In SOSEC's captured manifest, 48 safetensors total 166,886,535,336 bytes, or 155.43 GiB. The repository UI presents the same order of magnitude as about 167 decimal GB.
Starting topology
The official 0731 model card shows a single-node 4×GB300 vLLM command. The general V4 Flash recipe linked from that card now names 0731 as its default and also covers four-accelerator H200, B200, and B300 replicas and an eight-H200 4+4 prefill/decode layout. Confirm interconnect, memory, drivers, kernels, and container support on the exact provider SKU before purchase.
Runtime settings
Enable the DeepSeek V4 tokenizer, reasoning parser, tool parser, and DSpark configuration. For agent work, use the publisher's temperature=1.0 and top_p=0.95 starting point. High/max needs at least a 384K maximum model length; full 1M needs its own KV-cache, concurrency, and first-token test.
Supply chain
Pin digests for weights, encoding code, remote code, container, vLLM or SGLang, and the driver. Mount the artifact store read-only, give the service a least-privilege identity, separate model serving from tool execution, and require positive, negative, and rollback controls on an isolated node before promotion.

Common-task acceptance

Task set
Sample real defect repairs, cross-file features, test completion, migrations, dependency updates, and log diagnosis. Save the starting commit, allowed tools, expected tests, and forbidden side effects. Run each candidate at least twice so one sample does not decide procurement.
Outcome
The primary unit is a task that passes without a human rewrite. Also count human repair lines, extra turns, invalid tools, incomplete work, refusals, test escapes, and rollback. Generated line count, tone, and model self-assessment stay outside the success numerator.
SOSEC starting gate
Critical-task success may trail the incumbent by at most two percentage points, p95 completion time may regress by at most 20%, and all-in cost per successful task must improve by at least 30%. These are starting recommendations; each owner calibrates final values to business risk and capacity.
Local additions
Record available GPU-hours, successful jobs per node-hour, queueing, cold start, reload, GPU and host memory, KV eviction, network, storage, power, on-call work, and recovery. Both API and local routes are charged to successful tasks.

Canary and withdrawal

Routing
Mirror side-effect-free work first, then admit a small share of reversible jobs. Fix the routing key to task type, repository, and user cohort. Keep a retry on the original backend so a silent mid-task model switch cannot contaminate one result.
Automatic withdrawal
A changed fingerprint or local artifact digest, repeated failures, tool-schema errors, a critical test escape, or a cost/p95 breach closes the candidate route and sends new work to the incumbent. Compensate any file or external action already performed.
Recovery proof
After rollback, run a fixed positive set and a denied or unauthorized-tool negative set to confirm that routing, logs, billing, and permissions have returned. Retain candidate logs and failed artifacts until root cause, bill, and vendor case are closed.
Exit
End the route after two corrective cycles still miss the critical success gate or when low token price creates multitenancy, operations, or license duties the owner cannot sustain. The next release reuses the task set and gate, while the old conclusion expires.

2 Three identities now line up: API alias, 0731 backend, and open weights

DeepSeek documents deepseek-v4-flash as the request ID. On August 1, 2026, the pricing page mapped it to DeepSeek-V4-Flash-0731, while the API response exposed the actual model and a system_fingerprint representing backend configuration. Procurement, evaluation, and incident records need all three. The request alias identifies the entry point, the returned model identifies the served version, and the fingerprint catches a backend change under the same public name.

The 0731 open weights entered DeepSeek's official Hugging Face organization and V4 collection later the same release day. The model card calls this the official Flash release that supersedes Preview, retains the V4 Flash structure, and includes a DSpark speculative-decoding module. The repository and weights carry the MIT license. SOSEC pinned revision 7872f01b1d1fe23eabc4c98b48bffcef5a386062 after the repository's 2026-08-01 03:07:41 UTC update. API and local trials can therefore target the same named 0731 product version, while sampling, encoding, kernels, quantization, scheduling, and tool wrappers remain measurable implementation differences.

The older deepseek-ai/DeepSeek-V4-Flash repository continues to represent the Preview route. DeepSeek-V4-Flash-DSpark combines Preview weights with the speculative module. The new 0731 repository contains the re-post-trained official weights and DSpark. Reproduction records must name a complete repository and revision; the short label “V4 Flash” spans three distinct artifacts.

A warm hand-drawn blue-steel ticket machine feeds a replaceable API ticket into a backend machine carrying a fingerprint gauge and sealed receipt, while a separate rail carries a crate of open weights
Figure 1. The request alias, served backend, and local weights have separate identities. API trials retain the returned model and fingerprint; local trials retain repository revision, file digests, and serving stack so both results survive a later version change.

3 What public specifications can compare: price, modality, context, and lifecycle

The comparison admits current candidates that had an official entry point, callable identity, and public price on August 1, 2026. One representative per provider stays closest to general agent or coding work, preventing a single catalog from dominating the table. Rates remain in the publisher's per-million-token unit; CNY stays in CNY instead of receiving an ephemeral exchange-rate conversion. The reference job uses 200K uncached input and 20K output. It makes rates legible and leaves tokenizer, reasoning tokens, retries, tool fees, and success rate to the owned replay.

Model and freeze methodPublished modality and windowCache / input / output rateExample billSOSEC role and missing gate
DeepSeek V4 Flash 0731
deepseek-v4-flash moves; retain returned version and fingerprint; pin a revision locally
Text-only; 1M input context; 384K maximum output; thinking by default; Chat, Responses, Anthropic compatibility, and function tools$0.0028 / $0.14 / $0.28$0.0336First low-cost coding trial; public beta, moving alias, owned quality and stability still required
GPT-5.6 Terra
gpt-5.6-terra
Text output with text and image input; 1.05M / 128K; Responses, functions, structured output, and several hosted tools$0.20 / $2 / $12; requests above 272K input bill 2× input and 1.5× output across the full request$0.64Mature tool and vision comparator; long-prompt tiers, tool fees, and real cache-write billing enter the replay
Claude Sonnet 5
claude-sonnet-5 is a pinned snapshot under Anthropic's version rules
Text and image input, text output; 1M / 128K; adaptive thinking and tools; eligible ZDR arrangementsIntroductory $2 / $10 through August 31; standard $3 / $15; caching follows separate Anthropic rules$0.60 intro; $0.90 standardFixed-identity and data-governance comparator; the new tokenizer changes old workload token counts
Gemini 3.6 Flash
gemini-3.6-flash stable
Text, image, video, audio, and PDF input; 1,048,576 / 65,536; Search, Maps, code execution, file search, and functions$0.15 / $1.50 / $7.50; cache storage separate$0.45First multimodal and hosted-tool trial; grounding, tool, and storage fees follow the job
GLM 5.2
glm-5.2
Text-only; 1M / 128K; thinking, functions, structured output, caching, and MCP; MIT weights$0.26 / $1.40 / $4.40$0.368Managed plus open-weight comparator; a complete local route begins at eight-card scale
Kimi K3
kimi-k3
Native multimodal; 1M; always-thinking; OpenAI and Anthropic compatibility; open weights$0.30 / $3 / $15$0.90Very large open multimodal agent comparator; local topology, license, and parser costs are substantial
Grok 4.5
freeze a dated ID; the bare name and -latest move
Vendor's current code and general recommendation; 500K; adjustable reasoning; tools per xAI documentationShort context $0.30 / $2 / $6; long context $0.60 / $4 / $12$0.52–$1.04Pricing says long context starts at ≥200K; the model page says above 200K. This example is exactly 200K, so a live bill closes the boundary conflict
MiniMax M3
MiniMax-M3; the pricing page leads parts of the API guide
Native multimodal; 1M weight model; managed long-context tiers; open weights under a community licenseCurrent ≤512K: $0.06 / $0.30 / $1.20; 512K–1M: $0.12 / $0.60 / $2.40$0.084Low-price multimodal candidate; close pricing/API-guide synchronization, long-context availability, and commercial authorization first
Qwen3.7 Max
qwen3.7-max-2026-06-08 pinned snapshot
Text, image, and video input with text output; 1M; thinking/non-thinking, functions, built-in tools, and structured outputCache not listed in the standard price table / CNY 12 / CNY 36; one tier for 0<input≤1M; regional prices varyCNY 3.12Pinned Alibaba Cloud snapshot comparator; select region, currency, and exact snapshot before arithmetic

DeepSeek's rate is the visual outlier and the rate most likely to make procurement skip a trial. One successful task can consume several model calls, tool returns, and repair loops. Failed tasks also incur charges. Modality changes input metering, and hosted search, file retrieval, code execution, and computer use may add per-call cost. The comparable unit is all-in cost for one accepted task, combining tokens, tools, retries, and human correction under one denominator.

Lifecycle controls whether a result can be reproduced. DeepSeek accepts a moving public API name and therefore needs response identity. Anthropic explicitly treats the dateless Sonnet 5 ID as a pinned snapshot. Google labels 3.6 Flash stable. xAI's bare and -latest names move. Qwen exposes both a moving alias and dated snapshots. A model router converts these differences into policy: fixed evaluations use a fixed ID, local weights use a revision, and moving endpoints retain returned identity and automatically rerun regressions on change.

4 DeepSeek's release results admit 0731 to the first trial; production ranking still needs common evidence

DeepSeek published nine agent results for 0731: 82.7 on Terminal Bench 2.1, 54.2 on NL2Repo, 76.7 on Cybergym, 54.4 on DeepSWE, 70.3 on Toolathlon-Verified, 25.2 on Agents’ Last Exam, 25.1 on AutomationBench Public, 68.7 on DSBench-FullStack, and 59.6 on DSBench-Hard. Every result rises materially from Preview in the same release table; DeepSWE moves from 7.3 to 54.4. Those numbers explain why 0731 receives the first coding trial.

The model card also states the evidence boundary. Public code-agent tasks used the minimal mode of DeepSeek Harness, marked “to be released,” at max effort with temperature=1.0 and top_p=0.95. DSBench-FullStack and DSBench-Hard are internal. The release table includes GLM and Opus results, while the current page does not supply a single container, tool definition, task snapshot, turn cap, rerun policy, and complete log set for every compared model.

A cross-vendor league table would compress those implementation differences into one decimal column. A coding agent's outcome comes from the model, harness, shell authority, tests, context construction, retries, and stop policy together. SOSEC preserves the publisher results for their supported decision: choosing which candidate to evaluate. Production order comes from the common owned replay. A future fixed Harness and log release can raise reproducibility while the original result retains its version and settings.

4.1 A small, hard task set sits closer to procurement than a thousand generic questions

Stratify the task set by business failure. Small repository repairs test localization and test discipline. Cross-file features test planning, context, and tool loops. Migrations and dependency upgrades test documentation, compatibility, and rollback. Log diagnosis tests evidence use. Security tasks test whether the agent crosses authority or avoids a failing control. Each class keeps a positive control known to pass, a negative control that must be rejected, and one resource-failure path.

The grader accepts runnable outcomes. A candidate edits an isolated worktree, and native project tests, static checks, and task-specific assertions determine success. Human review then checks authority, scope, maintainability, and hidden side effects. A completion claim, higher line count, or a permissive test authored by the model adds nothing to the score. On failure, record the first divergence: missing context, faulty plan, invalid tool arguments, test escape, resource stop, refusal, or repair loop.

Run each model at least twice per task and rotate order so caching, provider load, and warm repository state cannot become a fixed advantage. Retain complete use and price, then calculate accepted-task cost, human minutes, and p95. This answers the purchasing question: for this team and repository, does 0731 deliver mergeable work faster and more cheaply than the incumbent?

5 From per-million tokens to total cost: cheap API service creates a high utilization bar for local nodes

Begin with a bill a person can read. The reference job uses 200K uncached input and 20K output. DeepSeek 0731 costs 0.2×0.14 + 0.02×0.28 = $0.0336. At the same token counts, MiniMax M3's current ≤512K rate costs $0.084, GLM 5.2 $0.368, Gemini 3.6 Flash $0.45, introductory Claude Sonnet 5 $0.60, GPT-5.6 Terra $0.64, and both Kimi K3 and standard Sonnet 5 $0.90. Grok 4.5 lands in a $0.52–$1.04 range because two official pages disagree at the exact 200K boundary. Qwen's documented global tier produces CNY 3.12.

The Grok range comes from a locatable documentation conflict. Its model detail says long-context pricing applies above 200K; the pricing matrix says ≥200K and doubles input, cache, and output rates together. The reference input is exactly 200K, so prose cannot settle the boundary. The POC retains usage and the actual bill as the procurement record. The other eight candidates have determinate tiers at 200K input: Terra remains below its >272K trigger; Sonnet 5 and Gemini 3.6 Flash publish one rate across their windows; MiniMax stays in its ≤512K tier; Qwen's snapshot covers 0<input≤1M; and the current DeepSeek, GLM, and Kimi pages publish no separate 200K step.

The arithmetic stays deliberately narrow. Tokenizers differ, so 200K identical characters become different token counts. Reasoning tokens, media, search, and tools alter the bill. One coding task commonly makes several calls. Caching can rewrite the result: DeepSeek cache-hit input costs only $0.0028 per million, so a stable shared prefix lowers the API side further. The evaluator retains original bytes, provider token counts, cache hit/miss, and tool charges; the example rate never substitutes for measured cost.

The local formula is cost per accepted task = (GPU + CPU/memory + storage + network + power/facility + platform + operations + redundancy per hour) / accepted tasks per node-hour. Accelerator utilization shows that equipment is busy. Accepted jobs show that it is delivering. Queueing, long-context KV cache, failed retries, weight updates, node failures, and idle valleys raise the numerator or lower the denominator.

Runpod's public Pod page, updated July 27, 2026, lists H200 at $4.39, B200 at $5.89, and B300 at $7.39 per GPU-hour. Four cards over a 730-hour month produce raw GPU rent of $12,818.80, $17,198.80, and $21,578.80. These numbers exclude CPU, memory, storage, network, image transfer, on-call labor, and redundancy. A quoted Pod must also place all four cards in a node with suitable interconnect. A B300 rate cannot stand in for a GB300 quote.

Dividing by the $0.0336 DeepSeek reference bill, 4×H200 needs about 523 accepted reference jobs per node-hour merely to reach raw rent, 4×B200 about 701, and 4×B300 about 880. These are required capacity thresholds; throughput must be measured at the target context, effort, output, concurrency, and quality gate. A second node, staff, and idle capacity push the threshold higher. API cache hits move it farther away.

The monthly view is even clearer. One hundred reference jobs per day cost about $100.80 on the DeepSeek API. One thousand cost about $1,008. Ten thousand cost about $10,080. Raw four-H200 rent approaches the third case, while the owner still must show that the node can deliver 10,000 accepted jobs every day and absorb peaks, failure, and upgrades. Data residency, egress restrictions, custom serving, or predictable reserved capacity can create an independent reason for local deployment; document that non-price value in the purchase rationale and exit gate.

6 Local deployment is real: 0731 starts at four-card scale, while one card establishes the control baseline

With the 0731 weight release, local and API routes can finally target the same product version. The 155.43 GiB figure covers repository weight files. Serving also consumes runtime, communication buffers, DSpark, KV cache, batches, and context. Full 1M and 384K high/max output require separate memory and latency measurements. The official four-GB300 example is an executable starting point; a purchase still follows the target concurrency and sequence-length calculation.

Local candidateFrozen weight and licensePublished artifact scaleOfficial or framework starting topologyPractical role and public cost boundary
DeepSeek V4 Flash 07317872f01b…; MIT; official 0731 plus DSpark155.43 GiB from SOSEC's manifest sum; about 167 GB in the HF UIModel card: 4×GB300; the current vLLM V4 Flash recipe defaults to 0731 and gives four-card H200/B200/B300 replicas plus 8×H200 4+4 PD; high/max at least 384KFirst exact-model local trial; 4×H200 raw rent $17.56/hour, with topology and full TCO requiring quote and measurement
Qwen3.6-35B-A3B FP8Official FP8; Apache-2.0; 35B total/3B activeAbout 34.89 GiB in the captured manifest; 262K native context with YaRN extensionvLLM provides a single H100/H200 or MI300X-class route; NVFP4 reaches RTX Pro 6000 and DGX Spark routesSingle-card serving, tool, and observability control for a small team; it carries no 0731 quality claim; 1×H200 raw rent $3,204.70/month
GLM 5.2Official weights; MITAbout 1,403.19 GiB in the captured BF16 manifestPractical FP8 default on 8×H200/H20; full 1M reference path on 8×B200 with FP8 KV cacheHeavier open coding/agent comparator; 8×H200 raw rent $25,637.60/month, with quantized quality and throughput still measured
MiniMax M3MiniMax Community License; legal review for commercial attribution and the stated revenue authorization boundaryAbout 854 GB BF16 in the captured official repository manifest; the card confirms 427/428B total and 23B activeNative multimodal; validate hardware and long-context configuration against current serving documentation and supplier quoteProcurement evaluation when local multimodal input matters; close license, documentation sync, and node cost first
Kimi K3Official weights; Kimi K3 License; 2.8T total/104B activeAbout 1,453.74 GiB in the captured manifestvLLM recipe requires at least 8×GB300; real production adds multi-node networking and tool-parser governanceVery large native multimodal local comparator; GB300 and multi-node cost requires a formal quote
A warm hand-drawn four-stage local deployment workshop grows from a one-card bench to a four-card cabinet, an eight-card rack with a KV reservoir, and two interconnected racks, with one, four, eight, and paired eight-card inventories shown below
Figure 2. Loading weights is the first cell. One card validates the serving contract; four enters the official 0731 scale; eight can separate prefill and decode or carry larger contexts; multi-node service adds networking, scheduling, cooling, redundancy, and independent on-call duty. The balance above compares API and local service only through accepted jobs.

6.1 Beyond four accelerators, the owner must build an internal model service

Download first solves provenance, resumption, verification, and mirroring for a roughly 167 GB object. Weights and code enter a read-only artifact repository, and production nodes pull from an internal mirror. Each release binds the Hugging Face revision, per-file digests, encoding directory, container, and driver. The DeepSeek card ships dedicated encoding code instead of a Jinja chat template. vLLM's DeepSeek V4 tokenizer, reasoning parser, and tool parser must be accepted with their fixed versions. Parser drift can change tool behavior while weights stay constant.

The service layer defines tenant isolation, queueing, fairness, rate limits, timeouts, cancellation, interrupted streams, and the resource-exhaustion contract. Long contexts consume KV cache, so one huge request can delay small requests. Partition scheduling by context and effort, cap in-flight tokens per tenant, and expose queue, prefill, decode, KV eviction, and failure metrics. Enable the complete one-million window only after real tasks show gain; size the default window from observed useful input.

Tool execution remains an application control. The model service emits parsed intent. A policy layer checks principal, tool, arguments, repository, path, and side-effect budget. An executor runs in a separate sandbox and returns a size- and sensitivity-bounded result. Multitenant use also validates prompt and cache isolation, log redaction, artifact access, and management-interface authentication. Local deployment brings the data path under organizational control and assigns these duties to that organization.

High availability needs a second capacity source or a managed fallback. Maintenance, driver failure, weight reload, and GPU faults can remove the whole single-node route. Two fully provisioned nodes roughly double most fixed costs. An owner can use one primary plus a smaller degradation pool, two availability-zone nodes, or local-first routing with DeepSeek API fallback. Each choice states whether data may leave the boundary, which jobs survive degradation, whether caches remain compatible, and how weight and log continuity is verified before returning.

6.2 Buying hardware, renting nodes, and staying on API make different commitments

API service fits variable demand, modest task volume, and rapidly changing models. The provider operates accelerators, kernels, and availability; the caller still owns evaluation, data governance, tool authority, and rollback. DeepSeek's rate stretches this route across a large volume range. The moving alias and public-beta state require stricter response identity and regression handling.

Hourly four-card rental fits data-location experiments, temporary capacity, and serving customization. The buyer can fix weights and container and release the node later. Persistent downloads, idle time, and cross-zone storage add charges. Before contracting, confirm same-host interconnect, shared storage, image pull time, quota, and replacement behavior. A public per-GPU-hour rate creates a usable node only when the topology qualifies.

Owned hardware fits stable high utilization and organizations with an existing facility and GPU operations. Total ownership includes depreciation, capital, warranty, racks, power, cooling, switching, spares, staff, and utilization opportunity cost. A three-year amortization cannot erase the possibility that the preferred model changes in six months. The exit path lets the node carry other inference work and preserves the managed-API compatibility layer. Accepted-job throughput, policy value, and reusable infrastructure together decide the purchase.

7 Close the decision: treat 0731 as a reversible product change

For most teams, the next action is straightforward: keep the current production model and run the same real coding tasks through the managed DeepSeek API. Its $0.0336 reference bill makes quality, tool, latency, and failure evidence cheap to collect before reserving a four-card node. Add Gemini 3.6 Flash when the task is multimodal and Sonnet 5 when fixed identity or ZDR eligibility matters. Other models enter when the current workload or procurement boundary calls for them.

The 0731 weight release changes the data boundary and customization space. A public revision, MIT license, and serving documentation create an executable local POC. Four-card hardware, 155.43 GiB of weights, KV cache, DSpark, parsers, and model-service operations form the actual solution. Hundreds or thousands of long jobs per day commonly retain lower direct cost through the API. Sustained high throughput, data residency, or an existing GPU platform can move local service to the first trial.

The production answer belongs in a versioned decision record: task-set revision, candidate identity, success, human repair, p50/p95, all-in cost, tool side effects, data path, capacity, recovery, and rollback result. Canary 0731 when it passes and withdraw it when a critical gate fails. Repeat the same contract for the next release. This versioned contract becomes the durable system for choosing models; a static leaderboard expires with the next alias or backend update.

Research record

8Evidence, objects, and sources

The material below preserves the identifiers and references used in this report.

8.1Research objects

Products, actors, techniques, affected objects, and control points discussed in the report.

DeepSeek API IDdeepseek-v4-flash

Moving request ID mapped to DeepSeek-V4-Flash-0731 on the 2026-08-01 pricing page

0731 weight revision7872f01b1d1fe23eabc4c98b48bffcef5a386062

Official Hugging Face revision pinned for SOSEC's blobs-manifest read

Weight bytes166886535336

Manifest sum of 48 safetensors, equal to 155.43 GiB

Reference job200K uncached input + 20K output

Billing arithmetic that carries no cross-tokenizer quality or throughput claim

8.2Event chronology

  1. DeepSeek V4 Preview weights released

    V4 Flash Preview established the 284B/13B, one-million-context, MIT-weight route; 0731 later re-post-trained the same architecture.

  2. V4 Flash 0731 API enters public beta

    DeepSeek updated the moving API alias, agent results, Responses API, and rate, with explicit Harness and internal-set boundaries.

  3. Official 0731 weights enter Hugging Face

    The official repository and V4 collection added 0731; SOSEC subsequently pinned the revision and calculated weight bytes.

  4. SOSEC closes the public-information and cost snapshot

    API rates, model identity, local topologies, public GPU rent, and comparator entry points were reviewed in one snapshot.

8.3Sources and material

  1. DeepSeek API update noticehttps://api-docs.deepseek.com/updates/
  2. DeepSeek models, features, pricing, and concurrencyhttps://api-docs.deepseek.com/quick_start/pricing/
  3. DeepSeek Chat Completions parameters, finish reasons, usage, and fingerprinthttps://api-docs.deepseek.com/api/create-chat-completion
  4. DeepSeek rate limits and user isolationhttps://api-docs.deepseek.com/quick_start/rate_limit
  5. Official DeepSeek V4 Flash 0731 card, evaluation, deployment, and MIT licensehttps://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
  6. Pinned official 0731 weight revisionhttps://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/tree/7872f01b1d1fe23eabc4c98b48bffcef5a386062
  7. Official DeepSeek V4 model collectionhttps://huggingface.co/collections/deepseek-ai/deepseek-v4
  8. General vLLM V4 Flash deployment recipe, currently defaulting to 0731https://recipes.vllm.ai/deepseek-ai/DeepSeek-V4-Flash
  9. SGLang DeepSeek V4 deployment and tuning cookbookhttps://docs.sglang.io/cookbook/autoregressive/DeepSeek/DeepSeek-V4
  10. Official OpenAI GPT-5.6 Terra model pagehttps://developers.openai.com/api/docs/models/gpt-5.6-terra
  11. Current Anthropic Claude models, pricing, and contexthttps://platform.claude.com/docs/en/about-claude/models/overview
  12. Anthropic model ID and version ruleshttps://platform.claude.com/docs/en/about-claude/models/model-ids-and-versions
  13. Anthropic full pricing and Sonnet 5 full-window standard rateshttps://platform.claude.com/docs/en/about-claude/pricing
  14. Anthropic Sonnet 5 tokenizer, context, and ZDR noteshttps://platform.claude.com/docs/en/about-claude/models/whats-new-sonnet-5
  15. Anthropic API data-retention and ZDR eligibility boundarieshttps://platform.claude.com/docs/en/manage-claude/api-and-data-retention
  16. Official Google Gemini 3.6 Flash model pagehttps://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash
  17. Official Google Gemini API pricinghttps://ai.google.dev/gemini-api/docs/pricing
  18. Z.ai GLM 5.2 model and interfacehttps://docs.z.ai/guides/llm/glm-5.2
  19. Official Z.ai model pricinghttps://docs.z.ai/guides/overview/pricing
  20. Official GLM 5.2 open weightshttps://huggingface.co/zai-org/GLM-5.2
  21. vLLM GLM 5.2 deployment recipehttps://recipes.vllm.ai/zai-org/GLM-5.2
  22. Kimi API models and pricinghttps://platform.kimi.ai/
  23. Official Kimi K3 card and licensehttps://huggingface.co/moonshotai/Kimi-K3
  24. vLLM Kimi K3 deployment recipehttps://recipes.vllm.ai/moonshotai/Kimi-K3
  25. Alibaba Cloud Model Studio text models and snapshotshttps://help.aliyun.com/en/model-studio/text-generation-model
  26. Alibaba Cloud Qwen3.7 Max snapshot and visual-capability release recordhttps://help.aliyun.com/en/model-studio/newly-released-models
  27. Official Alibaba Cloud Model Studio pricinghttps://help.aliyun.com/en/model-studio/model-pricing
  28. Current xAI models, identities, and priceshttps://docs.x.ai/developers/models
  29. xAI Grok 4.5 model detail and 200K boundary wordinghttps://docs.x.ai/developers/models/grok-4.5
  30. xAI model pricing and ≥200K long-context matrixhttps://docs.x.ai/developers/pricing
  31. Current MiniMax API pricinghttps://platform.minimax.io/subscribe/token-plan?tab=api-enterprise
  32. Official MiniMax M3 weights and licensehttps://huggingface.co/MiniMaxAI/MiniMax-M3
  33. Full MiniMax M3 Community Licensehttps://huggingface.co/MiniMaxAI/MiniMax-M3/blob/main/LICENSE
  34. Official Qwen3.6-35B-A3B model cardhttps://huggingface.co/Qwen/Qwen3.6-35B-A3B
  35. Official Qwen3.6-35B-A3B FP8 weightshttps://huggingface.co/Qwen/Qwen3.6-35B-A3B-FP8
  36. vLLM single-card Qwen3.6-35B-A3B recipehttps://recipes.vllm.ai/Qwen/Qwen3.6-35B-A3B
  37. Runpod public GPU Pod pricinghttps://www.runpod.io/pricing