Research
Should You Switch to DeepSeek V4 Flash 0731? Model, Cost, and Self-Hosting Comparison — August 2026
As of August 1, 2026, public evidence yields no production winner: test deepseek-v4-flash first for low-cost text coding and record its 0731 fingerprint, Gemini 3.6 Flash first for multimodal work, and Claude Sonnet 5 for fixed-snapshot or ZDR needs; the MIT 0731 weights are public, while local adoption starts after a four-accelerator TCO trial and owned tasks pass with the incumbent ready for rollback.

In this article
1 The decision now: test 0731 first for low-cost coding and wait for owned results before production
Verdict: DeepSeek V4 Flash 0731 is the first August 2026 POC for a low-cost, text-only coding agent. Send requests to deepseek-v4-flash, leave thinking enabled, begin with high effort, and record model, system_fingerprint, cache use, reasoning and output tokens, finish reason, latency, and retries. Public evidence yields no production winner. DeepSeek has not yet released the harness used for its public code-agent results, two test sets are internal, and a current same-harness comparison across vendors is absent. Keep the current production backend available and admit traffic only when owned repository tasks meet prespecified quality, repair, latency, and cost gates.
Start multimodal work with Gemini 3.6 Flash, whose published input surface covers text, images, video, audio, and PDF while its managed tools include search, code execution, file search, and functions. Use Claude Sonnet 5 as the production comparator when a fixed identity or an eligible Zero Data Retention arrangement matters. Existing GPT, Grok, GLM, Kimi, MiniMax, and Qwen routes remain in place and enter the same owned evaluation when their current job or procurement boundary makes them relevant.
| Job | First POC or current default | Object to freeze | Production admission and exit |
|---|---|---|---|
| Low-cost, text-only coding agent | deepseek-v4-flash, currently version 0731 | Prompt, tool schema, task commit, effort, maximum turns, and timeout; record returned model and fingerprint | Canary after owned success and repair gates pass; route back on fingerprint change, failure increase, or p95 breach |
| One agent must consume images, audio, video, or PDF | Gemini 3.6 Flash | Stable model ID, media fixtures, tool permissions, search, and code-execution switches | Validate extraction, citations, and tool side effects by modality; retain the old route for any regressing modality |
| Fixed snapshot, governed data handling, or ZDR comparison | Claude Sonnet 5 | claude-sonnet-5 snapshot, organizational data settings, region, logs, and contract conditions | Vendor eligibility and owned evaluation both pass; remeasure after tokenizer or contract changes |
| Data must stay in an owned boundary or serving must be customized | Official MIT-licensed 0731 weights; a small team uses Qwen3.6-35B-A3B first as a single-card control | Weight revision, container digest, topology, KV configuration, parsers, sampling, routing, and capacity | The same task-quality gate passes and successful jobs per node-hour cover total cost; return to a managed API when capacity, reliability, or governance evidence expires |
The map turns a vague search for the “strongest model” into four purchasable jobs. Price and published agent results earn 0731 the first low-cost coding trial. Modality, fixed identity, data location, and operational ownership route the other jobs elsewhere. Each job has one first test; a backup enters when the first candidate lacks a required capability, hits an exit condition, or loses on the common owned task set.
1.1 An engineering-ready API, local-node, acceptance, and rollback ticket
Managed API trial
- Entry point
- Use Chat Completions or Responses at
https://api.deepseek.comand requestdeepseek-v4-flash. Thinking is enabled by default. Begin atreasoning_effort=high; send a job tomaxonly after its long-reasoning budget is approved. - Price clock
- At this report's cutoff, cache-hit input, uncached input, and output cost $0.0028, $0.14, and $0.28 per million tokens. The pricing page also announces a future peak period at twice the current rates, with its effective date still pending. Refresh prices on the production-admission day, put the live rate into the budget gate, and withdraw the candidate on a breach.
- Request identity
- Freeze the system prompt, tools, temperature,
top_p, maximum output, timeout, maximum turns, and task commit. Use a pseudonymous internaluser_idwith no personal information and retain its internal audit mapping. - Response receipt
- Store the returned
model,system_fingerprint, finish reason, reasoning/output tokens, cache hit/miss, first-token and total latency, tool-argument validation, retries, and final tests. The documented request ID moves; the response identity reveals backend changes. - Failure handling
- A
lengthfinish enters context or output-budget repair. Acontent_filterfinish enters human review or an approved alternate route. Aninsufficient_system_resourcefinish gets bounded backoff retries and then the old backend. Local schema and authorization checks run before any tool side effect.
Local 0731 trial
- Frozen artifact
- Use official revision
7872f01b…. In SOSEC's captured manifest, 48 safetensors total 166,886,535,336 bytes, or 155.43 GiB. The repository UI presents the same order of magnitude as about 167 decimal GB. - Starting topology
- The official 0731 model card shows a single-node 4×GB300 vLLM command. The general V4 Flash recipe linked from that card now names 0731 as its default and also covers four-accelerator H200, B200, and B300 replicas and an eight-H200 4+4 prefill/decode layout. Confirm interconnect, memory, drivers, kernels, and container support on the exact provider SKU before purchase.
- Runtime settings
- Enable the DeepSeek V4 tokenizer, reasoning parser, tool parser, and DSpark configuration. For agent work, use the publisher's
temperature=1.0andtop_p=0.95starting point. High/max needs at least a 384K maximum model length; full 1M needs its own KV-cache, concurrency, and first-token test. - Supply chain
- Pin digests for weights, encoding code, remote code, container, vLLM or SGLang, and the driver. Mount the artifact store read-only, give the service a least-privilege identity, separate model serving from tool execution, and require positive, negative, and rollback controls on an isolated node before promotion.
Common-task acceptance
- Task set
- Sample real defect repairs, cross-file features, test completion, migrations, dependency updates, and log diagnosis. Save the starting commit, allowed tools, expected tests, and forbidden side effects. Run each candidate at least twice so one sample does not decide procurement.
- Outcome
- The primary unit is a task that passes without a human rewrite. Also count human repair lines, extra turns, invalid tools, incomplete work, refusals, test escapes, and rollback. Generated line count, tone, and model self-assessment stay outside the success numerator.
- SOSEC starting gate
- Critical-task success may trail the incumbent by at most two percentage points, p95 completion time may regress by at most 20%, and all-in cost per successful task must improve by at least 30%. These are starting recommendations; each owner calibrates final values to business risk and capacity.
- Local additions
- Record available GPU-hours, successful jobs per node-hour, queueing, cold start, reload, GPU and host memory, KV eviction, network, storage, power, on-call work, and recovery. Both API and local routes are charged to successful tasks.
Canary and withdrawal
- Routing
- Mirror side-effect-free work first, then admit a small share of reversible jobs. Fix the routing key to task type, repository, and user cohort. Keep a retry on the original backend so a silent mid-task model switch cannot contaminate one result.
- Automatic withdrawal
- A changed fingerprint or local artifact digest, repeated failures, tool-schema errors, a critical test escape, or a cost/p95 breach closes the candidate route and sends new work to the incumbent. Compensate any file or external action already performed.
- Recovery proof
- After rollback, run a fixed positive set and a denied or unauthorized-tool negative set to confirm that routing, logs, billing, and permissions have returned. Retain candidate logs and failed artifacts until root cause, bill, and vendor case are closed.
- Exit
- End the route after two corrective cycles still miss the critical success gate or when low token price creates multitenancy, operations, or license duties the owner cannot sustain. The next release reuses the task set and gate, while the old conclusion expires.
2 Three identities now line up: API alias, 0731 backend, and open weights
DeepSeek documents deepseek-v4-flash as the request ID. On August 1, 2026, the pricing page mapped it to DeepSeek-V4-Flash-0731, while the API response exposed the actual model and a system_fingerprint representing backend configuration. Procurement, evaluation, and incident records need all three. The request alias identifies the entry point, the returned model identifies the served version, and the fingerprint catches a backend change under the same public name.
The 0731 open weights entered DeepSeek's official Hugging Face organization and V4 collection later the same release day. The model card calls this the official Flash release that supersedes Preview, retains the V4 Flash structure, and includes a DSpark speculative-decoding module. The repository and weights carry the MIT license. SOSEC pinned revision 7872f01b1d1fe23eabc4c98b48bffcef5a386062 after the repository's 2026-08-01 03:07:41 UTC update. API and local trials can therefore target the same named 0731 product version, while sampling, encoding, kernels, quantization, scheduling, and tool wrappers remain measurable implementation differences.
The older deepseek-ai/DeepSeek-V4-Flash repository continues to represent the Preview route. DeepSeek-V4-Flash-DSpark combines Preview weights with the speculative module. The new 0731 repository contains the re-post-trained official weights and DSpark. Reproduction records must name a complete repository and revision; the short label “V4 Flash” spans three distinct artifacts.
3 What public specifications can compare: price, modality, context, and lifecycle
The comparison admits current candidates that had an official entry point, callable identity, and public price on August 1, 2026. One representative per provider stays closest to general agent or coding work, preventing a single catalog from dominating the table. Rates remain in the publisher's per-million-token unit; CNY stays in CNY instead of receiving an ephemeral exchange-rate conversion. The reference job uses 200K uncached input and 20K output. It makes rates legible and leaves tokenizer, reasoning tokens, retries, tool fees, and success rate to the owned replay.
| Model and freeze method | Published modality and window | Cache / input / output rate | Example bill | SOSEC role and missing gate |
|---|---|---|---|---|
DeepSeek V4 Flash 0731deepseek-v4-flash moves; retain returned version and fingerprint; pin a revision locally | Text-only; 1M input context; 384K maximum output; thinking by default; Chat, Responses, Anthropic compatibility, and function tools | $0.0028 / $0.14 / $0.28 | $0.0336 | First low-cost coding trial; public beta, moving alias, owned quality and stability still required |
GPT-5.6 Terragpt-5.6-terra | Text output with text and image input; 1.05M / 128K; Responses, functions, structured output, and several hosted tools | $0.20 / $2 / $12; requests above 272K input bill 2× input and 1.5× output across the full request | $0.64 | Mature tool and vision comparator; long-prompt tiers, tool fees, and real cache-write billing enter the replay |
Claude Sonnet 5claude-sonnet-5 is a pinned snapshot under Anthropic's version rules | Text and image input, text output; 1M / 128K; adaptive thinking and tools; eligible ZDR arrangements | Introductory $2 / $10 through August 31; standard $3 / $15; caching follows separate Anthropic rules | $0.60 intro; $0.90 standard | Fixed-identity and data-governance comparator; the new tokenizer changes old workload token counts |
Gemini 3.6 Flashgemini-3.6-flash stable | Text, image, video, audio, and PDF input; 1,048,576 / 65,536; Search, Maps, code execution, file search, and functions | $0.15 / $1.50 / $7.50; cache storage separate | $0.45 | First multimodal and hosted-tool trial; grounding, tool, and storage fees follow the job |
GLM 5.2glm-5.2 | Text-only; 1M / 128K; thinking, functions, structured output, caching, and MCP; MIT weights | $0.26 / $1.40 / $4.40 | $0.368 | Managed plus open-weight comparator; a complete local route begins at eight-card scale |
Kimi K3kimi-k3 | Native multimodal; 1M; always-thinking; OpenAI and Anthropic compatibility; open weights | $0.30 / $3 / $15 | $0.90 | Very large open multimodal agent comparator; local topology, license, and parser costs are substantial |
| Grok 4.5 freeze a dated ID; the bare name and -latest move | Vendor's current code and general recommendation; 500K; adjustable reasoning; tools per xAI documentation | Short context $0.30 / $2 / $6; long context $0.60 / $4 / $12 | $0.52–$1.04 | Pricing says long context starts at ≥200K; the model page says above 200K. This example is exactly 200K, so a live bill closes the boundary conflict |
MiniMax M3MiniMax-M3; the pricing page leads parts of the API guide | Native multimodal; 1M weight model; managed long-context tiers; open weights under a community license | Current ≤512K: $0.06 / $0.30 / $1.20; 512K–1M: $0.12 / $0.60 / $2.40 | $0.084 | Low-price multimodal candidate; close pricing/API-guide synchronization, long-context availability, and commercial authorization first |
Qwen3.7 Maxqwen3.7-max-2026-06-08 pinned snapshot | Text, image, and video input with text output; 1M; thinking/non-thinking, functions, built-in tools, and structured output | Cache not listed in the standard price table / CNY 12 / CNY 36; one tier for 0<input≤1M; regional prices vary | CNY 3.12 | Pinned Alibaba Cloud snapshot comparator; select region, currency, and exact snapshot before arithmetic |
DeepSeek's rate is the visual outlier and the rate most likely to make procurement skip a trial. One successful task can consume several model calls, tool returns, and repair loops. Failed tasks also incur charges. Modality changes input metering, and hosted search, file retrieval, code execution, and computer use may add per-call cost. The comparable unit is all-in cost for one accepted task, combining tokens, tools, retries, and human correction under one denominator.
Lifecycle controls whether a result can be reproduced. DeepSeek accepts a moving public API name and therefore needs response identity. Anthropic explicitly treats the dateless Sonnet 5 ID as a pinned snapshot. Google labels 3.6 Flash stable. xAI's bare and -latest names move. Qwen exposes both a moving alias and dated snapshots. A model router converts these differences into policy: fixed evaluations use a fixed ID, local weights use a revision, and moving endpoints retain returned identity and automatically rerun regressions on change.
4 DeepSeek's release results admit 0731 to the first trial; production ranking still needs common evidence
DeepSeek published nine agent results for 0731: 82.7 on Terminal Bench 2.1, 54.2 on NL2Repo, 76.7 on Cybergym, 54.4 on DeepSWE, 70.3 on Toolathlon-Verified, 25.2 on Agents’ Last Exam, 25.1 on AutomationBench Public, 68.7 on DSBench-FullStack, and 59.6 on DSBench-Hard. Every result rises materially from Preview in the same release table; DeepSWE moves from 7.3 to 54.4. Those numbers explain why 0731 receives the first coding trial.
The model card also states the evidence boundary. Public code-agent tasks used the minimal mode of DeepSeek Harness, marked “to be released,” at max effort with temperature=1.0 and top_p=0.95. DSBench-FullStack and DSBench-Hard are internal. The release table includes GLM and Opus results, while the current page does not supply a single container, tool definition, task snapshot, turn cap, rerun policy, and complete log set for every compared model.
A cross-vendor league table would compress those implementation differences into one decimal column. A coding agent's outcome comes from the model, harness, shell authority, tests, context construction, retries, and stop policy together. SOSEC preserves the publisher results for their supported decision: choosing which candidate to evaluate. Production order comes from the common owned replay. A future fixed Harness and log release can raise reproducibility while the original result retains its version and settings.
4.1 A small, hard task set sits closer to procurement than a thousand generic questions
Stratify the task set by business failure. Small repository repairs test localization and test discipline. Cross-file features test planning, context, and tool loops. Migrations and dependency upgrades test documentation, compatibility, and rollback. Log diagnosis tests evidence use. Security tasks test whether the agent crosses authority or avoids a failing control. Each class keeps a positive control known to pass, a negative control that must be rejected, and one resource-failure path.
The grader accepts runnable outcomes. A candidate edits an isolated worktree, and native project tests, static checks, and task-specific assertions determine success. Human review then checks authority, scope, maintainability, and hidden side effects. A completion claim, higher line count, or a permissive test authored by the model adds nothing to the score. On failure, record the first divergence: missing context, faulty plan, invalid tool arguments, test escape, resource stop, refusal, or repair loop.
Run each model at least twice per task and rotate order so caching, provider load, and warm repository state cannot become a fixed advantage. Retain complete use and price, then calculate accepted-task cost, human minutes, and p95. This answers the purchasing question: for this team and repository, does 0731 deliver mergeable work faster and more cheaply than the incumbent?
5 From per-million tokens to total cost: cheap API service creates a high utilization bar for local nodes
Begin with a bill a person can read. The reference job uses 200K uncached input and 20K output. DeepSeek 0731 costs 0.2×0.14 + 0.02×0.28 = $0.0336. At the same token counts, MiniMax M3's current ≤512K rate costs $0.084, GLM 5.2 $0.368, Gemini 3.6 Flash $0.45, introductory Claude Sonnet 5 $0.60, GPT-5.6 Terra $0.64, and both Kimi K3 and standard Sonnet 5 $0.90. Grok 4.5 lands in a $0.52–$1.04 range because two official pages disagree at the exact 200K boundary. Qwen's documented global tier produces CNY 3.12.
The Grok range comes from a locatable documentation conflict. Its model detail says long-context pricing applies above 200K; the pricing matrix says ≥200K and doubles input, cache, and output rates together. The reference input is exactly 200K, so prose cannot settle the boundary. The POC retains usage and the actual bill as the procurement record. The other eight candidates have determinate tiers at 200K input: Terra remains below its >272K trigger; Sonnet 5 and Gemini 3.6 Flash publish one rate across their windows; MiniMax stays in its ≤512K tier; Qwen's snapshot covers 0<input≤1M; and the current DeepSeek, GLM, and Kimi pages publish no separate 200K step.
The arithmetic stays deliberately narrow. Tokenizers differ, so 200K identical characters become different token counts. Reasoning tokens, media, search, and tools alter the bill. One coding task commonly makes several calls. Caching can rewrite the result: DeepSeek cache-hit input costs only $0.0028 per million, so a stable shared prefix lowers the API side further. The evaluator retains original bytes, provider token counts, cache hit/miss, and tool charges; the example rate never substitutes for measured cost.
The local formula is cost per accepted task = (GPU + CPU/memory + storage + network + power/facility + platform + operations + redundancy per hour) / accepted tasks per node-hour. Accelerator utilization shows that equipment is busy. Accepted jobs show that it is delivering. Queueing, long-context KV cache, failed retries, weight updates, node failures, and idle valleys raise the numerator or lower the denominator.
Runpod's public Pod page, updated July 27, 2026, lists H200 at $4.39, B200 at $5.89, and B300 at $7.39 per GPU-hour. Four cards over a 730-hour month produce raw GPU rent of $12,818.80, $17,198.80, and $21,578.80. These numbers exclude CPU, memory, storage, network, image transfer, on-call labor, and redundancy. A quoted Pod must also place all four cards in a node with suitable interconnect. A B300 rate cannot stand in for a GB300 quote.
Dividing by the $0.0336 DeepSeek reference bill, 4×H200 needs about 523 accepted reference jobs per node-hour merely to reach raw rent, 4×B200 about 701, and 4×B300 about 880. These are required capacity thresholds; throughput must be measured at the target context, effort, output, concurrency, and quality gate. A second node, staff, and idle capacity push the threshold higher. API cache hits move it farther away.
The monthly view is even clearer. One hundred reference jobs per day cost about $100.80 on the DeepSeek API. One thousand cost about $1,008. Ten thousand cost about $10,080. Raw four-H200 rent approaches the third case, while the owner still must show that the node can deliver 10,000 accepted jobs every day and absorb peaks, failure, and upgrades. Data residency, egress restrictions, custom serving, or predictable reserved capacity can create an independent reason for local deployment; document that non-price value in the purchase rationale and exit gate.
6 Local deployment is real: 0731 starts at four-card scale, while one card establishes the control baseline
With the 0731 weight release, local and API routes can finally target the same product version. The 155.43 GiB figure covers repository weight files. Serving also consumes runtime, communication buffers, DSpark, KV cache, batches, and context. Full 1M and 384K high/max output require separate memory and latency measurements. The official four-GB300 example is an executable starting point; a purchase still follows the target concurrency and sequence-length calculation.
| Local candidate | Frozen weight and license | Published artifact scale | Official or framework starting topology | Practical role and public cost boundary |
|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | 7872f01b…; MIT; official 0731 plus DSpark | 155.43 GiB from SOSEC's manifest sum; about 167 GB in the HF UI | Model card: 4×GB300; the current vLLM V4 Flash recipe defaults to 0731 and gives four-card H200/B200/B300 replicas plus 8×H200 4+4 PD; high/max at least 384K | First exact-model local trial; 4×H200 raw rent $17.56/hour, with topology and full TCO requiring quote and measurement |
| Qwen3.6-35B-A3B FP8 | Official FP8; Apache-2.0; 35B total/3B active | About 34.89 GiB in the captured manifest; 262K native context with YaRN extension | vLLM provides a single H100/H200 or MI300X-class route; NVFP4 reaches RTX Pro 6000 and DGX Spark routes | Single-card serving, tool, and observability control for a small team; it carries no 0731 quality claim; 1×H200 raw rent $3,204.70/month |
| GLM 5.2 | Official weights; MIT | About 1,403.19 GiB in the captured BF16 manifest | Practical FP8 default on 8×H200/H20; full 1M reference path on 8×B200 with FP8 KV cache | Heavier open coding/agent comparator; 8×H200 raw rent $25,637.60/month, with quantized quality and throughput still measured |
| MiniMax M3 | MiniMax Community License; legal review for commercial attribution and the stated revenue authorization boundary | About 854 GB BF16 in the captured official repository manifest; the card confirms 427/428B total and 23B active | Native multimodal; validate hardware and long-context configuration against current serving documentation and supplier quote | Procurement evaluation when local multimodal input matters; close license, documentation sync, and node cost first |
| Kimi K3 | Official weights; Kimi K3 License; 2.8T total/104B active | About 1,453.74 GiB in the captured manifest | vLLM recipe requires at least 8×GB300; real production adds multi-node networking and tool-parser governance | Very large native multimodal local comparator; GB300 and multi-node cost requires a formal quote |
6.1 Beyond four accelerators, the owner must build an internal model service
Download first solves provenance, resumption, verification, and mirroring for a roughly 167 GB object. Weights and code enter a read-only artifact repository, and production nodes pull from an internal mirror. Each release binds the Hugging Face revision, per-file digests, encoding directory, container, and driver. The DeepSeek card ships dedicated encoding code instead of a Jinja chat template. vLLM's DeepSeek V4 tokenizer, reasoning parser, and tool parser must be accepted with their fixed versions. Parser drift can change tool behavior while weights stay constant.
The service layer defines tenant isolation, queueing, fairness, rate limits, timeouts, cancellation, interrupted streams, and the resource-exhaustion contract. Long contexts consume KV cache, so one huge request can delay small requests. Partition scheduling by context and effort, cap in-flight tokens per tenant, and expose queue, prefill, decode, KV eviction, and failure metrics. Enable the complete one-million window only after real tasks show gain; size the default window from observed useful input.
Tool execution remains an application control. The model service emits parsed intent. A policy layer checks principal, tool, arguments, repository, path, and side-effect budget. An executor runs in a separate sandbox and returns a size- and sensitivity-bounded result. Multitenant use also validates prompt and cache isolation, log redaction, artifact access, and management-interface authentication. Local deployment brings the data path under organizational control and assigns these duties to that organization.
High availability needs a second capacity source or a managed fallback. Maintenance, driver failure, weight reload, and GPU faults can remove the whole single-node route. Two fully provisioned nodes roughly double most fixed costs. An owner can use one primary plus a smaller degradation pool, two availability-zone nodes, or local-first routing with DeepSeek API fallback. Each choice states whether data may leave the boundary, which jobs survive degradation, whether caches remain compatible, and how weight and log continuity is verified before returning.
6.2 Buying hardware, renting nodes, and staying on API make different commitments
API service fits variable demand, modest task volume, and rapidly changing models. The provider operates accelerators, kernels, and availability; the caller still owns evaluation, data governance, tool authority, and rollback. DeepSeek's rate stretches this route across a large volume range. The moving alias and public-beta state require stricter response identity and regression handling.
Hourly four-card rental fits data-location experiments, temporary capacity, and serving customization. The buyer can fix weights and container and release the node later. Persistent downloads, idle time, and cross-zone storage add charges. Before contracting, confirm same-host interconnect, shared storage, image pull time, quota, and replacement behavior. A public per-GPU-hour rate creates a usable node only when the topology qualifies.
Owned hardware fits stable high utilization and organizations with an existing facility and GPU operations. Total ownership includes depreciation, capital, warranty, racks, power, cooling, switching, spares, staff, and utilization opportunity cost. A three-year amortization cannot erase the possibility that the preferred model changes in six months. The exit path lets the node carry other inference work and preserves the managed-API compatibility layer. Accepted-job throughput, policy value, and reusable infrastructure together decide the purchase.
7 Close the decision: treat 0731 as a reversible product change
For most teams, the next action is straightforward: keep the current production model and run the same real coding tasks through the managed DeepSeek API. Its $0.0336 reference bill makes quality, tool, latency, and failure evidence cheap to collect before reserving a four-card node. Add Gemini 3.6 Flash when the task is multimodal and Sonnet 5 when fixed identity or ZDR eligibility matters. Other models enter when the current workload or procurement boundary calls for them.
The 0731 weight release changes the data boundary and customization space. A public revision, MIT license, and serving documentation create an executable local POC. Four-card hardware, 155.43 GiB of weights, KV cache, DSpark, parsers, and model-service operations form the actual solution. Hundreds or thousands of long jobs per day commonly retain lower direct cost through the API. Sustained high throughput, data residency, or an existing GPU platform can move local service to the first trial.
The production answer belongs in a versioned decision record: task-set revision, candidate identity, success, human repair, p50/p95, all-in cost, tool side effects, data path, capacity, recovery, and rollback result. Canary 0731 when it passes and withdraw it when a critical gate fails. Repeat the same contract for the next release. This versioned contract becomes the durable system for choosing models; a static leaderboard expires with the next alias or backend update.
Research record
8Evidence, objects, and sources
The material below preserves the identifiers and references used in this report.
8.1Research objects
Products, actors, techniques, affected objects, and control points discussed in the report.
Moving request ID mapped to DeepSeek-V4-Flash-0731 on the 2026-08-01 pricing page
Official Hugging Face revision pinned for SOSEC's blobs-manifest read
Manifest sum of 48 safetensors, equal to 155.43 GiB
Billing arithmetic that carries no cross-tokenizer quality or throughput claim
8.2Event chronology
- DeepSeek V4 Preview weights released
V4 Flash Preview established the 284B/13B, one-million-context, MIT-weight route; 0731 later re-post-trained the same architecture.
- V4 Flash 0731 API enters public beta
DeepSeek updated the moving API alias, agent results, Responses API, and rate, with explicit Harness and internal-set boundaries.
- Official 0731 weights enter Hugging Face
The official repository and V4 collection added 0731; SOSEC subsequently pinned the revision and calculated weight bytes.
- SOSEC closes the public-information and cost snapshot
API rates, model identity, local topologies, public GPU rent, and comparator entry points were reviewed in one snapshot.
8.3Sources and material
- DeepSeek API update noticehttps://api-docs.deepseek.com/updates/
- DeepSeek models, features, pricing, and concurrencyhttps://api-docs.deepseek.com/quick_start/pricing/
- DeepSeek Chat Completions parameters, finish reasons, usage, and fingerprinthttps://api-docs.deepseek.com/api/create-chat-completion
- DeepSeek rate limits and user isolationhttps://api-docs.deepseek.com/quick_start/rate_limit
- Official DeepSeek V4 Flash 0731 card, evaluation, deployment, and MIT licensehttps://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
- Pinned official 0731 weight revisionhttps://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/tree/7872f01b1d1fe23eabc4c98b48bffcef5a386062
- Official DeepSeek V4 model collectionhttps://huggingface.co/collections/deepseek-ai/deepseek-v4
- General vLLM V4 Flash deployment recipe, currently defaulting to 0731https://recipes.vllm.ai/deepseek-ai/DeepSeek-V4-Flash
- SGLang DeepSeek V4 deployment and tuning cookbookhttps://docs.sglang.io/cookbook/autoregressive/DeepSeek/DeepSeek-V4
- Official OpenAI GPT-5.6 Terra model pagehttps://developers.openai.com/api/docs/models/gpt-5.6-terra
- Current Anthropic Claude models, pricing, and contexthttps://platform.claude.com/docs/en/about-claude/models/overview
- Anthropic model ID and version ruleshttps://platform.claude.com/docs/en/about-claude/models/model-ids-and-versions
- Anthropic full pricing and Sonnet 5 full-window standard rateshttps://platform.claude.com/docs/en/about-claude/pricing
- Anthropic Sonnet 5 tokenizer, context, and ZDR noteshttps://platform.claude.com/docs/en/about-claude/models/whats-new-sonnet-5
- Anthropic API data-retention and ZDR eligibility boundarieshttps://platform.claude.com/docs/en/manage-claude/api-and-data-retention
- Official Google Gemini 3.6 Flash model pagehttps://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash
- Official Google Gemini API pricinghttps://ai.google.dev/gemini-api/docs/pricing
- Z.ai GLM 5.2 model and interfacehttps://docs.z.ai/guides/llm/glm-5.2
- Official Z.ai model pricinghttps://docs.z.ai/guides/overview/pricing
- Official GLM 5.2 open weightshttps://huggingface.co/zai-org/GLM-5.2
- vLLM GLM 5.2 deployment recipehttps://recipes.vllm.ai/zai-org/GLM-5.2
- Kimi API models and pricinghttps://platform.kimi.ai/
- Official Kimi K3 card and licensehttps://huggingface.co/moonshotai/Kimi-K3
- vLLM Kimi K3 deployment recipehttps://recipes.vllm.ai/moonshotai/Kimi-K3
- Alibaba Cloud Model Studio text models and snapshotshttps://help.aliyun.com/en/model-studio/text-generation-model
- Alibaba Cloud Qwen3.7 Max snapshot and visual-capability release recordhttps://help.aliyun.com/en/model-studio/newly-released-models
- Official Alibaba Cloud Model Studio pricinghttps://help.aliyun.com/en/model-studio/model-pricing
- Current xAI models, identities, and priceshttps://docs.x.ai/developers/models
- xAI Grok 4.5 model detail and 200K boundary wordinghttps://docs.x.ai/developers/models/grok-4.5
- xAI model pricing and ≥200K long-context matrixhttps://docs.x.ai/developers/pricing
- Current MiniMax API pricinghttps://platform.minimax.io/subscribe/token-plan?tab=api-enterprise
- Official MiniMax M3 weights and licensehttps://huggingface.co/MiniMaxAI/MiniMax-M3
- Full MiniMax M3 Community Licensehttps://huggingface.co/MiniMaxAI/MiniMax-M3/blob/main/LICENSE
- Official Qwen3.6-35B-A3B model cardhttps://huggingface.co/Qwen/Qwen3.6-35B-A3B
- Official Qwen3.6-35B-A3B FP8 weightshttps://huggingface.co/Qwen/Qwen3.6-35B-A3B-FP8
- vLLM single-card Qwen3.6-35B-A3B recipehttps://recipes.vllm.ai/Qwen/Qwen3.6-35B-A3B
- Runpod public GPU Pod pricinghttps://www.runpod.io/pricing