Research
Should You Switch to DeepSeek V4 Flash 0731? Model, Cost, and Self-Hosting Comparison — August 2026
As of August 1, 2026, public evidence yields no production winner: test deepseek-v4-flash first for low-cost text coding and record its 0731 fingerprint, Gemini 3.6 Flash first for multimodal work, and Claude Sonnet 5 for fixed-snapshot or ZDR needs; the MIT 0731 weights are public, while local adoption starts after a four-accelerator TCO trial and owned tasks pass with the incumbent ready for rollback.

In this article
1 The answer first: trial 0731 for low-cost coding and keep production on the incumbent
Verdict: DeepSeek V4 Flash 0731 is the strongest first candidate to trial in August 2026 for a low-cost, text-only coding agent. Request deepseek-v4-flash, leave thinking enabled, and begin at high effort. Keep production on the incumbent until a set of real repository tasks returns quality, latency, tool reliability, and successful-task cost together. Public material determines trial order; it does not yet identify a cross-vendor production winner.
Start jobs containing images, audio, video, or PDF with Gemini 3.6 Flash, which exposes those inputs alongside managed search, code execution, file search, and functions. Use Claude Sonnet 5 as the comparator when a fixed model identity or an eligible Zero Data Retention arrangement is mandatory. The MIT-licensed 0731 weights provide an executable route when data must remain inside an owned boundary; budget that route from a four-accelerator node, while a smaller one-card model validates only the serving, tool, and observability contract.
GPT, Grok, GLM, Kimi, MiniMax, and Qwen still occupy useful capability, price, and deployment boundaries. Put the relevant candidates through the same tasks and let the observed work decide. The practical choice is the route that repeatedly completes this team's jobs and can return immediately to a known-good backend when a gate fails.
2 Turn the low price into a reversible production decision with thirty real tasks
The first round is a comparison, not a production migration. Sample roughly thirty real tasks across defect repair, cross-file features, test completion, migrations, dependency updates, and log diagnosis. Freeze each starting commit, system prompt, tool schema, permission set, maximum turns, timeout, expected tests, and prohibited side effects. Run the incumbent and 0731 at least twice with alternating order. Add Gemini 3.6 Flash or Claude Sonnet 5 under the same task contract when multimodal input or data governance is material.
Retain one receipt for the request and response. Use Chat Completions or Responses at https://api.deepseek.com and request deepseek-v4-flash. At the cutoff, cache-hit input, uncached input, and output cost $0.0028, $0.14, and $0.28 per million tokens. The pricing page also announces a future peak period at twice the current rates without an effective date, so refresh the rate on the canary-admission day. Store the returned model, system_fingerprint, finish reason, reasoning/output tokens, cache hit/miss, first-token and total latency, tool-argument validation, retries, and final tests. A length finish enters context or output-budget repair; content_filter enters human review or an approved alternate route; insufficient_system_resource receives bounded backoff and then the incumbent. Local schema and authorization checks precede every side effect.
Admission uses executable outcomes. The primary unit is a task that passes without a human rewrite. Also count human repair lines, extra turns, invalid tool calls, incomplete work, refusals, test escapes, and rollback. For critical work, require the candidate's lower confidence bound on success to meet the incumbent, keep unauthorized tool calls and critical test escapes at zero, and hold p95 inside the current service objective. All-in savings per accepted task must also cover migration, observability, and incremental on-call work. Record the baseline, sample size, interval, and business zero-tolerance conditions before the replay, then let the evidence determine canary admission.
A local trial is pinned to bytes and topology. Use official revision 7872f01b…; its 48 safetensors total 166,886,535,336 bytes, or 155.43 GiB. The official model card starts at single-node 4×GB300 vLLM, while its linked recipe also covers four-card H200, B200, and B300 replicas plus an eight-H200 4+4 prefill/decode layout. Pin weights, encoding and remote code, container, vLLM or SGLang, driver, and parsers. Measure interconnect, memory, kernel, KV cache, concurrency, and first-token latency on the exact provider SKU.
Canary with side-effect-free work first. Mirror jobs before admitting a small reversible share. A changed fingerprint or local artifact digest, repeated failures, tool-schema errors, a critical test escape, or a cost/p95 breach closes the candidate route and sends new work to the incumbent. After rollback, run a fixed positive set and a denied or unauthorized-tool negative set to prove routing, logs, billing, and permissions have returned. End the trial when two corrective cycles still miss the critical success gate or the local route's multitenancy, operational, or license obligations exceed the team's capacity.
3 Three identities now line up: API alias, 0731 backend, and open weights
DeepSeek documents deepseek-v4-flash as the request ID. On August 1, 2026, the pricing page mapped it to DeepSeek-V4-Flash-0731, while the API response exposed the actual model and a system_fingerprint representing backend configuration. Procurement, evaluation, and incident records need all three. The request alias identifies the entry point, the returned model identifies the served version, and the fingerprint catches a backend change under the same public name.
The 0731 open weights entered DeepSeek's official Hugging Face organization and V4 collection later the same release day. The model card calls this the official Flash release that supersedes Preview, retains the V4 Flash structure, and includes a DSpark speculative-decoding module. The repository and weights carry the MIT license. This report pins revision 7872f01b1d1fe23eabc4c98b48bffcef5a386062 after the repository's 2026-08-01 03:07:41 UTC update. API and local trials can therefore target the same named 0731 product version, while sampling, encoding, kernels, quantization, scheduling, and tool wrappers remain measurable implementation differences.
The older deepseek-ai/DeepSeek-V4-Flash repository continues to represent the Preview route. DeepSeek-V4-Flash-DSpark combines Preview weights with the speculative module. The new 0731 repository contains the re-post-trained official weights and DSpark. Reproduction records must name a complete repository and revision; the short label “V4 Flash” spans three distinct artifacts.
4 Public prices differ sharply across nine candidates, while price still answers only the bill
The comparison fixes 200K uncached input and 20K output and applies each vendor's public rates on August 1, 2026. It makes the order of magnitude visible while preserving three boundaries: CNY is not converted through a transient exchange rate, Grok has conflicting official tier language at exactly 200K input, and no amount includes tokenizer differences, reasoning tokens, tool calls, retries, or failed jobs.
DeepSeek V4 Flash 0731 costs $0.0336 in the example. It is text-only, publishes a 1M input context and 384K maximum output, and exposes both a moving API alias and MIT weights. MiniMax M3 costs $0.084 and publishes native multimodality, a 1M context, and open weights under a community license. GLM 5.2 costs $0.368 and publishes text-only 1M/128K operation, functions and MCP, plus MIT weights. Price admits them to different-budget trials; license, API synchronization, and local topology remain separate gates.
Gemini 3.6 Flash costs $0.45 and publishes text, image, video, audio, and PDF input with a 1,048,576/65,536 window plus search, code execution, file search, and functions. Grok 4.5's 500K route lands at $0.52–$1.04 because two official pages disagree on the exact 200K tier boundary. Claude Sonnet 5 accepts text and images at 1M/128K and costs $0.60 at the introductory rate through August 31 or $0.90 at standard rates; it also supplies fixed-snapshot rules and eligible ZDR arrangements. GPT-5.6 Terra costs $0.64 and publishes a 1.05M/128K window, image input, and a broad hosted-tool surface.
Kimi K3's 1M native-multimodal open-weight route costs $0.90 in the example, while its local artifact and minimum topology occupy the very-large-model class. The dated Qwen3.7 Max snapshot accepts text, image, and video and costs CNY 3.12 in the global CNY tier; procurement records must fix region, currency, and exact snapshot. One successful coding job commonly consumes several model calls, tool returns, and repair loops, while failed calls are billable. Total cost for one accepted job is the comparison unit that survives those differences.
Lifecycle determines reproducibility. DeepSeek uses a moving public API alias, so the caller retains response identity. Anthropic defines the undated Sonnet 5 ID as a fixed snapshot. Google marks 3.6 Flash stable. xAI's bare name and -latest move. Qwen provides both moving aliases and dated snapshots. Fixed evaluations use a pinnable ID, local weights use a revision, and moving entry points automatically rerun regression when returned identity changes.
5 DeepSeek's nine release scores buy a trial slot; real tasks decide elimination
DeepSeek published nine agent results for 0731: 82.7 on Terminal Bench 2.1, 54.2 on NL2Repo, 76.7 on Cybergym, 54.4 on DeepSWE, 70.3 on Toolathlon-Verified, 25.2 on Agents' Last Exam, 25.1 on AutomationBench Public, 68.7 on DSBench-FullStack, and 59.6 on DSBench-Hard. Every result rises materially from Preview in the same release table; DeepSWE moves from 7.3 to 54.4. Those numbers explain why 0731 receives the first coding trial.
The model card also states the evidence boundary. Public code-agent tasks used the minimal mode of DeepSeek Harness, marked “to be released,” at max effort with temperature=1.0 and top_p=0.95. DSBench-FullStack and DSBench-Hard are internal. The release table includes GLM and Opus results, while the current page does not supply a single container, tool definition, task snapshot, turn cap, rerun policy, and complete log set for every compared model.
A cross-vendor league table would compress those implementation differences into one decimal column. A coding agent's outcome comes from the model, harness, shell authority, tests, context construction, retries, and stop policy together. Publisher results choose which candidate enters evaluation; production order comes from the common owned replay. A future fixed Harness and log release can raise reproducibility while the original result retains its version and settings.
Procurement turns on a small, hard set of real tasks. Stratify the task set by business failure. Small repository repairs test localization and test discipline. Cross-file features test planning, context, and tool loops. Migrations and dependency upgrades test documentation, compatibility, and rollback. Log diagnosis tests evidence use. Security tasks test whether the agent crosses authority or avoids a failing control. Each class keeps a positive control known to pass, a negative control that must be rejected, and one resource-failure path.
The grader accepts runnable outcomes. A candidate edits an isolated worktree, and native project tests, static checks, and task-specific assertions determine success. Human review then checks authority, scope, maintainability, and hidden side effects. A completion claim, higher line count, or a permissive test authored by the model adds nothing to the score. On failure, record the first divergence: missing context, faulty plan, invalid tool arguments, test escape, resource stop, refusal, or repair loop.
Run each model at least twice per task and rotate order so caching, provider load, and warm repository state cannot become a fixed advantage. Retain complete use and price, then calculate accepted-task cost, human minutes, and p95. This answers the purchasing question: for this team and repository, does 0731 deliver mergeable work faster and more cheaply than the incumbent?
6 From per-million tokens to total cost: cheap API service creates a high utilization bar for local nodes
Before calculating local break-even, turn only the DeepSeek rate into a bill. The reference job uses 200K uncached input and 20K output, yielding 0.2×0.14 + 0.02×0.28 = $0.0336. Figure 2 retains the other candidates' amounts on the same workload. The hardware calculation below uses this DeepSeek bill as its denominator without reciting the same price row again.
The Grok range comes from a locatable documentation conflict. Its model detail says long-context pricing applies above 200K; the pricing matrix says ≥200K and doubles input, cache, and output rates together. The reference input is exactly 200K, so prose cannot settle the boundary. The POC retains usage and the actual bill as the procurement record. The other eight candidates have determinate tiers at 200K input: Terra remains below its >272K trigger; Sonnet 5 and Gemini 3.6 Flash publish one rate across their windows; MiniMax stays in its ≤512K tier; Qwen's snapshot covers 0<input≤1M; and the current DeepSeek, GLM, and Kimi pages publish no separate 200K step.
The arithmetic stays deliberately narrow. Tokenizers differ, so 200K identical characters become different token counts. Reasoning tokens, media, search, and tools alter the bill. One coding task commonly makes several calls. Caching can rewrite the result: DeepSeek cache-hit input costs only $0.0028 per million, so a stable shared prefix lowers the API side further. The evaluator retains original bytes, provider token counts, cache hit/miss, and tool charges; the example rate never substitutes for measured cost.
The local formula is cost per accepted task = (GPU + CPU/memory + storage + network + power/facility + platform + operations + redundancy per hour) / accepted tasks per node-hour. Accelerator utilization shows that equipment is busy. Accepted jobs show that it is delivering. Queueing, long-context KV cache, failed retries, weight updates, node failures, and idle valleys raise the numerator or lower the denominator.
Runpod's public Pod page, updated July 27, 2026, lists H200 at $4.39, B200 at $5.89, and B300 at $7.39 per GPU-hour. Four cards over a 730-hour month produce raw GPU rent of $12,818.80, $17,198.80, and $21,578.80. These numbers exclude CPU, memory, storage, network, image transfer, on-call labor, and redundancy. A quoted Pod must also place all four cards in a node with suitable interconnect. A B300 rate cannot stand in for a GB300 quote.
Dividing by the $0.0336 DeepSeek reference bill, 4×H200 needs about 523 accepted reference jobs per node-hour merely to reach raw rent, 4×B200 about 701, and 4×B300 about 880. These are required capacity thresholds; throughput must be measured at the target context, effort, output, concurrency, and quality gate. A second node, staff, and idle capacity push the threshold higher. API cache hits move it farther away.
The monthly view is even clearer. One hundred reference jobs per day cost about $100.80 on the DeepSeek API. One thousand cost about $1,008. Ten thousand cost about $10,080. Raw four-H200 rent approaches the third case, while the owner still must show that the node can deliver 10,000 accepted jobs every day and absorb peaks, failure, and upgrades. Data residency, egress restrictions, custom serving, or predictable reserved capacity can create an independent reason for local deployment; document that non-price value in the purchase rationale and exit gate.
7 Beyond four accelerators, the team is buying a model service
With the 0731 weight release, local and API routes can finally target the same product version. The 155.43 GiB figure covers repository weight files. Serving also consumes runtime, communication buffers, DSpark, KV cache, batches, and context. Full 1M and 384K high/max output require separate memory and latency measurements. The official four-GB300 example is an executable starting point; a purchase still follows the target concurrency and sequence-length calculation.
0731 is the exact-model local path. Official revision 7872f01b… is MIT-licensed and combines the official 0731 weights with DSpark. The manifest holds 155.43 GiB of weights, while the Hugging Face UI reports roughly 167 decimal GB. The model card gives 4×GB300; the current vLLM recipe that defaults to 0731 gives four-card H200, B200, and B300 replicas plus an 8×H200 4+4 prefill/decode layout. At the public H200 rate, four cards cost $17.56 per raw GPU-hour; interconnect, CPU, memory, storage, network, redundancy, and staff still require quotes and measurement.
Qwen3.6-35B-A3B FP8 is a one-card control. This Apache-2.0 model has 35B total and 3B active parameters, a captured 34.89 GiB manifest, and 262K native context. vLLM provides a single H100, H200, or MI300X-class route. It exposes container, parser, tool, queue, and observability defects early and establishes the roughly $3,204.70 per month 1×H200 operating baseline. It carries no 0731 quality claim.
The other three open routes remain outside the first purchase. GLM 5.2 has a captured 1,403.19 GiB BF16 manifest; practical FP8 starts at 8×H200/H20, while its full-1M reference uses 8×B200 and FP8 KV cache. MiniMax M3 has a captured roughly 854 GB BF16 manifest; its community license's commercial attribution and revenue boundary, serving-document synchronization, and node quote remain open. Kimi K3 has a captured 1,453.74 GiB manifest, 2.8T total and 104B active parameters, and a vLLM recipe starting at 8×GB300 before multi-node networking and tool-parser governance. They enter a later round only when owned tasks require their corresponding capabilities.
Deployment identity extends from weights into parsers. Download first solves provenance, resumption, verification, and mirroring for a roughly 167 GB object. Weights and code enter a read-only artifact repository, and production nodes pull from an internal mirror. Each release binds the Hugging Face revision, per-file digests, encoding directory, container, and driver. The DeepSeek card ships dedicated encoding code instead of a Jinja chat template. vLLM's DeepSeek V4 tokenizer, reasoning parser, and tool parser must be accepted with their fixed versions. Parser drift can change tool behavior while weights stay constant.
The service layer defines tenant isolation, queueing, fairness, rate limits, timeouts, cancellation, interrupted streams, and the resource-exhaustion contract. Long contexts consume KV cache, so one huge request can delay small requests. Partition scheduling by context and effort, cap in-flight tokens per tenant, and expose queue, prefill, decode, KV eviction, and failure metrics. Enable the complete one-million window only after real tasks show gain; size the default window from observed useful input.
Tool execution remains an application control. The model service emits parsed intent. A policy layer checks principal, tool, arguments, repository, path, and side-effect budget. An executor runs in a separate sandbox and returns a size- and sensitivity-bounded result. Multitenant use also validates prompt and cache isolation, log redaction, artifact access, and management-interface authentication. Local deployment brings the data path under organizational control and assigns these duties to that organization.
High availability needs a second capacity source or a managed fallback. Maintenance, driver failure, weight reload, and GPU faults can remove the whole single-node route. Two fully provisioned nodes roughly double most fixed costs. An owner can use one primary plus a smaller degradation pool, two availability-zone nodes, or local-first routing with DeepSeek API fallback. Each choice states whether data may leave the boundary, which jobs survive degradation, whether caches remain compatible, and how weight and log continuity is verified before returning.
API service, hourly rental, and owned hardware make three different commitments. API service fits variable demand, modest task volume, and rapidly changing models. The provider operates accelerators, kernels, and availability; the caller still owns evaluation, data governance, tool authority, and rollback. DeepSeek's rate stretches this route across a large volume range. The moving alias and public-beta state require stricter response identity and regression handling.
Hourly four-card rental fits data-location experiments, temporary capacity, and serving customization. The buyer can fix weights and container and release the node later. Persistent downloads, idle time, and cross-zone storage add charges. Before contracting, confirm same-host interconnect, shared storage, image pull time, quota, and replacement behavior. A public per-GPU-hour rate creates a usable node only when the topology qualifies.
Owned hardware fits stable high utilization and organizations with an existing facility and GPU operations. Total ownership includes depreciation, capital, warranty, racks, power, cooling, switching, spares, staff, and utilization opportunity cost. A three-year amortization cannot erase the possibility that the preferred model changes in six months. The exit path lets the node carry other inference work and preserves the managed-API compatibility layer. Accepted-job throughput, policy value, and reusable infrastructure together decide the purchase.
8 Close the decision: treat 0731 as a reversible product change
For most teams, the next action is straightforward: keep the current production model and run the same real coding tasks through the managed DeepSeek API. Its $0.0336 reference bill makes quality, tool, latency, and failure evidence cheap to collect before reserving a four-card node. Add Gemini 3.6 Flash when the task is multimodal and Sonnet 5 when fixed identity or ZDR eligibility matters. Other models enter when the current workload or procurement boundary calls for them.
The 0731 weight release changes the data boundary and customization space. A public revision, MIT license, and serving documentation create an executable local POC. Four-card hardware, 155.43 GiB of weights, KV cache, DSpark, parsers, and model-service operations form the actual solution. Hundreds or thousands of long jobs per day commonly retain lower direct cost through the API. Sustained high throughput, data residency, or an existing GPU platform can move local service to the first trial.
The production answer belongs in a versioned decision record: task-set revision, candidate identity, success, human repair, p50/p95, all-in cost, tool side effects, data path, capacity, recovery, and rollback result. Canary 0731 when it passes and withdraw it when a critical gate fails. Repeat the same contract for the next release. This versioned contract becomes the durable system for choosing models; a static leaderboard expires with the next alias or backend update.
Research basis
Research basisAs of 2026-08-01 14:00 CST, this report reviewed the public evaluations, API identity, pricing, context, tool interfaces, and same-day MIT weight release for DeepSeek V4 Flash 0731; pinned Hugging Face revision 7872f01b1d1fe23eabc4c98b48bffcef5a386062 and summed the exact bytes of 48 safetensors from the blobs manifest; calculated a billing example with 200K uncached input and 20K output for DeepSeek plus eight current managed comparators; and recorded official or framework-validated hardware topologies, artifact sizes, licenses, and public GPU rent for five local routes. No common-task cross-model quality, throughput, latency, energy, or security test was run for this report, so public evidence determines POC order only; each production decision must come from the reader's own repositories, tools, and acceptance criteria.
SourceDeepSeek API documentation, update notice, official Hugging Face weights, and fixed repository manifest / official model and pricing material from OpenAI, Anthropic, Google, Z.ai, Moonshot AI, Alibaba Cloud, xAI, MiniMax, and Qwen / official vLLM, SGLang, and Runpod deployment and pricing material / SOSEC identity, price, weight-byte, and break-even review on 2026-08-01
Evidence confidence Medium-high
9Evidence and sources
9.1What we examined
Moving request ID mapped to DeepSeek-V4-Flash-0731 on the 2026-08-01 pricing page
Official Hugging Face revision pinned for SOSEC's blobs-manifest read
Manifest sum of 48 safetensors, equal to 155.43 GiB
Billing arithmetic that carries no cross-tokenizer quality or throughput claim
9.2Timeline
- DeepSeek V4 Preview weights released
V4 Flash Preview established the 284B/13B, one-million-context, MIT-weight route; 0731 later re-post-trained the same architecture.
- V4 Flash 0731 API enters public beta
DeepSeek updated the moving API alias, agent results, Responses API, and rate, with explicit Harness and internal-set boundaries.
- Official 0731 weights enter Hugging Face
The official repository and V4 collection added 0731; SOSEC subsequently pinned the revision and calculated weight bytes.
- SOSEC closes the public-information and cost snapshot
API rates, model identity, local topologies, public GPU rent, and comparator entry points were reviewed in one snapshot.
9.3Sources and material
- DeepSeek API update noticehttps://api-docs.deepseek.com/updates/
- DeepSeek models, features, pricing, and concurrencyhttps://api-docs.deepseek.com/quick_start/pricing/
- DeepSeek Chat Completions parameters, finish reasons, usage, and fingerprinthttps://api-docs.deepseek.com/api/create-chat-completion
- DeepSeek rate limits and user isolationhttps://api-docs.deepseek.com/quick_start/rate_limit
- Official DeepSeek V4 Flash 0731 card, evaluation, deployment, and MIT licensehttps://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
- Pinned official 0731 weight revisionhttps://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/tree/7872f01b1d1fe23eabc4c98b48bffcef5a386062
- Official DeepSeek V4 model collectionhttps://huggingface.co/collections/deepseek-ai/deepseek-v4
- General vLLM V4 Flash deployment recipe, currently defaulting to 0731https://recipes.vllm.ai/deepseek-ai/DeepSeek-V4-Flash
- SGLang DeepSeek V4 deployment and tuning cookbookhttps://docs.sglang.io/cookbook/autoregressive/DeepSeek/DeepSeek-V4
- Official OpenAI GPT-5.6 Terra model pagehttps://developers.openai.com/api/docs/models/gpt-5.6-terra
- Current Anthropic Claude models, pricing, and contexthttps://platform.claude.com/docs/en/about-claude/models/overview
- Anthropic model ID and version ruleshttps://platform.claude.com/docs/en/about-claude/models/model-ids-and-versions
- Anthropic full pricing and Sonnet 5 full-window standard rateshttps://platform.claude.com/docs/en/about-claude/pricing
- Anthropic Sonnet 5 tokenizer, context, and ZDR noteshttps://platform.claude.com/docs/en/about-claude/models/whats-new-sonnet-5
- Anthropic API data-retention and ZDR eligibility boundarieshttps://platform.claude.com/docs/en/manage-claude/api-and-data-retention
- Official Google Gemini 3.6 Flash model pagehttps://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash
- Official Google Gemini API pricinghttps://ai.google.dev/gemini-api/docs/pricing
- Z.ai GLM 5.2 model and interfacehttps://docs.z.ai/guides/llm/glm-5.2
- Official Z.ai model pricinghttps://docs.z.ai/guides/overview/pricing
- Official GLM 5.2 open weightshttps://huggingface.co/zai-org/GLM-5.2
- vLLM GLM 5.2 deployment recipehttps://recipes.vllm.ai/zai-org/GLM-5.2
- Kimi API models and pricinghttps://platform.kimi.ai/
- Official Kimi K3 card and licensehttps://huggingface.co/moonshotai/Kimi-K3
- vLLM Kimi K3 deployment recipehttps://recipes.vllm.ai/moonshotai/Kimi-K3
- Alibaba Cloud Model Studio text models and snapshotshttps://help.aliyun.com/en/model-studio/text-generation-model
- Alibaba Cloud Qwen3.7 Max snapshot and visual-capability release recordhttps://help.aliyun.com/en/model-studio/newly-released-models
- Official Alibaba Cloud Model Studio pricinghttps://help.aliyun.com/en/model-studio/model-pricing
- Current xAI models, identities, and priceshttps://docs.x.ai/developers/models
- xAI Grok 4.5 model detail and 200K boundary wordinghttps://docs.x.ai/developers/models/grok-4.5
- xAI model pricing and ≥200K long-context matrixhttps://docs.x.ai/developers/pricing
- Current MiniMax API pricinghttps://platform.minimax.io/subscribe/token-plan?tab=api-enterprise
- Official MiniMax M3 weights and licensehttps://huggingface.co/MiniMaxAI/MiniMax-M3
- Full MiniMax M3 Community Licensehttps://huggingface.co/MiniMaxAI/MiniMax-M3/blob/main/LICENSE
- Official Qwen3.6-35B-A3B model cardhttps://huggingface.co/Qwen/Qwen3.6-35B-A3B
- Official Qwen3.6-35B-A3B FP8 weightshttps://huggingface.co/Qwen/Qwen3.6-35B-A3B-FP8
- vLLM single-card Qwen3.6-35B-A3B recipehttps://recipes.vllm.ai/Qwen/Qwen3.6-35B-A3B
- Runpod public GPU Pod pricinghttps://www.runpod.io/pricing