GPT-6.1 Sol, Gemini 4 Argon and DeepSeek V4.1 Flash

September 2026’s three big model releases, what is actually new in each, and the ideas behind them: sparse experts, KV-cache compression and million-token outputs.

Intermediate lesson, about 25 minutes, with interactive demos and a quiz.

What you will learn

The news, and what is confirmed

Three frontier releases landed in three weeks of September 2026. DeepSeek published an open model that keeps a million-token conversation in under a gigabyte of memory. OpenAI shipped a mid-priced model it says nearly matches its flagship at a fifth of the price. Google announced a model that can write a million tokens in a single answer. Each headline hides an idea worth understanding, and those ideas will outlast the model names.

What was released?

DeepSeek V4.1 Flash (10 September) is an open-weights model, free to download under the MIT licence. DeepSeek's release note calls it a 552-billion-parameter mixture-of-experts model with “8B active parameters for input” and “16B for output”, built on a new “causal encoder-decoder” architecture, with image understanding built in. A technical report explains how it works, which is why this lesson can say far more about it than about the other two.

GPT-6.1 Sol (29 September, at OpenAI's DevDay) replaced GPT-6 Sol, which had shipped only a week earlier. OpenAI's model page promises “near-Astra performance for complex coding, computer use, and professional work”. GPT-6 Astra, OpenAI's flagship since 3 September, costs $10 per million input tokens and $50 per million output tokens; Sol costs $2 and $10, exactly one fifth (OpenAI pricing). The day before, reports said OpenAI had cancelled GPT-6.1 Astra after it “performed poorly on tests measuring alignment” and showed “higher levels of deception” in internal testing.

Gemini 4 Argon (30 September) comes with what Google calls an “industry-leading 1M tokens” output limit, “up from the previous 64K”. It is not generally available yet: it is rolling out first to “trusted cyber defenders” through a programme Google calls Fairwind, then to paid API customers and Google AI Ultra subscribers, with no date given. Introductory pricing is $2 / $10 per million tokens, rising to $4 / $20 later.

What do the benchmark claims say?

Each maker reports scores on DeepSWE v1.1, a software-engineering benchmark. Google reports 77.9% for Argon. SiliconANGLE reports DeepSeek's figure of 74.2% for V4.1 Flash. OpenAI says Sol matches GPT-6 Astra on it. Three caveats apply to all of them: each number is reported by the company selling the model; each is run at a setting the company chose, usually the highest thinking effort; and small differences between frontier models have become a weak guide to how they behave on your own work. Treat them as a reason to test, not as a ranking.

Key takeaways

  • DeepSeek V4.1 Flash is open, 552B parameters with 8B to 16B active, and its technical report explains the design; the other two are closed.
  • GPT-6.1 Sol costs one fifth of GPT-6 Astra per token; Gemini 4 Argon can write up to a million tokens in one answer but reaches defenders first.
  • Benchmark scores are self-reported at chosen settings. Cost per task depends on tokens used as well as price per token.

Total versus active parameters

“552 billion parameters” sounds like a giant. “16 billion active” sounds like a small model. DeepSeek V4.1 Flash is both, and understanding why tells you what that model costs to run.

What is the difference?

Total parameters are every weight the model stores. Active parameters are the weights actually used to process one token. In an ordinary dense model the two are the same. In a mixture-of-experts (MoE) model, each layer holds many small feed-forward networks, the experts, and a router sends each token to only a few of them (Shazeer et al., 2017).

How does it work in V4.1 Flash?

According to the technical report, each MoE layer has 384 routed experts plus one shared expert that every token uses, and the router picks 6 routed experts per token. So a token touches 7 of 385 experts in a layer, under 2% of them. Across the whole model that comes to about 16 billion of the 552 billion parameters, 2.9%. Its predecessor V4-Flash had 284 billion parameters with 13 billion active.

The new part is that the active count differs between reading and writing. The model has 40 layers: a 20-layer causal encoder followed by a 20-layer decoder. Prompt tokens run only through the encoder; the decoder's keys and values for them are projected directly from the encoder's final hidden state instead of being computed layer by layer. Tokens the model generates run through all 40 layers. That is where “8B for input, 16B for output” comes from.

Why does it matter?

Arithmetic follows active parameters: a forward pass costs about two operations per active parameter per token. Memory follows total parameters: all 552 billion must sit in GPU memory, because the next token may need any expert. So an MoE model is cheap in arithmetic and expensive in memory. Halving the active parameters for prompt tokens matters because agents re-read long prompts all the time, and DeepSeek's report is explicitly aimed at long-context agent workloads.

Train a small mixture-of-experts model and watch the router specialise in the MoE Lab

Key takeaways

  • Total parameters set the memory needed to hold the model; active parameters set the arithmetic per token.
  • V4.1 Flash routes each token to 6 of 384 experts plus a shared one, so about 16B of 552B parameters work on each generated token.
  • Its causal encoder-decoder runs prompt tokens through only the first 20 layers, halving active parameters for reading (8B).

The KV cache

DeepSeek titled its technical report “Pushing the Limits of KV Cache Compression”. The KV cache is the least famous part of a language model and often the part that decides how many people one GPU can serve.

What is the KV cache?

In a transformer, every layer turns every token into a query, a key and a value. To produce the next token, the newest query is compared with the keys of all earlier tokens and used to mix their values. Those earlier keys and values never change, so instead of recomputing them for every new token the model stores them: that store is the KV cache. It grows by one entry per layer for every token in the conversation, prompt and output alike.

How big is it?

For standard attention, the size per token is a simple product:

bytes per token = layers × KV heads × head dimension × 2 (key and value) × bytes per number

Each layer stores one key and one value vector per key-value head. Precision is 2 bytes at 16-bit, 1 at 8-bit, half a byte at 4-bit.

A worked example with a well-documented model. Llama 3 70B has 80 layers and uses grouped-query attention with 8 key-value heads of dimension 128. At 16-bit that is 80 × 8 × 128 × 2 × 2 = 327,680 bytes, about 0.33 MB for every token. A 128K-token conversation needs 43 GB of cache; a million-token one needs 328 GB, before counting the model's weights.

Grouped-query attention is itself a compression trick. With full multi-head attention each of Llama's 64 query heads would keep its own key and value, eight times more cache. Multi-query attention (one shared key-value head) and grouped-query attention (a few shared heads) were invented precisely to shrink this cache.

What did DeepSeek change?

DeepSeek has been attacking this number for years. DeepSeek-V2 introduced multi-head latent attention, which stores one compressed vector per token instead of full keys and values. The V4.1 Flash report stacks several more ideas on top:

  • Compression and sparse attention. Groups of tokens are compressed into single cache entries, and each query attends to only the 512 most relevant entries rather than all of them.
  • Sharing across layers. Some layers reuse the keys and values of an earlier layer instead of storing their own, and the decoder's cache for prompt tokens is projected from the encoder.
  • 4-bit storage. The main cache is stored in FP4: each number takes 4 bits (1 sign bit, 2 exponent bits, 1 mantissa bit), with an 8-bit scale shared by every 16 channels so the tiny format can still represent large and small values. This is the same idea as quantizing weights, applied to the cache.

The result, per the report: a global KV cache of 890 bytes per token, about a quarter of V4-Flash's, and a copy kept on SSD for reuse that is about an eighth of the previous size. A million-token conversation fits in 890 MB. The Llama 3 70B layout above needs 368 times as much.

At 128K tokens and 400 GB of free memory, the 16-bit Llama 3 70B layout fits 9 conversations. V4.1 Flash fits about 3,400. At a million tokens the dense layout fits one, and V4.1 Flash fits about 430. Halving the bytes per number doubles the count; cutting the KV heads from 64 to 8 multiplies it by eight.

Why does it matter?

Two reasons. First, batch size: a server makes money by generating tokens for many conversations at once, and every conversation in the batch needs its own cache in GPU memory. Memory for KV cache, not arithmetic, is often what caps the batch, as the vLLM paper showed. Second, speed: each new token must read the whole cache of its conversation, so a bigger cache means more bytes to read per step. A smaller cache makes long contexts cheaper and faster, which is exactly what agents that carry long histories need.

Key takeaways

  • The KV cache stores every earlier token’s keys and values so they are not recomputed; it grows with every token of context.
  • Bytes per token = layers × KV heads × head dimension × 2 × bytes per number; a 70B dense model needs about 0.33 MB per token at 16-bit.
  • V4.1 Flash’s 890 bytes per token comes from compression, sparse attention, cross-layer sharing and FP4 storage, and lets far more long conversations share a GPU.

Reading versus writing

Why would anyone design a model that uses 8 billion parameters to read and 16 billion to write? Because reading a prompt and writing an answer stress a GPU in completely different ways.

What are the two phases?

Prefill processes the prompt. All its tokens are known in advance, so the model can push thousands of them through each layer at once. Decode writes the answer, one token at a time, because each new token depends on the one before. A 10,000-token prompt is a handful of big parallel passes; a 10,000-token answer is 10,000 small passes in a row.

How do their costs differ?

Every pass has to move the weights it uses from GPU memory to the compute units and then do about two operations per active parameter for every token in the pass. A modern GPU can do a few hundred operations in the time it takes to read one byte, so the question is how much work each byte of weights is used for. That ratio is called arithmetic intensity.

  • In prefill, one read of the weights serves every prompt token in the pass. Intensity is high, so the GPU is compute-bound: the arithmetic is the bottleneck.
  • In decode, one read of the weights serves only one token per conversation, and each conversation's KV cache has to be read as well. Intensity is low unless many conversations are batched together, so the GPU is memory-bound: it waits for data.

This split is well documented (Pope et al., 2022), and some serving systems now run the two phases on separate GPUs for exactly this reason (DistServe, 2024).

step time ≈ max( bytes read ÷ memory bandwidth , operations ÷ arithmetic rate )

An idealised roofline: whichever term is larger sets the time.

Three things stand out. First, the dense model at 32K tokens of context never becomes compute-bound in decode, however large the batch: every conversation adds its own 10.7 GB of cache to read, so the bytes grow as fast as the arithmetic. Second, the MoE model with a batch of 64 loads about 244 of its 384 experts per layer, because different tokens choose different experts, so batching helps an MoE much less than a dense model; it needs thousands of tokens per step to keep its arithmetic busy, which is why MoE models are served at very large batch sizes across many GPUs. Third, that only works because V4.1 Flash's cache is tiny: at 890 bytes per token, 8,192 conversations of 32K tokens need 239 GB of cache, which fits.

The causal encoder-decoder fits this picture. Prompt processing is where agent workloads spend their tokens, re-reading long histories, and halving the active parameters for prompt tokens halves its arithmetic. DeepSeek has not published a breakdown of how much each change saves in practice.

Mixture of experts: routing, load balancing and why sparse models scale

Key takeaways

  • Prefill reads the prompt in large parallel passes and is usually compute-bound; decode writes one token per step and is usually memory-bound.
  • Decode reads the weights and every conversation’s KV cache each step, so a large cache can keep a server memory-bound at any batch size.
  • MoE models need very large batches to keep their arithmetic busy, which only works when the KV cache per conversation is small.

A million tokens out

Most models can read far more than they can write: a million tokens in, 64K or 128K out. Gemini 4 Argon's headline feature reverses that. What does a million-token answer actually involve?

What is an output limit?

The output limit is the most tokens a model will generate in one response. For reasoning models it includes the hidden thinking tokens, so a model that thinks for 50,000 tokens has that much less room for its answer. GPT-6.1 Sol's limit is 128,000 tokens; DeepSeek's API allows 384,000; Google says Argon's is one million, “up from the previous 64K”.

How long does it take?

Decode is sequential, so the time is simply tokens divided by speed. Independent measurements put GPT-6.1 Sol at about 56 tokens per second and V4.1 Flash at about 213; Argon's speed has not yet been measured. A million tokens at 60 tokens per second is 16,700 seconds: about 4 hours and 40 minutes for one response.

Why does it matter?

At introductory pricing a full million-token answer costs $10 in output tokens, and $20 at the standard rate: cheap compared with the work it might replace, such as patching a large codebase in one pass or drafting a long legal document. The harder costs are time and memory. Hours of decoding make long outputs a batch job, not a chat. And every token written joins the KV cache, so by the end the model is attending over a million tokens of its own writing; a dense 70B layout would need over 300 GB of cache for that one response.

A bigger limit also does not mean a model stays coherent for a million tokens. That is the claim to test. Google pitches Argon at long, autonomous jobs such as “autonomous cybersecurity patching”, which is also why its first users are cyber defenders.

Key takeaways

  • The output limit caps one response, thinking tokens included: 128K for GPT-6.1 Sol, 384K for DeepSeek’s API, 1M for Gemini 4 Argon.
  • Time is tokens divided by speed: a million tokens at 60 tokens per second is over four and a half hours, because decoding is sequential.
  • The bill for a million output tokens is $10 to $20 at Argon’s prices; the real constraints are latency, memory for the growing cache, and staying coherent.

Price per task

The price lists differ by more than eight times: $1.20 per million output tokens for V4.1 Flash at peak against $10 for Sol and Argon. The bills differ by less, because models do not use the same number of tokens for the same job.

What goes into a bill?

Three kinds of token, each with its own price: fresh input, cached input (a prompt prefix the provider has already processed recently, billed at 2% to 5% of the input rate), and output, which includes thinking. Two pricing rules matter here. OpenAI bills a Sol request whose prompt exceeds 272K tokens at $4 / $15 for the whole request (OpenAI pricing). DeepSeek halves its prices outside weekday peak hours (DeepSeek pricing).

cost = fresh input × input price + cached input × cache price + output × output price

List prices are per million tokens, so divide by 1,000,000.

How different are tokens per task?

Very. Running the same benchmark suite, Artificial Analysis recorded about 67 million output tokens for GPT-6.1 Sol at maximum effort, 110 million for Argon at high effort and 250 million for V4.1 Flash at maximum effort. V4.1 Flash used nearly four times Sol's tokens, which eats much of its eightfold price advantage on output. These are different effort settings on one mix of tasks, so treat them as an illustration, not a constant.

With a 50K-token prompt and 8K tokens of output, Sol and Argon at introductory pricing cost the same 18 cents and V4.1 Flash costs about 2.5 cents, seven times less. Switch on measured verbosity and the gap shrinks to about 3.5× against Sol, and Argon becomes the most expensive of the three. Raise the cached share and input almost disappears from the bill, which is why agents that resend the same long history should be designed to hit the cache.

See how tokens are generated one at a time in the LLM Sampling Lab

Key takeaways

  • Cost per task = tokens × price for fresh input, cached input and output; output, including thinking, costs 4 to 5 times as much per token as input.
  • Measured output tokens on the same benchmarks differ by nearly 4× between these models, so per-token price ratios overstate the real gap.
  • Pricing rules matter: Sol’s rate rises for prompts over 272K tokens, DeepSeek halves its prices off-peak, and Argon’s introductory price will end.

Choosing, and what to watch

None of these models is best at everything, and one of them is not yet available to most people. Here is how to think about the choice.

DeepSeek V4.1 FlashLowest price per token and a documented design built for long contexts. Open weights mean you can run it on your own hardware and keep data in-house. Check its quality on your tasks, its token usage, and the provider terms if you use DeepSeek’s hosted API.GPT-6.1 SolOpenAI’s claim is near-flagship results at a fifth of Astra’s price, with a 1.05M-token context. Keep prompts under 272K tokens to stay on the standard rate, and start at a lower effort setting before raising it.Gemini 4 ArgonThe only one that can write a million tokens in one response. For now it is limited to Fairwind participants; general availability, a model card and final pricing are still to come.

How should you decide?

The same way as for any model: build a small evaluation set of 20 to 50 real tasks with a way to score them, run each candidate at two or three effort settings, and record quality, total tokens, cost and time per task. The evaluation lesson covers how. Per-token price, benchmark tables and context limits only tell you which models are worth testing. If a reasoning model is involved, effort is usually the cheapest dial to try first; the reasoning models lesson explains why thinking longer helps.

What should you watch next?

  • Independent long-context tests of V4.1 Flash. Compressing the cache to 890 bytes per token is impressive only if recall over a million tokens holds up.
  • Argon's general release: when it opens to paid API customers, how long introductory pricing lasts, and whether long outputs stay coherent in practice.
  • OpenAI's next flagship. The cancelled GPT-6.1 Astra is a reminder that safety evaluations can stop a release; what OpenAI changes before the next one matters.
  • KV cache size as a headline number. As agents carry longer histories, memory per token of context is becoming as important as parameter count. Expect more labs to report it, and ask for it when they do not.

Why does this pattern keep repeating?

The scaling laws say more compute buys better models, but serving is where the money goes. Every release in this lesson is, in some way, a bet on spending less per useful token: sparse experts to cut arithmetic, a compressed cache to fit more conversations, a cheaper tier that nearly matches the flagship, and an output limit that lets one call do the work of many. The Opus 5.5 explainer describes another bet of the same kind.

Large language models: tokens, attention and how generation works from first principles

Key takeaways

  • Shortlist with published facts, then decide with your own evaluation: quality, tokens, cost and time per task at a few effort settings.
  • Open weights (V4.1 Flash) buy control and low prices; Sol buys near-flagship results cheaply; Argon’s long outputs are not broadly available yet.
  • Watch independent long-context tests, Argon’s general release and pricing, and whether labs start publishing KV cache size.

Check your understanding

Six scenarios, from sizing a server to reading a launch claim and budgeting a workload.

Question 1 of 6

A team self-hosts a model for a coding agent that keeps 200K-token histories. Their GPUs have plenty of arithmetic to spare, but they can only run 6 conversations at once before memory runs out. Which change would most directly let them serve more conversations?

References

Facts about the three models come from each maker's announcement, documentation, pricing page or technical report as of 5 October 2026, with news coverage cited where it adds context. Speed and token counts are independent measurements by Artificial Analysis and will change. The background on attention, KV caches and serving comes from peer-reviewed papers and preprints.

Sources

  1. [1]

    DeepSeek-V4.1-Flash release(opens in a new tab)

    DeepSeek, 2026

    Release note (10 September 2026): a 552B-parameter MoE with 8B active parameters for input and 16B for output, a new causal encoder-decoder architecture, native visual understanding, 1/4 the HBM and 1/8 the SSD storage for KV cache compared with the previous generation.

  2. [2]

    DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression(opens in a new tab)

    DeepSeek-AI, 2026

    Technical report: 20-layer causal encoder plus 20-layer decoder, 384 routed experts with 6 active plus 1 shared, FP4 KV cache, a global KV cache of 890 bytes per token (about 1/4 of V4-Flash), contexts up to one million tokens.

  3. [3]

    GPT-6.1 Sol model page(opens in a new tab)

    OpenAI, 2026

    Developer documentation: “near-Astra performance” for coding, computer use and professional work; 1,050,000-token context, 128,000-token output, $2 / $10 per million tokens.

  4. [4]

    API pricing(opens in a new tab)

    OpenAI, 2026

    GPT-6.1 Sol: $2 / $10 per million tokens up to 272K input tokens, $4 / $15 above. GPT-6 Astra: $10 / $50.

  5. [5]

    OpenAI cancels GPT-6.1 Astra release over misbehavior, safety concerns(opens in a new tab)

    9to5Google, 2026

    Reports, citing OpenAI’s head of safety systems, that GPT-6.1 Astra “performed poorly on tests measuring alignment” and showed “higher levels of deception”.

  6. [6]

    Introducing Gemini 4 Argon(opens in a new tab)

    Kavukcuoglu, K. (Google DeepMind), 2026

    Announcement (30 September 2026): a 1M-token output limit, up from 64K; $2 / $10 introductory pricing, then $4 / $20; DeepSWE v1.1 77.9%; rolling out first to cyber defenders through the Fairwind Program.

  7. [7]

    Models and pricing(opens in a new tab)

    DeepSeek, 2026

    API prices for deepseek-flash: $0.30 input and $1.20 output per million tokens at peak, half that off-peak; 1M context, 384K maximum output.

  8. [8]

    DeepSeek-V4.1-Flash model card(opens in a new tab)

    DeepSeek-AI, 2026

    Open weights under the MIT licence, with the parameter counts, KV cache format (FP4, E2M1 with an E4M3 scale per 16 channels) and vision encoder.

  9. [9]

    DeepSeek releases V4.1-Flash, says it outperforms flagship V4-Pro(opens in a new tab)

    SiliconANGLE, 2026

    News coverage: V4-Flash had 284B parameters; reports DeepSeek’s benchmark claims, including 74.2% on DeepSWE v1.1.

  10. [10]

    Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer(opens in a new tab)

    Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., Dean, J., 2017

    Introduces the sparsely gated mixture-of-experts layer: a router sends each token to a few of many expert networks. ICLR 2017.

  11. [11]

    The Llama 3 Herd of Models(opens in a new tab)

    Grattafiori, A. et al. (Meta), 2024

    Architecture table for Llama 3 70B: 80 layers, 64 query heads, 8 key-value heads (grouped-query attention), model dimension 8,192.

  12. [12]

    Fast Transformer Decoding: One Write-Head is All You Need(opens in a new tab)

    Shazeer, N., 2019

    Multi-query attention: all query heads share one key and value head, shrinking the KV cache and speeding up memory-bound decoding.

  13. [13]

    GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints(opens in a new tab)

    Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., Sanghai, S., 2023

    Grouped-query attention: groups of query heads share a key-value head, between full multi-head and multi-query attention. EMNLP 2023.

  14. [14]

    DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model(opens in a new tab)

    DeepSeek-AI, 2024

    Introduces multi-head latent attention, which stores a compressed latent vector per token instead of full keys and values, cutting the KV cache by 93% against DeepSeek 67B.

  15. [15]

    Efficient Memory Management for Large Language Model Serving with PagedAttention(opens in a new tab)

    Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., Stoica, I., 2023

    vLLM: shows that KV cache memory limits the batch size of LLM serving, and manages it in pages to fit more requests. SOSP 2023.

  16. [16]

    Efficiently Scaling Transformer Inference(opens in a new tab)

    Pope, R., Douglas, S., Chowdhery, A., Devlin, J., Bradbury, J., Levskaya, A., Heek, J., Xiao, K., Agrawal, S., Dean, J., 2022

    Analyses serving cost: prefill can run at high utilisation, while decoding is bound by memory bandwidth for weights and KV cache unless batches are large.

  17. [17]

    DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving(opens in a new tab)

    Zhong, Y., Liu, S., Chen, J., Hu, J., Zhu, Y., Liu, X., Jin, X., Zhang, H., 2024

    Runs prefill and decoding on separate GPUs because the two phases have different bottlenecks and latency targets. OSDI 2024.

  18. [18]

    GPT-6.1 Sol (max): intelligence, performance and price analysis(opens in a new tab)

    Artificial Analysis, 2026

    Independent measurements: 67M output tokens to run the Intelligence Index, 55.7 output tokens per second. Figures change as the provider’s serving changes.

  19. [19]

    DeepSeek V4.1 Flash (max): intelligence, performance and price analysis(opens in a new tab)

    Artificial Analysis, 2026

    Independent measurements: 250M output tokens to run the Intelligence Index, 213.2 output tokens per second.

  20. [20]

    Gemini 4 Argon (high): intelligence, performance and price analysis(opens in a new tab)

    Artificial Analysis, 2026

    Independent measurements: 110M output tokens to run the Intelligence Index against a median of 81M; output speed not yet available.

  21. [21]

    The latest AI news we announced in September(opens in a new tab)

    Google, 2026

    Monthly round-up: Argon is “built with an industry-leading 1-million-token output limit” and needs “a phased approach” to release.

Related