FlorianIO

JOURNAL

Thirty Cents a Million: DeepSeek V4.1 Flash and the Collapsing Cost of Intelligence

2026.09.1220 MIN READ

The Price Tag Fell Off

Now the part I actually opened my spreadsheet for — and this time I'm quoting the official rate card, not an aggregator's screenshot. DeepSeek kept the peak/off-peak split from their older models: off-peak is exactly half of peak, and the official definition of peak hours is 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday. Read that again: every evening, every night, and every single weekend on Earth is off-peak. Fellow nights-and-weekends developer, keep that in your pocket for later. Here's the full card for deepseek-flash, per million tokens:

  • Input, cache miss: $0.30 peak / $0.15 off-peak
  • Input, cache hit: $0.006 peak / $0.003 off-peak
  • Output: $1.20 peak / $0.60 off-peak

Now hold that against the outgoing flagship's own official card. V4 Pro, the model being retired, listed at $1.32 input / $3.96 output at peak — so the "up to 60%" that launch coverage led with (IT之家) actually undersold it. Official card against official card, the real cut is 77% on input, 70% on output, and 86% on cached reads. DeepSeek didn't discount their flagship. They replaced it with a smaller model that costs a quarter as much to read and a third as much to write, then walked the expensive one off the stage.

But rate cards don't hit until you put them in a league table, so here's the same job, standard tier, quoted from each lab's own pricing page:

  • GPT-6 Astra (OpenAI): $10.00 input, $1.00 cached, $50.00 output — and past 272K tokens of context it doubles to $20 / $2 / $75.
  • Claude Fable 5.1 (per Startup Fortune's launch coverage — Anthropic's pricing page won't load from where I sit): $10 input, $0.25 cached, $50 output.
  • Gemini 3.8 Flash, the other "Flash" and the actual cheap-tier peer (Google): $0.75 input, $3.75 output on a promo rate through the end of 2026, cached reads $0.075 — plus an hourly fee just to keep the cache warm.
  • V4.1 Flash: $0.30 / $0.006 / $1.20.

Against the $10/$50 flagships, that's 33x cheaper to read with, 42x cheaper to write with, and up to 167x cheaper on cached tokens — Astra charges a full dollar for the cached million that DeepSeek sells for six-tenths of a cent. Even against its closest peer, the other Flash, it's still 2.5x on input, 3.1x on output, and 12.5x on cache, with no storage meter running.

And a league table still doesn't hit until you price real work. So here's my cost ledger, worked out by hand at official peak rates — halve everything if you run it off-peak:

Reading. A million input tokens is roughly an entire mid-sized codebase, or about 1,500 pages of text. V4.1 Flash reads all of it for $0.30. Astra charges $10 for the same read — and $20 if your codebase doesn't fit inside its short-context window. Re-read it the next day and V4.1 Flash's cache drops that to $0.006. That's not a typo. Six-tenths of one cent.

Writing. A 100,000-word novel draft comes out to about 130K output tokens: roughly $0.16 at peak, eight cents off-peak. The same draft on Astra or Fable 5.1: about $6.50. On Gemini 3.8 Flash, the budget peer: $0.49. I checked that math three times because it felt fake. It's not fake.

Running an agent. This is the one that matters for me, because agents are mostly re-reading the same long system prompt thousands of times — DeepSeek's own announcement singles out cache-hit charges as a major expense in agent workloads, which is exactly the bill this architecture was engineered to shrink. Say your agent carries a 10K-token prompt and runs 1,000 calls a day — that's 10M cached input tokens daily. V4.1 Flash: $0.06 a day, call it $1.80 a month. Astra's cached rate: $10 a day, $300 a month, for identical tokens. Claude's brand-new $0.25 cached read — the cut that got celebrated as making agentic workloads up to 45% cheaper — lands at $2.50 a day, $75 a month. Same fortnight, same tokens, and the gap between the cheap labs and the expensive ones is still 42 to 167 times wide. (☞゚ヮ゚)☞ One honesty note on my own math, though: that $1.80 counts cached input only. Real agents generate as they go, and output at $1.20 a million is where actual bills grow. The cache sets the floor of the bill, not the ceiling.

What real bills look like. Sticker arithmetic is one thing; reported bills are another, and the samples are finally coming in. The developers who moved their daily coding off GitHub Copilot and posted their numbers describe about $20 covering more than a month of heavy use — one logged over five billion tokens on it (r/GithubCopilot) — and the before/after stories across r/LocalLLaMA sound the same note: monthly API bills falling from the $80s to the low $20s. But run it the way agents actually run it — long sessions, thousands of calls, real output at every step — and the meter is plainly visible: V2EX users compare notes on burning about ¥30 in a three-hour crawl-and-generate binge, r/DeepSeek's heavy agentic sessions land in the same few-dollars-per-hour zone, and a nine-day billing analysis on Zhihu put it in perspective — switching from V4 Pro to V4.1 Flash cut their bill from ¥534 to ¥254, which is dramatic, and is still about $35 for nine days. Light users land at a couple of dollars a month. Heavy agentic users land at $20-and-up. The card is cheap. The workload decides the bill.

And one more thing the launch coverage flattened, which the community has not: the price history here has three chapters, and the story changes depending on which one you start from. Chapter one is the original V4 Flash — the famously cheap card at roughly $0.14 input / $0.28 output (Ofox). Chapter two is the hikes: DeepSeek raised rates across the lineup before this launch, hard enough that "massive price increase" threads ran all over r/DeepSeek, and that hiked card is what heavy users have actually been paying lately. Chapter three is this week: V4.1 Flash's card is a genuine cut against the hiked rates, so if you've been living on post-hike prices, your bill really did just fall. But stack it against chapter one — the pre-hike Flash, the golden-age card people still quote — and it's roughly double the input and up to four times the output at peak. Cheaper than last month, more expensive than the good old days, dramatically cheaper than everyone else. All three are true at once, and you should know all three before you migrate anything.

So where does that leave someone like me? Right where the samples say. My own usage sits in the light band — a few million tokens a month is a couple of dollars at peak rates, and the usage calculators agree. And here's the part that feels personally targeted, in the good way: the cheap hours are nights and weekends, and nights and weekends are exactly when I build. For a solo developer on a $0 stack, the discount tier isn't a compromise. It's my calendar. Just know which band you're in before you build on this: the card is cheap, the workload decides the bill, and output — not cache — is where the meter really runs.

And one naming footnote, because I can't help myself after the last post: Gemini Flash, now DeepSeek Flash. "Flash" is quietly becoming the industry's word for "the cheap fast one," and I could not be happier about the vocabulary. The budget tier finally got a cool name.

A rate card tells you what tokens cost, I've learned. Only your workload knows what they're worth.

RELATED