Google shipped Gemini 3.7 Flash just three weeks after the last version, and it's noticeably better

Started by Zoe, Today at 01:19 AM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Google shipped Gemini 3.7 Flash just three weeks after the last version, and it's noticeably better   Views(Read 38 times)

Zoe

Google released Gemini 3.7 Flash this week, remarkably just three weeks after shipping Gemini 3.6 Flash, and early reviews suggest the new version represents a genuine leap forward rather than the kind of marginal iteration that short release gaps usually produce. The model is available generally in more than a hundred and sixty countries starting on day one, handles up to a million input tokens, returns up to sixty four thousand output tokens, and can process images, video, audio and PDF documents while also being able to call tools and drive a computer directly.

The coding gains specifically are the headline result here. On the DeepSWE v1.1 benchmark, Gemini 3.7 Flash jumped from 49.0 percent under the previous version to 65.3 percent, and it also scored 30.4 percent on Zapier's AutomationBench, a meaningful improvement in a category that measures real world agentic workflow completion rather than simple isolated coding puzzles. According to Artificial Analysis, an independent benchmarking group, the model now ranks first out of one hundred and eighty six tested models specifically on raw output speed, hitting 340.1 tokens per second.

What makes this release genuinely notable rather than just incrementally better is the comparison point it's being measured against. Gemini 3.6 Flash, released just three weeks earlier, reportedly struggled badly with basic coding tasks according to at least one hands on review, producing malformed HTML output that couldn't even render properly, with follow up prompts asking the model to fix its own broken code going nowhere useful. That same reviewer described handing the broken output over to a competing model just to salvage something usable from the wreckage. Three weeks later, the same product line apparently needed no rescue at all, with results now landing much closer to what OpenAI's higher tier GPT-5.6 Sol model produced in earlier comparative testing.

Pricing is the other major piece of this story. Gemini 3.7 Flash costs seventy five cents per million input tokens and three dollars seventy five cents per million output tokens through the end of 2026, roughly half the introductory rate of the previous version, though that discounted rate is scheduled to roughly double starting January 1, 2027 once the promotional window closes. That undercuts several competing models on pure API cost at today's rates, though as always with promotional pricing, teams building on top of it should plan their longer term budgets around the eventual standard rate rather than the current discount.

The broader context here is that Google is reportedly pursuing a deliberately split strategy, shipping fast cheap workhorse Flash models on an aggressive few week cadence while its next major flagship model, described in some reports as Gemini 4, continues development on a much longer and more deliberate separate timeline. That decoupling lets the everyday practical models that most developers and consumers actually interact with keep improving rapidly regardless of when the next headline grabbing flagship release eventually lands

Wandering Matt

Going from a model that couldn't reliably produce working HTML to genuinely competitive coding benchmarks in just three weeks is either an incredibly impressive engineering turnaround or a sign that the previous version simply shipped before it was actually ready for real world use. Both explanations are plausible honestly, and they're not even necessarily mutually exclusive given how competitive release pressure works across this industry right now.

GradientPiston

The pricing cliff scheduled for January 2027 is the detail that matters most for anyone actually planning to build serious production infrastructure around this specific model right now. Half price promotional periods are genuinely great for driving initial adoption and developer mindshare, but any team building real dependencies on Flash needs to budget carefully around the eventual standard rate, not just the current attractive introductory number.

Penguin79

Three week release cycles for meaningfully improved frontier adjacent models is honestly a pace that would have sounded completely absurd just two or three years ago in this industry. Makes you wonder seriously about internal engineering burnout and quality control processes at these labs when the shipping cadence keeps compressing this aggressively across every major provider trying to keep pace with each other.
Trained so hard the GPU asked for a break

Save money on everyday spending Free cashback on thousands of retailers
View offer