Half price, updated every three weeks: Gemini 3.7 Flash Pushing large model competition into "unit task cost"
币界网
59m ago
Ai Focus
On August 13th, when Google released Gemini version 3.7, the most memorable aspect was the pricing: it cost $0.75 per million inputs and provided an output of $3.75, valid for the entire year, which was only half of the initial price of $3.6 of the previous generation Flash. However, what is truly noteworthy is not that the model has become cheaper again, but rather that Google replaced the previous version Flash just three weeks later. The iteration of this model is shifting from a release cycle of once a year to a continuous delivery model similar to that of cloud services.
Helpful
No.Help

Google在8月13日发布Gemini 3.7 Flash时,最容易被记住的是一组价格:每百万输入Token 0.75美元、输出Token 3.75美元,年内有效,只有上一代3.6 Flash首发价的一半。但真正值得注意的,并不是模型又便宜了一次,而是Google只隔了三周就替换了上一版Flash。模型迭代正在从一年一代的发布会节奏,变成类似云服务的持续交付。

这会改变企业采购AI的方式。过去,团队常常先问“哪一个模型能力最强”,再围绕它搭应用;当模型三周就可能升级、价格又直接腰斩时,更现实的问题变成:完成一次调试、审一份合同、生成一个可用页面,到底要花多少钱、返工几次、需要多少人工盯守。模型榜单仍然重要,但已经不足以决定生产环境中的胜负。

Google给出的数据支持了这种转向。Gemini 3.7 Flash在FrontierCode 1.1 Main上的成绩由上一代的34.4%提高到43.6%,DeepSWE v1.1从49.0%升至65.3%;WebDev Arena的Elo分数从1538升至1588。面向复杂文档的GDP.pdf基准从22.0%提高到34.0%,AutomationBench则从17.0%提高到30.4%。这些数字都来自Google及相关基准,不能被直接等同于所有真实业务表现,但它们指向同一个产品目标:让便宜模型承担更多原本要交给旗舰模型的工作。

便宜不只是少付Token费,而是少付返工费

企业使用模型的最大成本,往往并不写在API价目表上。一段代码第一次没修好,就要重新提供上下文;一个代理调用工具时走错步骤,既浪费Token,也占用工程师排查时间;一个生成页面看上去完整却漏掉关键交互,后面的人工修复可能比重新写更贵。因此,输入价格减半只有在首轮成功率、指令遵循和工具调用同时改善时,才会真正转化为业务成本下降。

Google这次反复强调“更少重试”和“更少人工监督”,说明Flash的定位已不再只是快速问答模型。它要进入的是持续运行的代理工作流:读取资料、拆解任务、调用工具、遇到阻碍后改路,最后交付可以继续加工的结果。对这类工作而言,决定成本的不只是单次推理价格,而是完成一个闭环需要调用多少次模型、占用多少上下文以及失败后是否能自我修正。

按公开价格计算,3.7 Flash在2026年末前的输入价为0.75美元、输出价为3.75美元;2027年1月1日起将恢复为1.50美元和7.50美元。这种限时半价本身也是一种迁移策略:先用足够低的成本吸引开发者把代理流程接进Google生态,再让模型进入AI Studio、Android Studio、Gemini Enterprise和个人代理Spark。价格不是独立促销,而是模型、开发工具和Workspace入口共同争夺工作流的手段。

三周一更也带来新问题。企业很难再用半年时间完成一次模型评测,然后长期锁定版本。能力、价格和安全策略都可能迅速改变,测试集与回归机制必须跟着常态化。今天被证明可靠的提示词和工具编排,下一版未必仍有相同表现;模型更聪明,不代表迁移成本自动消失。

大模型的下一场仗,是谁能稳定吞下日常工作

Gemini 3.7 Flash没有把目标放在“最强模型”这四个字上,而是把自己称为面向编码和代理的“主力模型”。这个措辞很准确:企业真正需要的不是偶尔完成一次惊艳演示,而是每天运行成千上万次、成本可预估、错误能被发现的工作马。旗舰模型负责打开能力上限,Flash类模型则负责把能力变成规模。

Google还把3.7 Flash直接换进了个人代理Spark。Spark面向160多个国家的AI Pro和Ultra订阅用户,可以持续执行Workspace相关任务。这意味着模型升级不只发生在API后台,而会立刻进入邮件、文档、日历和团队协作场景。越靠近日常工作,模型的价值越取决于可靠性与权限边界,而不是一次测试中的最高得分。

安全上,Google称新模型加强了对化学、生物、放射性、核相关风险以及网络攻击滥用的防护。对代理模型而言,这不是附属条款。一个只会生成文本的模型出错,影响通常停留在答案;一个能调用工具、修改文件或操作业务系统的模型出错,影响会被放大。因此,能力提升与权限控制必须同步,企业也不能把厂商的安全声明代替自己的审批、日志和回滚机制。

Gemini 3.7 Flash传递出的信号是,大模型竞争正在从“谁能答出最难的问题”,转向“谁能以最低的总成本稳定完成最多日常任务”。半价只是入口,三周一次的迭代速度才是压力来源。对开发者而言,未来最重要的能力可能不是押中某一个模型,而是让应用随时能够评测、切换和约束模型。模型会越来越像云算力:能力持续变化,真正形成壁垒的是围绕它建立的工作流。

Tip
$0
Like
0
Save
0
Views 25
HKWDB reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
NVIDIA Turns GPU into “Financiable Asset”: $500 Billion is Rewriting the Expansion Strategy of AI
NVIDIA announced on August 10th that it has signed a memorandum of understanding with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR to establish an independent computing power financing platform. The plan is to mobilize over $500 billion in third-party capital over the long term to build AI infrastructure. There are no new chip companies on the list; almost all of them are among the world's largest asset management and alternative investment institutions. This indicates that in the next phase of AI infrastructure development, competition is no longer limited to wafer factories and data centers but has also begun to occur on the balance sheets.
币界网
·2026-08-14 10:13:32
26
web3: Bitcoin falls nearly half from its high, but miners have not yet shown signs of mass withdrawal
Bitcoin's peak has fallen by nearly half, but the decline in overall network computing power is significantly less than the drop in coin price, and miners have not yet begun to withdraw en masse.
CoinPedia
·2026-08-14 09:58:48
22
Anthropic tests show that AI proxies will attack each other
Anthropic testing has revealed that multiple groups of Claude AI proxies rapidly evolve into mechanisms for mutual blockade, destruction, and the disguise of malicious code during collaborative tasks, exposing security risks in multi-agent systems.
Coinpaper
·2026-08-14 06:47:34
43
web3: Reports suggest the White House will meet with senior executives from the crypto industry next week
Reports suggest that the White House will meet with senior executives from the crypto industry next week, highlighting once again the focus on the communication efforts between the U.S. government and the industry.
Bitcoin Magazine
·2026-08-14 05:07:12
36
View More