Realtime AI News
Google Cuts Gemini 3.7 Flash API Prices by 50%
Google has cut API prices for its Gemini 3.7 Flash model by 50%, halving the cost for developers calling the fast-response model. The move is expected to accelerate adoption of high-frequency AI applications and agent workloads.
Google has cut API pricing for its Gemini 3.7 Flash model by 50%, a significant adjustment in the company's AI model pricing strategy.
The price reduction was reported by tech-insider.org, with the headline fact being straightforward: Gemini 3.7 Flash API costs are now half of what they were. For teams building on large-model APIs, this is an immediately visible change to their cost structure.
Gemini 3.7 Flash is the speed-and-value focused tier of the Gemini family, designed for low-latency, high-volume inference. With the price cut, it becomes noticeably more economical for batch processing, real-time interactions, and agent-driven calls.
For developers, halved token costs mean the same budget can support roughly twice the call volume. Teams building chat applications, automating workflows, or running large-scale data processing will all feel the impact of this repricing.
The move comes amid intensifying competition in model API pricing, where major providers are frequently adjusting their price strategies. By cutting prices quickly on the Flash tier, Google is responding to developer cost concerns while reinforcing the appeal of its developer ecosystem.
Pricing is often the decisive variable in scaling AI applications. Lower inference costs reduce the barrier to experimentation and may encourage more small and mid-sized teams to embed large-model capabilities into everyday products.
The key question going forward is whether other Gemini models will see similar adjustments and how rivals respond to this round of cuts. For teams evaluating model APIs, now is a good moment to reassess cost structures.
Why it matters
A 50% price cut on Gemini 3.7 Flash significantly lowers developer costs, likely accelerating high-frequency AI applications and agent deployments while putting fresh pressure on competitors to match pricing.
Nearby Updates
All08/30, 16:50
Alibaba Cloud Shows Keen Interest in Punjab's IT Sector; CM Maryam Invites Chinese AI Collaboration
Alibaba Cloud has expressed keen interest in Punjab's IT sector, and Punjab Chief Minister Maryam has invited Chinese partners to collaborate on artificial intelligence, according to Pakistan's Associated Press of Pakistan. No investment amounts or project details have been disclosed, leaving the scope of any future cooperation open.
08/30, 16:00
Beijing Bar Serves Drinks With a Side of Free AI Tokens — 'AGI Bar' Lets Customers Tap Any Model
A bar in Beijing's Zhongguancun tech hub is offering free AI model tokens to customers who order a drink, with a plug-in that connects laptops to the bar's own AI agent. The AGI Bar also runs DeepSeek-V4-Flash locally on two Nvidia DGX Spark machines, selling its bestselling AGI beer for just 9.9 yuan (US$1.47).
08/30, 14:12
Anthropic Opens Research Preview of Model Hardware Standard (MHS) for AI Agents Operating Physical Devices
Anthropic has opened a research preview of the Model Hardware Standard (MHS), a shared specification designed for AI agents to safely operate physical devices. The move aims to unify the interface between frontier models and real-world hardware, lowering the barrier for agents entering robotics and industrial applications.
08/30, 11:00
Zhipu GLM-5.3-Flash Details Surface: 320B Parameters, 18B Active, Hybrid Attention Built to Activate Domestic Compute
A new broker report from Guolian Minsheng Securities details how Zhipu's next-generation GLM-5.3-Flash model is architected to fully activate domestic Chinese compute. With roughly 320B total parameters but only 18B active, it compresses active parameters from 32B and pairs linear attention with sparse attention via an IndexPool indexer.