Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

Microsoft Halts 'Tokenmaxxing' With Strict Budget Caps; GPT-5.6 Becomes Internal Default

Chinese tech outlet QbitAI reports that Microsoft has halted internal "Tokenmaxxing" — pushing token usage to its limits — and locked AI budgets, with over-limit usage now at employees' own risk. The report adds that GPT-5.6 has become Microsoft's default internal model.

Published
微软叫停Tokenmaxxing:预算卡死、超限自负,内部默认用GPT-5.6
Image source: microsoft.com

Chinese tech outlet QbitAI reports that Microsoft has halted "Tokenmaxxing" internally — the practice of pushing token usage to its absolute limit — and locked down AI budgets.

According to the report, usage beyond the budget cap is now at employees' own risk, meaning internal teams can no longer burn tokens without restraint.

The report adds that GPT-5.6 is now Microsoft's default internal model for everyday AI work.

Token consumption is becoming the biggest variable in enterprise AI spending, with long-context and agent-driven tasks rapidly inflating the cost of a single run.

As a major OpenAI investor and one of the world's largest cloud providers, Microsoft's internal cost discipline is often read as a signal for AI spending across the industry.

For developers, hard budget caps mean prompt design, context length, and agent call counts now need careful planning.

The next question is whether Microsoft ships finer-grained quota management tools, and whether other tech giants follow with similar limits.

Why it matters

Microsoft's internal budget clampdown and default-model switch mark cost control as a hard constraint for big tech, likely prompting industry-wide copycats.

MicrosoftGPT-5.6AI Budget
Back to realtime news

Nearby Updates

All

08/05, 15:05

Sand.ai Open-Sources What It Calls the First 100B-Parameter MoE Video Generation Model

Sand.ai has open-sourced a Mixture-of-Experts video generation model it bills as the world's first 100-billion-parameter MoE video model, with 114B total parameters and just 6B active. The model generates 10-second 1080p clips at a reported cost of about 0.5 yuan each, sharply lowering the cost barrier for high-quality AI video.

08/05, 13:59

ByteDance Seed Unveils SeedRealtime: Full-Duplex Audio-Video Model Debuts in Doubao

ByteDance's Seed team has released SeedRealtime, a full-duplex audio-video large model now integrated into the Doubao app. The model lets users watch, listen, and speak simultaneously, removing the awkward pauses of turn-based voice assistants and pushing real-time multimodal interaction into a mainstream consumer product.

08/05, 13:53

AI Cracks NVIDIA's 20-Year CUDA Moat in About 10 Hours, QbitAI Reports

QbitAI reports that AI has cracked NVIDIA's CUDA moat — built over 20 years by Jensen Huang — in about 10 hours. The report asks whether "CUDA is in danger again," arguing that AI is sharply lowering the barrier to bypassing or replacing the CUDA ecosystem.

08/05, 16:04

NVIDIA Nemotron 3 Ultra Highlights Major AI Agent Advances

Coverage from Techgenyz spotlights NVIDIA's Nemotron 3 Ultra as a major step forward in AI agent capabilities. The release positions agent-focused abilities as the headline advance in NVIDIA's high-end open model line, with benchmark results and technical specifics still pending.