Realtime AI News
Microsoft Halts 'Tokenmaxxing' With Strict Budget Caps; GPT-5.6 Becomes Internal Default
Chinese tech outlet QbitAI reports that Microsoft has halted internal "Tokenmaxxing" — pushing token usage to its limits — and locked AI budgets, with over-limit usage now at employees' own risk. The report adds that GPT-5.6 has become Microsoft's default internal model.
Chinese tech outlet QbitAI reports that Microsoft has halted "Tokenmaxxing" internally — the practice of pushing token usage to its absolute limit — and locked down AI budgets.
According to the report, usage beyond the budget cap is now at employees' own risk, meaning internal teams can no longer burn tokens without restraint.
The report adds that GPT-5.6 is now Microsoft's default internal model for everyday AI work.
Token consumption is becoming the biggest variable in enterprise AI spending, with long-context and agent-driven tasks rapidly inflating the cost of a single run.
As a major OpenAI investor and one of the world's largest cloud providers, Microsoft's internal cost discipline is often read as a signal for AI spending across the industry.
For developers, hard budget caps mean prompt design, context length, and agent call counts now need careful planning.
The next question is whether Microsoft ships finer-grained quota management tools, and whether other tech giants follow with similar limits.
Why it matters
Microsoft's internal budget clampdown and default-model switch mark cost control as a hard constraint for big tech, likely prompting industry-wide copycats.
Nearby Updates
All08/05, 15:05
Sand.ai Open-Sources What It Calls the First 100B-Parameter MoE Video Generation Model
Sand.ai has open-sourced a Mixture-of-Experts video generation model it bills as the world's first 100-billion-parameter MoE video model, with 114B total parameters and just 6B active. The model generates 10-second 1080p clips at a reported cost of about 0.5 yuan each, sharply lowering the cost barrier for high-quality AI video.
08/05, 13:59
ByteDance Seed Unveils SeedRealtime: Full-Duplex Audio-Video Model Debuts in Doubao
ByteDance's Seed team has released SeedRealtime, a full-duplex audio-video large model now integrated into the Doubao app. The model lets users watch, listen, and speak simultaneously, removing the awkward pauses of turn-based voice assistants and pushing real-time multimodal interaction into a mainstream consumer product.
08/05, 13:53
AI Cracks NVIDIA's 20-Year CUDA Moat in About 10 Hours, QbitAI Reports
QbitAI reports that AI has cracked NVIDIA's CUDA moat — built over 20 years by Jensen Huang — in about 10 hours. The report asks whether "CUDA is in danger again," arguing that AI is sharply lowering the barrier to bypassing or replacing the CUDA ecosystem.
08/05, 16:04
NVIDIA Nemotron 3 Ultra Highlights Major AI Agent Advances
Coverage from Techgenyz spotlights NVIDIA's Nemotron 3 Ultra as a major step forward in AI agent capabilities. The release positions agent-focused abilities as the headline advance in NVIDIA's high-end open model line, with benchmark results and technical specifics still pending.