Realtime AI News
Qualcomm Launches Two New Smartphone Chips With AI Focus
Qualcomm has launched two new smartphone chips, placing on-device AI at the centre of the announcement. The company says its top-end chip can run a 30B mixture-of-experts model locally, pushing larger model inference onto the phone itself.

Qualcomm has launched two new smartphone chips, placing on-device AI at the centre of the announcement.
The company says its top chip can run a 30B mixture-of-experts model locally. The operative word is locally: inference happens on the handset rather than round-tripping to a data centre.
Mixture-of-experts architectures are well matched to that constraint. The model as a whole is large, but only a fraction of it is activated for any given request, so capacity can grow while per-request compute stays relatively contained.
The 30B figure deserves its own look. Models that ran reliably on phones have generally sat in the single-digit billions, so pushing to 30 billion widens the range of tasks a device can plausibly handle.
For users, the most immediate consequences are privacy and availability. Local inference does not require uploading input to a server and keeps working offline or on poor connections; for device makers, it moves some inference cost off the cloud.
Chip specifications and real experience are still different things. Throughput, power draw, heat and the degree of quantisation applied to the model all determine whether that capability is genuinely usable day to day, which shipping devices will have to demonstrate.
What to watch next is which handset makers adopt the two chips first, and whether developers build applications that exploit a local 30B-class model in ways cloud inference cannot match.
Why it matters
A higher on-device model ceiling will pull more inference away from the cloud and onto handsets. That reshapes both the privacy and availability story and the cost split between phone makers and cloud providers.
Nearby Updates
All09/23, 03:09
Meta admits Muse's likeness to OpenClaw isn't a coincidence
TechCrunch reports that Meta has acknowledged the resemblance between its AI assistant Muse and OpenClaw is not a coincidence, saying Muse was heavily inspired by OpenClaw. According to the report, that similarity extends to some workspace filenames and content.
09/23, 05:00
OpenAI Improves Prompt Caching for GPT-6
OpenAI has detailed a set of prompt caching improvements for GPT-6, including higher cache hit rates, new diagnostics, explicit breakpoints and additional controls. The changes aim to cut latency and cost for repeated prompts while making cache behaviour easier to observe and steer.
09/23, 02:00
OpenAI launches GPT-6 Sol and Luna, pitching lower cost and fewer mistakes
OpenAI has launched two new models, GPT-6 Sol and GPT-6 Luna, describing them as cut from the same cloth as Astra while pitching lower cost and fewer mistakes. TechCrunch reports the pair arrived together, a sign that price and reliability, not just benchmark scores, are now the headline claims of a frontier release.
09/23, 00:30
Anthropic releases Opus 5.5 with lower prices and Fable-level performance
Anthropic has released Opus 5.5, describing it as the strongest-performing model the company has tested to date. The new flagship arrives with lower pricing and performance the report places at Fable level, putting cost alongside capability as a headline claim.