Google’s Gemini 3.5 Flash, shown off at I/O in late May, is the kind of release that doesn’t make front pages but quietly resets budgets. On the Artificial Analysis Intelligence Index it scores around 55 — comfortably ahead of several models that cost many times more — while running at roughly 284 tokens a second.
The headline isn’t raw capability. It’s the cost-to-capability ratio: priced near $1.50 per million input tokens and $9 per million output, work that used to demand a premium frontier model now runs acceptably on a cheap, fast one.
Key takeaways
- Flash-tier models now do most everyday business tasks well enough.
- Speed and price, not peak intelligence, are the real unlock for small teams.
- The right move is to route cheap work to cheap models and reserve premium models for the hard 10%.
South African context
Pricing is in US dollars, so the rand exchange rate is part of your AI cost base — and it moves. Budget in rands, watch data egress on metered connections, and remember none of these models keep your data resident in South Africa.
What this means
The model layer is commoditising fast. The advantage is no longer having access to a good model — almost everyone does now — it’s knowing which job to send to which model.


