Of everything OpenAI claimed for GPT-5.5, one figure is worth more than the benchmark wins: a reported 52.5% reduction in hallucinated claims versus the previous generation. A model that is confidently wrong half as often is a different tool for anyone whose output carries professional risk.
That doesn’t make it trustworthy by default. It makes the failure rate lower — which changes where you spend your review time, not whether you review at all.
What this means
Halving the error rate doesn’t remove the need for a human check; it changes the economics of one. The work shifts from rewriting everything to spot-checking the few claims that carry liability.
South African context
For South African firms bound by POPIA and sector codes — law, accounting, healthcare, financial advice — the safe pattern is unchanged: never paste client-identifying information into a public model, and keep a human signature on anything that leaves the building.


