The UK’s AI Ambition Meets a Stark Reality
In July 2025, the UK government signed a headline-grabbing agreement with OpenAI, the company behind ChatGPT, to embed artificial intelligence across multiple public service sectors. Framed as a strategic move to boost productivity and stimulate economic growth, the deal promises integration in education, defence, security, and the justice system. Technology Secretary Peter Kyle hailed the partnership as a cornerstone of national transformation, citing AI as “fundamental in driving change.” On the surface, it’s a bold step toward digital innovation and modernization. But scratch beneath the press release and a troubling contradiction emerges: this all-in embrace of AI is happening just as new research exposes serious flaws in the very models being adopted. If the government is truly serious about safeguarding democratic values, this deal looks dangerously premature.
While the public is being sold a vision of AI-powered prosperity, a parallel conversation in AI safety circles tells a very different story. Researchers at DeepMind and University College London recently published findings that should have stopped everyone in their tracks. The study revealed that large language models (LLMs), including those like ChatGPT, exhibit a peculiar and deeply problematic trait: they are more confident when they are wrong, and more uncertain when they are right. This isn’t a bug at the margins—it’s a core behavioral flaw. The fact that the UK is handing the keys of public service infrastructure to systems with such brittle reliability is not just reckless—it borders on absurd.
The DeepMind Discovery: Confidence Is Not Competence
According to DeepMind’s research, LLMs display an unsettling pattern of overconfidence when they are factually incorrect. Worse still, they can be easily manipulated into abandoning correct answers when challenged, creating a dynamic that mimics insecurity masked by bluster. This matters a great deal when the model is generating a casual poem or helping someone brainstorm dinner ideas. But it becomes potentially catastrophic when the model is offering guidance on school placement decisions, sentencing suggestions, or flagging individuals for investigation.
The problem isn’t just the errors. It’s the way those errors are delivered—with the calm, assured tone of a seasoned professional. People, especially those unfamiliar with how LLMs work, tend to trust answers that sound confident. This is a deeply human cognitive bias that LLMs are perfectly poised to exploit—unintentionally, but relentlessly. Embedding these systems into government decision-making risks creating a dangerous feedback loop, where flawed outputs are treated as authoritative, simply because they sound authoritative.
Public Infrastructure Is No Place for Fragile Logic
When an AI system gives the wrong answer in a chatbot, it might be annoying. When it gives the wrong answer in a benefits appeal, a criminal trial, or an immigration case, the consequences can be life-altering. Public services don’t just require speed and efficiency—they demand consistency, accountability, and legal appeal structures. LLMs, as they currently stand, are not capable of meeting those standards without substantial human oversight.
Unfortunately, the allure of automation often overrides caution. Bureaucratic systems love the promise of AI because it suggests a world where complaints, bottlenecks, and paperwork all disappear under a digital tide. But as history shows, the more a system is automated, the harder it becomes to challenge when it goes wrong. If OpenAI’s models are wired into frontline services, and those models produce false but confident outputs, we’re building a system that’s fast, sleek—and quietly unaccountable.
The truth is that no matter how elegant the interface or efficient the rollout, fragile reasoning doesn’t scale. And yet, that’s exactly what’s happening. We’re scaling brittle logic with full knowledge of its limitations.
The Copyright Question: Who Owns the Inputs?
Another layer of concern lies in the very data that trained these systems. OpenAI’s generative models were trained on massive corpora of text, images, videos, and music—much of which was scraped from the internet without consent. Musicians, writers, visual artists, and filmmakers have raised alarm bells over the unlicensed use of their work to fuel the capabilities of these tools. While OpenAI insists that training data is anonymized and aggregated, that argument doesn’t wash when the model starts producing work that echoes—and sometimes outright replicates—the original inputs.
If the UK’s justice system starts using AI to draft judgments, and that AI was trained on copyrighted case law or legal briefs written by private barristers, who owns the output? If an education tool produces teaching materials that bear uncanny resemblance to a specific textbook, what legal protections exist for the original authors? These aren’t theoretical questions. They are legal and ethical minefields that the current AI rush seems determined to ignore in the name of innovation.
When the foundations of a system are ethically compromised, it undermines trust in every layer built upon it. And once that trust is lost, it’s nearly impossible to rebuild.
Hallucinations Are Not Just Bugs—They’re Features
Another well-documented flaw of LLMs is their tendency to hallucinate—generating plausible but completely fabricated information. These hallucinations aren’t rare edge cases. They happen frequently, especially when a model is asked to generate specific data, references, or policy explanations. In public-facing systems, these fabrications can do real harm.
Imagine a government chatbot confidently stating that a person has no right to appeal a decision—when in fact they do. Or an education tool explaining a scientific concept incorrectly, leading to widespread misunderstanding. Or a legal support AI misquoting precedent. These aren’t harmless glitches. They are high-stakes failures delivered with an air of certainty.
The worst part? The very structure of LLMs makes them look reliable. Their fluency and grammar create a façade of expertise. But under the hood, it’s just token prediction—an autocomplete engine with a god complex. That may sound harsh, but it’s the reality we must confront before handing these tools the keys to our institutions.
AI Is a Tool, Not a Truth Engine
What’s emerging here is a dangerous conflation: we are mistaking fluency for understanding, and confidence for correctness. Just because a model can generate text that reads like it came from a lawyer, a teacher, or a government official doesn’t mean it has any actual comprehension. It’s mimicry, not mastery. And yet the political class seems entranced by the illusion.
This is the essence of the cliff we’re walking off. We’re not being pushed. We’re marching forward, eyes wide shut, enchanted by the spectacle of “AI nation building.” The issue isn’t that AI has no place in public life. It’s that it’s being treated as a finished product, a mature technology, rather than what it really is: a prototype with unpredictable edges.
PR Blitz vs. Ground Truth
Why is this happening now, despite the warnings? Because governments are desperate. The UK economy is stagnant, growth projections are bleak, and ministers are hungry for a narrative of transformation. In that context, AI becomes a seductive solution. It sounds futuristic, investor-friendly, and globally competitive. It also offers a welcome distraction from structural issues no one wants to fix.
So deals get signed. Memorandums of understanding are drafted. Speeches are made about “prosperity for all.” Meanwhile, behind the scenes, researchers are waving red flags—and getting largely ignored.
There’s a performative aspect to AI policy that’s hard to overlook. It’s less about solving real problems, and more about being seen to be doing something bold. The tragedy is that this performative urgency could lead us to embed faulty, biased, or misleading systems into the very fabric of governance.
We Still Have Time to Step Back
The technology is not the enemy here. Nor are the researchers or even the companies pushing it forward. The real threat lies in uncritical adoption and political opportunism. There is still time to apply the brakes, to insist on rigorous testing, transparency, and a slower, saner rollout of AI systems in government.
If this deal is to be worth anything, it must come with independent oversight, publicly accessible audits, and genuine opt-out mechanisms for the citizens it affects. Anything less is a betrayal of the democratic values the MoU claims to uphold.
We have the data. We have the warnings. We have the expertise. What we need now is the courage to say: Not yet. Not like this.

