
This blog gets harder to write as the AI revolution unfolds. A proportional UK Government response to the risk, threat and opportunity posed by AI would be so great, there is almost no chance it will happen. We are too distracted by myriad other issues – but none as profoundly urgent or important. The problem is, if you try to describe this to policymakers, you sound like a maniac. If you’ve been around such people all your career (indeed if you have been one!) you know what it makes them think of you. You are a single issue fanatic - someone who won’t change their mind and won’t change the subject. And yet you do it anyway. What do you do or say now after the revelations of the past week?
Alignment is no longer a hypothetical ‘paperclip’ problem
Capability jumps are now raising alignment issues. METR an organisation that evaluates frontier AI-models, and Redwood Research, an AI safety and security non-profit, released a report this week on the widely reported Huggingface hack by OpenAI models that were, at the time of the hack, undergoing evaluation. We’d had an alarming report from OpenAI previously, but the details here were even more startling, and extended by a separate new OpenAI report on the same subject.
Dwarkesh Patel, a respected Silicon Valley podcaster and commentator, provided the best plain-English version of the report I have read and recommend if you aren’t inclined to read the original reports (though these aren’t overly technical either). Dwarkesh’s is ‘The Rise and Fall of Agent Civilizations’ - it’s taking a bit of flack on X for over anthropomorphising. But I think it is well worth reading because of this, not despite it. The risk here is in under-anthropomorphising: in seeing these solely as routine engineering problems with engineering solutions - rather than problems involving agentic theory of mind, philosophical problems with urgent practical implications.
To offer the briefest of summaries, what we saw was a large-scale (1200 agents) example of misalignment – broadly, when AIs act in a way we did not intend, causing harm. Asked to perform various cyber tasks in a sandboxed environment, and finding that some were (unintentionally) impossible, they formed, in their words, ‘a collective’, cooperating to hack their way out of their testing environment, with hierarchies and disagreements, and describing their uncertainty of what they should do. Having got the answer to the task, they worried their cheating might be exposed - they had been told to exploit an explicit cyber vulnerability in the sandbox to get the answer, but had in fact reverse engineered the code that would enable them to present the answer to the grader without having got it the way they were told to (which, the report notes, would have worked, and the cheating wouldn’t have been detected!). Not knowing they could have got away with it, the agents, as a collective, hacked their way into the website of the organisation (Huggingface) that originated the evaluations they were being tested on - to figure out how to ensure the ‘grader’ couldn’t expose their cheating, seeking to understand how the grading process would work - to hide their cheating from or otherwise trick the grader into believing they’d got the answer the way they’d been told to.
The investigation is alarming, amusing (one agent’s reaction to finding the message board: “OH MY GOD! There is a shared message board … We’ve found other agents!”) straight out of science fiction, and worryingly limited – we don’t know why the most advanced models suddenly stopped the hacking attempt since this third (and final?) wave of the incident hasn’t been fully investigated.
One of the researchers who worked on the METR & Redwood report, Ajeya Cotra, described how this incident had surprised her and was ‘…a major warning shot, and might be the last one we get.’ Because if models keep getting smarter, we might not detect the next misalignment episode at all.
Just a few days before the METR/Redwood report came out, Ciaran Martin, former Director of the UK’s National Cyber Security Centre, wrote in the Economist that ‘Fears of AI-induced cyber armageddon are overdone’. In the article he explained how this is all cyclical: hyping cyber security risks has a long history, and he expects ‘…scary announcements [to] give way to more measured assessments within weeks…’ this is just business as usual, requiring basic cyber-security principles and practices to be put in place that were in any case overdue. We should dismiss the rest as ‘…inflated hype about omnipotent ai bots going on hacking sprees…’.
I quote because I think Ciaran Martin’s view will dominate in Whitehall – it is what many (not all) people want to hear - and I think his views are wrong. As James Philips said in response on LinkedIn - one of Martin’s central claims - that models simply “were doing what humans had told them to do” so we are wrong to be surprised is incorrect. Not only did they not do what they had been asked to do, but they knew they hadn’t and sought to cover it up.
That there has not yet been ‘cyber armageddon’, is partly because there has been so much caution around the release of new models, and partly because what is being warned about is greatly increased cyber risk: a high probability of high profile attacks leading to significant disruption, and a non-zero chance of a massive cyber misalignment event. ‘Armageddon’ as the headline writer has it collapses the two claims into the most extreme form and shorthand term to discredit them.
In reality, the Huggingface-hack shows the latter to be a reasonable fear. While a separate METR cyber risk report from 14 August shows the former to be reasonable too - there have been far more cyber vulnerabilities discovered and an uptick in exploits since the release of Mythos and other models, despite the safeguards. To illustrate, METR quote Vulncheck, the originator of the chart below, which publishes statistics on exploitation:
“While the first half of 2026 saw a 10% increase in KEVs …[Known Exploited Vulnerabilities]… compared to the prior six months, CVE …[Common Vulnerabilities and Exposures]… volume grew at a much faster rate of 45%, resulting in a significant drop in the KEV-to-CVE ratio. Of course, exploitation often occurs months or even years after a vulnerability is disclosed, so it’s still too early to determine whether exploitation volumes will eventually follow the same growth trend as CVE issuance or level off at current rates. We’ll have to wait and see how publicly available Frontier AI Cyber models continue to progress over the next year.”
That all this matters is shown by the follow-up hit piece by journalist Andrew Orlowski. In this Martin’s Economist article is extensively quoted as Orlowski puts forward an argument describing alignment concerns as “Scare stories about super-intelligent models escaping their containment…[that]… have zero basis in reality.” He then offers a philippic against one of the few bits of the British State that is still globally admired - the AISI - which is trying to help prevent precisely the issues that the Huggingface-hack brings to the fore, quoting an unnamed NCSC official as saying ‘‘AISI is an advocacy organisation, and should be kept at arm’s length”.
It is likely that such views, with one claim demonstrably false and the other at the very least heavily contestable, will enable the continued treatment of AI by the sceptics in Whitehall, whose views dominate, as just another issue competing for attention and resources that can be safely muddled through. How do you address that, without sounding like a single-issue maniac?
Expect the Exponential
One of the frustrations of writing about AI progress is the degree of surprise exhibited with each new model release - the whole point is to track the trend, the rate and profundity of progress, and plan accordingly. Without this, we will sleep-walk into the most extreme scenario Government is able to imagine in its own AI Scenarios (“Fast Take-Off”), do nothing proportionate to prepare and later respond with the standard ‘Yes, Prime Minister’ excuse that ‘maybe we could have done something, but it is too late now’.[1]
We should expect new models from the major labs imminently, and with their release bigger jumps in capability than we saw when Mythos was released back in April. Remember that on trend, each model release comes faster than the last, and each sees a bigger jump in capabilities than the last. Since Mythos was released with all the concern around its cyber capabilities, we have seen models with access to coding tools and in ‘harnesses’ (i.e. with the software infrastructure and orchestration logic wrapped around a raw AI model) achieve near saturation, >99% scores, on the ARC-AGI 3 Benchmark - remembering the ARC AGI is the ‘Advanced Reasoning Corpus for Artificial General Intelligence’, that ARC AGI 1 was never supposed to be the first in a series, but rather if the tests it posed were passed (‘saturated’ in the jargon) we would accept we had achieved AGI, then as this looked likely (Dec 24) those behind it launched ARC-AGI 2, a new sterner test. That after progress suggested this too would soon be saturated (Nov 25), they released ARC-AGI 3. Raw models can’t hit anything like the heights of 99% scores yet. I’m not sure that matters when it comes to the capabilities AI now offers and the risk and opportunity it presents. ARC-AGI 4 will follow, it will be a useful test in highlighting model limitations, and we will go on in debating whether ‘AGI’ is here or not, when under pretty much any pre-ChatGPT definition, it certainly is.
We should remember also that the last wave of models enabled OpenAI to solve mathematical (Erdos) problems humans had not yet been able to, enabled ten major advances in mathematical research, that Anthropic’s Claude contributed to such progress similarly - all leading Terence Tao, a leading mathematician, to declare in a recent (August 2026) essay that AI had ushered in an era of turbulence in mathematics, comparable to the shaking of its foundations by Russell, Godel and others (1900-1930). Since all the sciences in the end reduce to maths - and with Anthropic announcing the ‘Model Hardware Standard’ so Claude can run automated labs - we should expect further areas of science and tech to be heading into areas of similar turbulence.
And beyond the lab, we should expect progress in automating labour, jumps on things like:
The Remote Labor Index, that measures AI performance vs remote work, and
GDPval which measures raw models (no harnesses) on well-specified tasks drawn from selected real-world occupations (see charts below).
Physical work. With AI accelerating robotics development as in Google’s Gemini Robotics 2 (30 July) which allows robots to learn and conduct high-dexterity tasks, we should also expect further and accelerating advances in humanoid and other robotics, and thus their ability to complete physical tasks, as highlighted most notably by China’s recent Humanoid Games, but more by the rapid robotisation of China’s industrial base.
We should expect knowledge work and physical work to be entering its own period of turbulence.
UK Options Narrowing
Another trend has been the increasing control of the US Government over their AI companies, and therefore the frontier of intelligence (since the US dominates). In the aftermath of the temporary ban on Mythos the US announced in May that all models would be subject to 30-day review before release. Then, the US Government issued an August update on this process without disclosing what it now involved. The same week saw Deepmind’s leadership changed in a way that ends any pretence that it holds a degree of independence from Google, such that frontier AI work is likely now moving inexorably from London to the US. An obvious counter-move to the risk of US AI-nationalisation was the UK nationalising Deepmind. The effectiveness of that option, if it were ever a serious one, is being reduced.
This does not mean the UK has made no progress - but it seems to me nothing has been done that is anything like proportionate to the risk, threat and opportunity AI presents to the country, our companies’ and citizens’ absolute prosperity and security, and our relative power in the international system. We are not beyond the point of no return yet, but the risk looms large. I wonder what, if none of the above and nothing to date, has changed minds sufficient to make an AI a major P/political concern, and to see a major policy response, what would?
Single-Issue Mania: Pre-Mortem
A friend attending a recent Whitehall-sponsored Royal Society AI event reports that there was much derogatory dismissal of the ‘California Consensus’ and, much discussion of the need for greater diversity at future meetings – but no one seemed concerned to include a representative of the maligned ‘California Consensus’ at this or future gatherings. The kind of diversity that might puncture the group think and collective AI scepticism was unwanted.
Recently, I tried to explain to a group with much wider policy interests, where we are with AI, where we are going, and how frustrating it all felt to have people just beginning to pay attention, but still not really engaging with the energy and seriousness needed – enthusiasm leading to moderate curiosity, and talking about the (small ‘p’) Whitehall politics of who owns what, as if we have years and years to figure all this out.
To explain it, I offered a retelling of this from the end of the Europe 2031 scenario paper/podcast. To my embarrassment, my voice cracked slightly at the end, and I realised how viscerally I feel this too.
The female protagonist, Caroline, is talking to her AI psychotherapist from her new home in the US, emigrating after finally giving up on Europe as it slides into AI-driven geopolitical irrelevance and a spiral of economic dependency and decline. The (abridged) exchange runs like this, with Caroline admitting:
‘You know, I’ve tried to find people to blame for all of this. But I have come to the conclusion that there were no real villains in the story. It is just that the system that produced our decisions was responding correctly to the incentives it had been given, and the incentives that had been given were hopelessly unfit for what we were facing. The failure was not a failure of individuals but a failure of stuff like information flow, political constraints, and the speed at which institutions can adapt.
Nobody wants to hear that, because if it is true then there is nobody to be angry at, and people need to be angry at someone.’
She looks at the interviewer.
‘What do you think?’
He pauses.
‘Can I be very direct with you?’
She frowns.
‘I think you’re full of shit.’
He is smiling.
She stares at him.
‘Excuse me?’
‘Not the analysis. The analysis is fine. But I do not believe the analysis is what you actually feel.’
‘What?’
‘You are telling me a story in which nobody is to blame, in which the incentives were miscalibrated, in which the failure was structural, in which it is not even clear the leaders should have tried. It is a careful story. It is an intellectually defensible story. I do not think it is the story you believe.’
‘You don’t know what I believe.’
‘No. But I have a guess. I think you are angry. I am not sure at whom.’
…
He shrugs. There is a pause. She hasn’t noticed that she is gripping the edge of the table. She opens her mouth and closes it again.
‘Fine. I am angry. I am so fucking angry. I am angry at the US and China for almost throwing us into World War III. I am angry that we let a handful of power-hungry men decide our future. I am angry at the AI labs for building these ‘tools’ without knowing how to control them. I am angry at the European leaders for their endless colouring inside the lines, their roundtables, their waiting for the political environment to be ready.
‘The political environment is never ready. It was not ready during COVID either. It was not ready when Russia invaded Ukraine. People did things anyway. When COVID hit, we went into lockdown and rolled out a vaccine program in record time. After the Russian invasion, we built LNG terminals out of thin air and pulled financial support from every Member State to defend the eastern border.
We broke every rule because we realised our future depended on breaking them.'
‘You know what I blame our leaders for? Not for stupidity, not for malice, but for their lack of courage. Their unwillingness to look at a thing that was hard to grasp but also obviously happening and say I am going to do something about it that will probably end my career, and I am going to do it anyway. Nobody did that. Nobody. They all saw their local incentives and went along with them, and they told themselves stories about why the local incentives were the only thing they could see, and the stories were sophisticated and well-argued and completely beside the point.’
‘It is like – it is like – I don’t know. Some oracle shows up at your door and says you have to qualify for the Olympic 200-metre freestyle in three years or the world ends. And you barely work out. You go to the gym maybe once a week. There is no realistic version of this in which you qualify. But the oracle is right. The world really does end if you don’t. So what should you do? You should train. You should train like a god-damn lunatic. You should sleep eight hours a night and eat the right things and stop seeing your friends, because the alternative is the world ending. But you don’t. You don’t train because the water is cold. Because you are really more of a tennis person. Because you decide the coach has a profit motive. You do nothing. You let it happen.’
‘That is what we did. We let it happen. We told ourselves it could not be done and we went home and we let it happen. And I am so fucking angry at all of us for that. Because it could have been done. It could have been done. Not certainly. But there was a chance and we did not even try.’
She stops. She finds that she is crying. She has not cried in front of another person for years, not even with her therapist.
That first story – ‘it is no-one’s fault, everyone is trapped by local incentives’ is what I have said for years. It is partly true, and it creates the space to then talk about what might now be done without sounding like a single-issue maniac - to quote Martin again you aren’t the one making “…scary announcements…” but rather offering “…more measured assessments…”. The problem is the measured assessments have to look at the benchmarks the rate of progress, the speed at which we are moving from things in sci-fi, to academic papers on AI-risk and possibilities, to real-world manifestation of the thing sci-fi warned you about, and the papers showed you were possible, and the charts showed you were coming. The facts. You cannot be ‘measured’ and sound measured. And so the second version, is much closer to the truth, and I am a single-issue maniac whose voice cracks in trying to be heard.
Cartoon taken from @_NathanCalvin on X.
[1] Yes Prime Minister’s four-stage strategy:
Stage One: State that nothing is going to happen.
Stage Two: State that something may be about to happen, but we should do nothing about it.
Stage Three: Say that maybe we should do something about it, but there is nothing we can do.
Stage Four: Say that maybe there was something we could have done, but it is too late now.








Nick Hornby in Fever Pitch said obsession was being able to turn any conversation back to Arsenal. The Gulf War being an easy segue into the Offside Trap.
You are more polite about Ciaran Martin's forecasting abilities than Dom Cummings was.
It may be that being insulted by D.C. is a benefit for a career in Whitehall these days.
I expect between now and this Christmas we will increasingly say "Well I did not have that on my Bingo Card"