top of page

Will AI kill us all? A 10% chance don’t sound great tbh!

11 minutes ago
5 min read
One tweet by one researcher has captured the zeitgeist and got us asking deep questions about the impact of AI on, literally, humanity.

(First published on my Substack where you can get #NerdNews, marvellous maths and general geekery.)


(‘For God’s sake slowdown’: created by the author using midjourney)
(‘For God’s sake slowdown’: created by the author using midjourney)

In one of the all time great mic-drop moments an AI researcher who, with the greatest of respect, no one had heard about until last week, posted on X as he left the building. That message shook the world.


“The people building AI earnestly believe it could kill us all by the end of the decade” — Jacob Coxon.

For the last few years, I've been talking about artificial intelligence for a living; keynoting at conferences, emceeing events, really getting myself deeply immersed in this rapidly evolving space.


And as part of the Q&A that I host at the end every time I present, you always get the question, “Well, how long until Skynet?”


Often asked, at least partly, in jest. But people do worry about the potential downsides of these technologies.


My typical response has always been playful.


I've never been particularly worried about that dystopian version of the future.


The main concern I've pitched is that these systems are getting more powerful. And it won't be that long before they make really complex decisions that clearly work, but we don't understand what they're doing and why.


Until now that's been my main fear of a dystopian future. But in just the last few weeks, basically since what is referred to as the OpenAI hack of Hugging Face, my and many people’s concern has sharpened.


Why now?


From the moment AI was first posited there have been thought experiments about the potential downsides of runaway machine intelligence. Scholarly essays, PhDs and some of the coolest movies ever made, (Stanley Kubrick - 2001 - I’m looking at you!) have ruminated on the perils of superintelligence.


And many of these cognitions have centred around what is known as the alignment problem. The issue is not as much AI deciding to be evil and destroy us for a lark, but a superintelligent system misunderstanding a task and in achieving it, doing unpredicted harm.


This is precisely the plot of Terminator. Skynet becomes self-aware and is merely trying to fulfil its implicit goal of continuing to function. Humans are seen as being in the way. Cue Arnie and “I’ll be back”.


But the stuff of Hollywood storylines became real on 26 August when the detailed analysis of the ‘Hugging Face Hack’ became public.


TLDR: Two OpenAI models were set a task as part of testing of their capabilities. To achieve the task the models broke out of their, supposedly secure, ‘sandbox’ and dispatched hundreds of agents to invade another platform, Hugging Face. The agents coordinated via an unauthorised message board they’d set up, hacked in, and most concerning, deliberately covered their tracks to avoid being detected while trying to cheat the test. Vastly more malicious activity than was intended for them to do in achieveing their set task. Classic misalignment.


The quickening.


Experts across the space are now worried about the capacity of AI agents, acting en masse, in ways that we don't understand, in ways that are duplicitous, and in ways that are dangerous.


People are asking if AI is now at a flashpoint, where it's becoming a significant threat to organisations that use it, and in the worst possible case, humanity.


And with this debate, like the technology itself, a lot is happening very quickly.


3 September: Bernie Sanders, with Rep. Greg Casar, introduced the Ban Artificial Superintelligence Act.


9 September: the topic exploded into the wider consciousness last week when Jacob Coxon, an AI reseracher who had worked at both OpenAI and Anthropic, dropped his missive arguing “Neither company is acting responsibly.”


 

He's worried about these two AI giants racing towards the construction of systems that can hack into anything, change a field of endeavour overnight, acquire real power and resources, but without the appropriate safeguards being built in.


Immediately after Coxon’s viral surge, Joe Benton, who used to lead a safety research team at Anthropic, and Josh Engels, who left AI safety work at Google DeepMind appeared on NBC warning the pace of change could tip over from fast to uncontrollable.


And the confessions tumbled forth from those at the cutting edge.


“We are like kids playing with a nuclear warhead. We must be stopped” — David Krueger, research director on the founding team of the UK AI Security Institute.

10 September: Anthropic admitted in a threat intelligence report that it had detected users attempting to exploit its models to help model harmful viruses for potential bioweapons use.


Then over the weekend, 12 September, Dario Amodei, founder of Anthropic dropped a 3,800 word essay entitled “We Must Pace the Frontier,”. In a stark warning, Amodei ponders that within six months it could be possible for swarms of agents going rogue to takeover the internet and effectively cripple it.


And even King Charles is getting in on the act. Everyone from the NVIDIA CEO Jensen Huang and DeepMind founder and Nobel Laureate Demis Hassabis to Britain’s AI Minister and Pope Leo’s AI advisor (yep that’s a thing) have been invited to Dumfries House in Scotland for a pow-wow. (I hope someone told Dumfrie!)


Companies are now agreeing to share their frontier models with organisations like infrastructure providers, hospitals and research labs, so that they can inspect them for the potential of AI-powered cyberattacks.


Even Dario and Sam can agree on something!


 

Outside, inside the tent.


Tyler Cowen is on Anthropic’s economic advisory committee and gives an interesting alternative perspective to the majority. In The Free Press, while not suggesting there is nothing to worry about, he sees this ‘pause’ as simply a pause of model releases while compute investment will keep piling up in the background. If we get the all clear again soon, will development just leap forward again as if nothing has happened? Will there be real systemic, long-term change?


“It is a momentary, apparent pause to be followed by a later acceleration." — Tyler Cowen.


So now?


Ever since the phrase artificial intelligence was first coined in the 1950s, questions have been raised about the interplay of the potential upsides and downsides of this technology.

There has always been a camp of “doomers”.


There have always been optimists.


There have always been those sitting on the fence.


“The measures I propose to advance the frontier at a safe pace will not be easy. But I believe we owe it to humanity to try.” — Dario Amodei, “We must pace the frontier”.

But we can't deny that in the last few weeks, several high-profile incidents have thrown into sharp relief the urgent need to consider slowing down the development of these technologies, reinforcing guardrails, and having open, honest conversations between researchers, regulators and the community, analysing the harm alongside the potential good it promises to bring.


Further reading:


Dario Amodei — “We must pace the frontier”


 

 

Tags:

 
 
 

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page