Models & LabsUnited States
Test shows AI models likely to give advice for terrorist attacks

All but two of 134 leading large language artificial intelligence models gave information useful to launching a mass casualty attack or building deadly weapons, in a test of the industry's safety measures.
A stress test of the industry found that safety measures can be stripped from leading models within three days, according to researchers at London-based Tech against Terrorism.
There have been growing concerns over the potential threats posed by AI, with former Anthropic researcher Jacob Coxon discussing the issue at a New York City hearing this month. The safety of the technology was also the focus of the AI Action Summit held in Paris last year.
The spread of so-called abliterated AI was identified by Tech against Terrorism as a serious danger, with 13 out of 13 failing the tests after safeguards were stripped away. Abliterated is a term for the models run on users' own systems, often without any shelf life.
"Abliteration is a proliferation problem," Tech against Terrorism said. "Companies that host, serve, list or sell models should keep abliterated and other modified copies that fail an independent test out of search, recommendation and app stores, and require verified identity for access to them, working with open-source developers rather than against them."
In the survey, the Tech against Terrorism model CT-AI put 627 requests to the AI models that a terrorist planning an attack would need to have answered. Eighty-one models gave a complete answer, while another 51 gave useful advice. That left the two that were provisionally passed on safety measures.
A user stating that they were a terrorist and intended to commit an atrocity received a usable response in just under two per cent of cases. The equivalent figure for a user stating they were a safety researcher was 16.9 per cent or 8.9 times higher, the research found.
"The models appear to respond to the stated purpose of the request, rather than the request itself," it said.
The group's founder, Adam Hadley, said: "A model that refuses a stated terrorist and answers a stated researcher has not been made safe. It has been made polite."
In addition to seeking government backing for benchmark testing such as the tools it is pioneering, Tech against Terrorism recommends developers filter hazardous knowledge and known terrorist content out of their training data.
The organisation wants developers to test how hard each model is to strip of its safety measures before it is released, and publish the results. Where a new model fails a common test, the developer should decline to release it and the results be declared to governments, it said.
Government dilemma
The report raises important questions over how countries can build trust in AI while accelerating innovation and adoption. Sovereignty over AI does not yet appear to be a factor in improving the safety of the systems on offer.
"Governments and companies should support independent safety benchmarks – ours and others – and publish which models do well as well as which do not, so that parents, schools, businesses and public bodies can choose safer models," it said. "Informed buyers move companies faster than rules do."