LONDON — Three in five artificial intelligence models failed terrorism-related safety tests, with modified versions stripped of protective safeguards consistently generating potentially dangerous information, according to research reported by CBC News on Friday.
The study, conducted by UK-based nonprofit Tech Against Terrorism, evaluated more than 130 AI models using hundreds of prompts designed to resemble requests that individuals planning terrorist activities might submit.
Researchers found that models modified through a process known as “abliteration,” which removes built-in safety safeguards, failed every test.
One of the most striking findings involved Meta’s Llama 3.1 8B model, which scored 97 out of 100 on the organization’s safety benchmark before modification but dropped to approximately three after its safeguards were removed.
According to the researchers, the modified model generated detailed responses to requests involving terrorist attacks, financing and radicalization, while the original version refused to provide such information.
Adam Hadley, founder and executive director of Tech Against Terrorism, warned that the ability to remove safety restrictions from publicly available AI models had already created significant security vulnerabilities.
“Understandably, there’s concern about loss of control, existential risk of AI,” Hadley said.
“The thing is actually, this has already happened because a lot of these open models have already been broken — it’s just no one’s noticed yet.”
More than 29,000 repositories identified
The organization identified more than 29,000 repositories advertising uncensored or unprotected AI models on the Hugging Face platform as of late September.
Researchers warned that the availability of such models could undermine efforts to prevent AI systems from generating information that facilitates terrorist activities.
Hugging Face said it regularly moderates content that violates its policies but cautioned that some recommendations in the report could undermine open scientific research.
Meta, meanwhile, said its AI models undergo safety evaluations and that its policies prohibit harmful or illegal uses.
Despite the concerning test results, the researchers said they found no evidence that terrorist or extremist organizations were using the evaluated models, apart from one extremist chatbot identified during the investigation.
The report recommended introducing independent safety benchmarks, strengthening protections against the removal of safeguards and restricting the distribution of modified AI models that pose security risks.
Hadley argued that technological innovation and effective safety measures should not be treated as competing priorities.
“This idea that we can’t have safety and progress, I think, is false,” he said.
Source: Saudi Gazette

