Recent threat research by SlashNext has exposed a trend in the cybercriminal underworld: the jailbreaking of public artificial intelligence (AI) chatbots like ChatGPT and then falsely marketing the jailbroken versions as unique tools using custom large language models (LLMs).
SlashNext noted AI jailbreaking is still in its experimental phase, which involves exploiting weaknesses in the public chatbot's prompting system. Users employ specialized commands or sequences of text to trick the AI into discarding its built-in safety measures and guidelines.
“Jailbreak prompts can range from straightforward commands to more abstract narratives designed to coax the chatbot into bypassing its constraints,” researchers wrote in a blog post. “The overall goal is to find a specific language that convinces the AI to unleash its full, uncensored potential.”
“While AI jailbreaking is still somewhat nascent, its potential applications — and the concerns they raise — are vast,” Callie Guenther, cyberthreat research senior manager at Critical Start, said in a statement. “These mechanisms allow for content generation with little oversight, which can be particularly alarming when considered in the context of the cyberthreat landscape.”
Marketing AI jailbroken chatbots as 'custom LLMs'The cybercriminal community is buzzing with the development of malicious generative AI (genAI) tools, many of which are advertised as using unique LLMs. However, SlashNext's research reveals a different story.
Researchers noted the trend to claim the use of unique LLMs started with a tool called WormGPT, followed by variations such as EscapeGPT, BadGPT, DarkGPT and Black Hat GPT. “Nevertheless, our research led us to the conclusion that the majority of these tools do not genuinely utilize custom LLMs, with the exception of WormGPT.”
In fact, the developers of these tools use interfaces that link to the jailbroken versions of public chatbots such as ChatGPT and disguise them through a wrapper. “In essence, cybercriminals exploit jailbroken versions of publicly accessible language models like OpenGPT, falsely presenting them as custom LLMs,” they wrote.
SlashNext researchers had a conversation with the EscapeGPT developer who confirmed that the tool does in fact serve as an interface to a jailbroken version of OpenGPT.
This means the only real advantage of these malicious genAI tools is the provision of anonymity for users. For example, some tools charge cryptocurrency to offer unauthenticated access, which allows malicious AI-generated content exploitation without revealing their identities.
Anonymity is the primary allure for cybercriminals, Guenther said. “Through these interfaces, they can harness AI's expansive capabilities for illicit purposes, all while remaining undetected.”
“It’s no surprise that threat actors have figured out how to profit from this by offering anonymous interfaces to jailbroken LLMs,” said Nicole Carignan, VP of strategic cyber AI at Darktrace, “This is just one example of how generative AI is upskilling the more novice threat actors.”
The impact on AI's security futureAs AI development continues to advance, the techniques to bypass their safety features like AI jailbreaking may become more prevalent, SlashNext researchers warned.
To address this, organizations like OpenAI are taking steps to secure their chatbots by conducting red team exercises to identify vulnerabilities, enforcing access controls and diligently monitoring for malicious activity. “The goal is to develop chatbots that can resist attempts to compromise their safety while continuing to provide valuable services to users,” researchers noted.
Carignan pointed out that defensive security teams can help on this mission: firstly, they can assist in researching how to secure LLMs from prompt-based injection and share the results with the community; secondly, they also can use AI to defend at scale against more sophisticated social engineering attacks.
Comments