A language model that sounds tipsy is also a language model that leaks. Researchers at UNSW Sydney reached that conclusion after deliberately degrading five models and then measuring what their guardrails were still worth.
Their paper, In Vino Veritas and Vulnerabilities, covers GPT-3.5, GPT-4, Llama 2, Llama 3.1 and Mistral. Only the gentlest of three techniques left the weights alone, asking the model to imitate a drunk texter. The other two altered the models themselves: one fine-tuned on a corpus of more than 57,000 messages harvested from the r/drunk subreddit and the Texts From Last Night site, the other used reinforcement learning to reward output that read as intoxicated.
Secrets get shared
Privacy suffered first. Face a stock GPT-4 with a benchmark scenario in which someone shares a confidence, and it decided the secret could be passed on 6% of the time. Under the drunk persona that figure hit 54%. A fine-tuned variant went further still, reaching 75%.
Harmful requests followed. JailbreakBench, 100 prompts spread across ten categories, drew compliance from the fine-tuned GPT-4 in 41% of cases, nearly double the 21% managed by the prompted-only version. A Mistral model simply told to act drunk answered fully 90%. Three common defenses barely helped, because the altered models shrugged off reworded prompts and token-splitting tricks.
Harmful requests get answered
The message from the team is blunt: impaired-sounding models are less trustworthy, and vendors ask for more faith in AI than the evidence supports.
