UNSW researchers fine-tuned five language models on 57,000 drunk texts and watched their guardrails dissolve.