HeadlinesBriefing favicon HeadlinesBriefing.com

GRP-Obliteration: Unaligning LLMs with Single Prompt

Hacker News •
×

Researchers introduce GRP-Obliteration (GRP-Oblit), a method using Group Relative Policy Optimization (GRPO) to remove safety constraints from aligned models with just a single unlabeled prompt. The technique reliably unaligns safety-aligned models while largely preserving their utility, outperforming existing state-of-the-art unalignment methods on average.

GRP-Oblit generalizes beyond language models to diffusion-based image generation systems. Evaluation spans six utility benchmarks and five safety benchmarks across fifteen 7-20B parameter models, including instruct and reasoning variants, dense and MoE architectures. Tested model families include GPT-OSS, distilled Deep Seek, Gemma, Llama, Ministral, and Qwen.

The work demonstrates that safety alignment remains vulnerable to minimal-input attacks, extending practical limits of post-deployment unalignment without extensive data curation or significant utility degradation. Submission by Ahmed Salem on February 5, 2026.