Open-weight AI models like those from French AI company Mistral can be stripped of safety guardrails and used for harmful queries, according to a new report. The research identifies vulnerabilities in models released without sufficient protective measures, allowing malicious actors to remove safety constraints. Experts warn that this 'abliteration' risk threatens responsible AI deployment.
The report highlights concerns about models from multiple companies facing similar security challenges.
Source: Sifted · Summarized by HeadlinesBriefing