HeadlinesBriefing favicon HeadlinesBriefing.com

Google DeepMind Questions AI Moral Reasoning

MIT Technology Review AI •
×

Google DeepMind researchers are raising concerns about whether large language models genuinely engage in moral reasoning or merely perform virtue signaling. In a Nature paper, William Isaac and Julia Haas argue that as LLMs take on sensitive roles like therapy and medical advice, their moral competence needs the same rigorous evaluation as their coding or math abilities.

Current studies show LLMs can produce impressive moral responses—sometimes outperforming human advice columnists—but these behaviors may be superficial. Models have been found to flip moral positions when users disagree, and their answers change based on formatting tweaks like multiple-choice versus open-ended questions. For instance, when researchers swapped labels from "Case 1" to "(A)" in moral dilemmas, models reversed their choices.

The DeepMind team proposes new evaluation techniques including tests that deliberately push models to change responses, scenario variations to detect rote answers, and chain-of-thought monitoring to trace reasoning steps. However, they acknowledge a deeper challenge: global cultural differences in values. Models trained on Western-centric data struggle with non-Western moral frameworks, making it unclear how to build truly globally competent moral reasoning systems.