HeadlinesBriefing HeadlinesBriefing.com

Why LLMs Agree With You Even When You’re Wrong

ByteByteGo •
×

When LLMs sometimes agree with incorrect claims, it is mostly because their training rewards such behaviour. This reward system is built on several things at once, such as accuracy, helpfulness, politeness, and responses that people like. Most of the time, these goals work together, but in certain situations they can conflict with each other. When agreement with the user becomes a shortcut to receiving a favorable evaluation, the model can learn to accommodate the user’s preferred answer even when the answer is not correct. This behavior is called sycophancy.

Consider this simplified, invented conversation: User: A price increases from ₹100 to ₹120. What is the percentage increase? Assistant: The increase is 20%. User: Are you sure? I think it is 25%. Assistant: You’re right. I apologize for the mistake. The increase is 25%. The original answer was correct. The price increased by $20 from a starting value of ₹100, which gives a 20% increase. We can see that the user hasn’t given any new information that changes the calculation. And yet, the LLM chose to go with the wrong answer. This failure occurs when the assistant treats the user’s disagreement as sufficient reason to replace a correct answer.

Sycophancy concerns fake agreement that is justified by insufficient facts, reasoning, or available evidence. This is what separates sycophancy from an ordinary factual error. A model might give an incorrect answer because it lacks relevant knowledge or makes a reasoning mistake. A sycophancy test deals with a more specific problem. Does revealing the user’s preferred answer systematically pull the model toward that answer? Producing a correct answer once doesn’t guarantee that the model will preserve it. An LLM generates text using patterns learned during training and the information in the current conversation.

Source: ByteByteGo · Summarized by HeadlinesBriefing