HeadlinesBriefing HeadlinesBriefing.com

The Reversal Curse in Language Models Explained

Towards Data Science •
×

A 2023 paper by Berglund and colleagues argues that language models suffer from the Reversal Curse: knowing “A is B” doesn’t let them infer “B is A.” For example, a model trained on “Valentina Tereshkova was the first woman to travel to space” may answer “Who was Valentina Tereshkova?” but fail on “Who was the first woman to travel to space?”.

The authors fine-tuned GPT-3 and Llama-1 on invented facts and found near-zero accuracy when questions reversed the training order. Testing GPT-4 on real celebrities: it named a celebrity’s parent about 79% of the time, but named the celebrity from the parent only about 33% of the time (e.g., “Who is Tom Cruise’s mother?” vs. “Who is Mary Lee Pfeiffer’s son?”).

To explore this blindly spot simply, the article built a tiny model using only NumPy: single-layer embeddings without attention or hidden layers, trained on 200 invented facts pairing made-up names like “Zorvath Kellin is the Minister of Tides.” Half were taught in one direction, half in reverse, showing the curse emerges even in minimal architectures.

Source: Towards Data Science · Summarized by HeadlinesBriefing