This article explains how to build a decision model using a large language model (LLM) by constraining outputs to a fixed set of options, such as A through E. Unlike standard language models that generate tokens sequentially, decision models like Jev assume fixed choices and make predictions in a single pass. The author demonstrates this technique using the Qwen/Qwen3-1.7B model, applying constrained decoding to limit outputs to predefined tokens. By masking all other vocabulary items, the model can only emit valid options, and the highest probability token is selected as the answer.
The approach was tested on the CommonsenseQA dataset, achieving 59.38% accuracy without fine-tuning. After a quick fine-tune, accuracy improved to 62.41%. The article also highlights a limitation: output token probabilities do not necessarily reflect true confidence in correctness, especially on ambiguous questions. The author provides a full Python implementation using Hugging Face Transformers and PyTorch, showing how to format prompts, load models, and compute probabilities for each option. This method offers a lightweight way to repurpose LLMs for multiple-choice reasoning tasks.
Source: Hacker News · Summarized by HeadlinesBriefing