HeadlinesBriefing favicon HeadlinesBriefing.com

চ্যাট টেমপ্লেট LLM স্ব-রেফারেন্সাল Voix নিয়ন্ত্রণ করে

Hacker News •
×

Large Language Models (LLMs) often add disclaimers like 'I'm just an AI' when asked about themselves. This self-referential voice is used in debates about AI safety and model self-knowledge, but its drivers are unclear. We show that the chat template acts as a switch: when present, it increases disclaimer voice and decreases experiential voice like 'I feel' across 8 popular open-source instruct models up to 9B parameters.

Without the template, disclaimer voice decreases and experiential voice increases. In 3 models, we identified a specific activation direction that steers this behavior—removing it reduces disclaimers, adding it increases them, while random directions have little effect. Instruct models without chat templates, when given this direction, produce disclaimers as if the template were present.

This reveals that model self-descriptions are not purely intrinsic but are partially shaped by deployment artifacts like chat templates. Researchers studying model introspection must control for this confound. Our work provides a steerable direction to isolate or manipulate this voice, demonstrating that what LLMs say about themselves reflects not just their weights, but also how they are prompted and formatted.