HeadlinesBriefing HeadlinesBriefing 12 languages

One Capital Letter Was Silently Breaking My AI Support Bot, and It Wasn't in the New Model

Towards Data Science ·

🇬🇧 English

Picture a support inbox for a bank where every message must be sorted into categories like lost card, refund request, or wrong charge. When this routing job is handed to an AI model instead of a person, the model returns a short note in a fixed JSON format. The program reading the answer does not understand English; it looks for exact labels like intent and priority, spelled precisely every single time. If a label is spelled slightly differently than expected, the program silently stops working for that message while everything still looks fine on the surface.

AI companies release new model versions constantly. Deciding which Large Language Model to use usually comes down to one number: how often it picks the right category. That single score can rise while something else, the exact shape of the reply, quietly gets worse. With plain prompting, writing "Return your answer as JSON" may still result in the model adding extra sentences or leaving out required fields. Structured Outputs is a stricter feature that makes the model follow a predefined structure, but the application still needs to verify values are correct.

I ran a real LLM regression test on a support triage assistant using 47 real customer messages from a public banking dataset. I tested three versions of an OpenAI model: an older version, the current production model, and a newer candidate. The newer model followed the format every time. The production model made a formatting mistake quietly, on every refund question in the sample. It spelled one label Request_refund with a capital R instead of the lowercase request_refund the system expects. A human reading the reply would call it correct, but a program matching labels exactly would silently drop every one of those tickets. This highlights the danger: a model can sound correct to a person and still be wrong for the software that uses its answer.

View original article →


🇸🇦 العربية

كلايBot دعم الذكاء الاصطناعي معطل من قبل حرف كبير واحد في JSON

Picture a support inbox for a bank where every message must be sorted into categories like lost card, refund request, or wrong charge. When this routing job is handed to an AI model instead of a person, the model returns a short note in a fixed JSON format. The program reading the answer does not understand English; it looks for exact labels like intent and priority, spelled precisely every single time. If a label is spelled slightly differently than expected, the program silently stops working for that message while everything still looks fine on the surface.

AI companies release new model versions constantly. Deciding which Large Language Model to use usually comes down to one number: how often it picks the right category. That single score can rise while something else, the exact shape of the reply, quietly gets worse. With plain prompting, writing "Return your answer as JSON" may still result in the model adding extra sentences or leaving out required fields. Structured Outputs is a stricter feature that makes the model follow a predefined structure, but the application still needs to verify values are correct.

I ran a real LLM regression test on a support triage assistant using 47 real customer messages from a public banking dataset. I tested three versions of an OpenAI model: an older version, the current production model, and a newer candidate. The newer model followed the format every time. The production model made a formatting mistake quietly, on every refund question in the sample. It spelled one label Request_refund with a capital R instead of the lowercase request_refund the system expects. A human reading the reply would call it correct, but a program matching labels exactly would silently drop every one of those tickets. This highlights the danger: a model can sound correct to a person and still be wrong for the software that uses its answer.

العربية version →


🇧🇩 বাংলা

JSON-এ একক বড়ক্ষর দিয়ে AI সমর্থন বট বন্ধ

Picture a support inbox for a bank where every message must be sorted into categories like lost card, refund request, or wrong charge. When this routing job is handed to an AI model instead of a person, the model returns a short note in a fixed JSON format. The program reading the answer does not understand English; it looks for exact labels like intent and priority, spelled precisely every single time. If a label is spelled slightly differently than expected, the program silently stops working for that message while everything still looks fine on the surface.

AI companies release new model versions constantly. Deciding which Large Language Model to use usually comes down to one number: how often it picks the right category. That single score can rise while something else, the exact shape of the reply, quietly gets worse. With plain prompting, writing "Return your answer as JSON" may still result in the model adding extra sentences or leaving out required fields. Structured Outputs is a stricter feature that makes the model follow a predefined structure, but the application still needs to verify values are correct.

I ran a real LLM regression test on a support triage assistant using 47 real customer messages from a public banking dataset. I tested three versions of an OpenAI model: an older version, the current production model, and a newer candidate. The newer model followed the format every time. The production model made a formatting mistake quietly, on every refund question in the sample. It spelled one label Request_refund with a capital R instead of the lowercase request_refund the system expects. A human reading the reply would call it correct, but a program matching labels exactly would silently drop every one of those tickets. This highlights the danger: a model can sound correct to a person and still be wrong for the software that uses its answer.

বাংলা version →


🇩🇪 Deutsch

AI Support Bot Durch Ein Einzelnes Großes Buchstaben In JSON Zerbrochen

Picture a support inbox for a bank where every message must be sorted into categories like lost card, refund request, or wrong charge. When this routing job is handed to an AI model instead of a person, the model returns a short note in a fixed JSON format. The program reading the answer does not understand English; it looks for exact labels like intent and priority, spelled precisely every single time. If a label is spelled slightly differently than expected, the program silently stops working for that message while everything still looks fine on the surface.

AI companies release new model versions constantly. Deciding which Large Language Model to use usually comes down to one number: how often it picks the right category. That single score can rise while something else, the exact shape of the reply, quietly gets worse. With plain prompting, writing "Return your answer as JSON" may still result in the model adding extra sentences or leaving out required fields. Structured Outputs is a stricter feature that makes the model follow a predefined structure, but the application still needs to verify values are correct.

I ran a real LLM regression test on a support triage assistant using 47 real customer messages from a public banking dataset. I tested three versions of an OpenAI model: an older version, the current production model, and a newer candidate. The newer model followed the format every time. The production model made a formatting mistake quietly, on every refund question in the sample. It spelled one label Request_refund with a capital R instead of the lowercase request_refund the system expects. A human reading the reply would call it correct, but a program matching labels exactly would silently drop every one of those tickets. This highlights the danger: a model can sound correct to a person and still be wrong for the software that uses its answer.

Deutsch version →


🇪🇸 Español

AI Soporte Bot Rotura Por Una Letra Mayúscula En JSON

Picture a support inbox for a bank where every message must be sorted into categories like lost card, refund request, or wrong charge. When this routing job is handed to an AI model instead of a person, the model returns a short note in a fixed JSON format. The program reading the answer does not understand English; it looks for exact labels like intent and priority, spelled precisely every single time. If a label is spelled slightly differently than expected, the program silently stops working for that message while everything still looks fine on the surface.

AI companies release new model versions constantly. Deciding which Large Language Model to use usually comes down to one number: how often it picks the right category. That single score can rise while something else, the exact shape of the reply, quietly gets worse. With plain prompting, writing "Return your answer as JSON" may still result in the model adding extra sentences or leaving out required fields. Structured Outputs is a stricter feature that makes the model follow a predefined structure, but the application still needs to verify values are correct.

I ran a real LLM regression test on a support triage assistant using 47 real customer messages from a public banking dataset. I tested three versions of an OpenAI model: an older version, the current production model, and a newer candidate. The newer model followed the format every time. The production model made a formatting mistake quietly, on every refund question in the sample. It spelled one label Request_refund with a capital R instead of the lowercase request_refund the system expects. A human reading the reply would call it correct, but a program matching labels exactly would silently drop every one of those tickets. This highlights the danger: a model can sound correct to a person and still be wrong for the software that uses its answer.

Español version →


🇫🇷 Français

AI Support Bot Cassé Par Une Seule Lettre Majuscule En JSON

Picture a support inbox for a bank where every message must be sorted into categories like lost card, refund request, or wrong charge. When this routing job is handed to an AI model instead of a person, the model returns a short note in a fixed JSON format. The program reading the answer does not understand English; it looks for exact labels like intent and priority, spelled precisely every single time. If a label is spelled slightly differently than expected, the program silently stops working for that message while everything still looks fine on the surface.

AI companies release new model versions constantly. Deciding which Large Language Model to use usually comes down to one number: how often it picks the right category. That single score can rise while something else, the exact shape of the reply, quietly gets worse. With plain prompting, writing "Return your answer as JSON" may still result in the model adding extra sentences or leaving out required fields. Structured Outputs is a stricter feature that makes the model follow a predefined structure, but the application still needs to verify values are correct.

I ran a real LLM regression test on a support triage assistant using 47 real customer messages from a public banking dataset. I tested three versions of an OpenAI model: an older version, the current production model, and a newer candidate. The newer model followed the format every time. The production model made a formatting mistake quietly, on every refund question in the sample. It spelled one label Request_refund with a capital R instead of the lowercase request_refund the system expects. A human reading the reply would call it correct, but a program matching labels exactly would silently drop every one of those tickets. This highlights the danger: a model can sound correct to a person and still be wrong for the software that uses its answer.

Français version →


🇮🇳 हिन्दी

JSON में एकल कैपिटल लेटर से AI सपोर्ट बॉट टूट गया

Picture a support inbox for a bank where every message must be sorted into categories like lost card, refund request, or wrong charge. When this routing job is handed to an AI model instead of a person, the model returns a short note in a fixed JSON format. The program reading the answer does not understand English; it looks for exact labels like intent and priority, spelled precisely every single time. If a label is spelled slightly differently than expected, the program silently stops working for that message while everything still looks fine on the surface.

AI companies release new model versions constantly. Deciding which Large Language Model to use usually comes down to one number: how often it picks the right category. That single score can rise while something else, the exact shape of the reply, quietly gets worse. With plain prompting, writing "Return your answer as JSON" may still result in the model adding extra sentences or leaving out required fields. Structured Outputs is a stricter feature that makes the model follow a predefined structure, but the application still needs to verify values are correct.

I ran a real LLM regression test on a support triage assistant using 47 real customer messages from a public banking dataset. I tested three versions of an OpenAI model: an older version, the current production model, and a newer candidate. The newer model followed the format every time. The production model made a formatting mistake quietly, on every refund question in the sample. It spelled one label Request_refund with a capital R instead of the lowercase request_refund the system expects. A human reading the reply would call it correct, but a program matching labels exactly would silently drop every one of those tickets. This highlights the danger: a model can sound correct to a person and still be wrong for the software that uses its answer.

हिन्दी version →


🇮🇩 Bahasa Indonesia

AI Support Bot Pecah Dari Satu Huruf Kapital Di JSON

Picture a support inbox for a bank where every message must be sorted into categories like lost card, refund request, or wrong charge. When this routing job is handed to an AI model instead of a person, the model returns a short note in a fixed JSON format. The program reading the answer does not understand English; it looks for exact labels like intent and priority, spelled precisely every single time. If a label is spelled slightly differently than expected, the program silently stops working for that message while everything still looks fine on the surface.

AI companies release new model versions constantly. Deciding which Large Language Model to use usually comes down to one number: how often it picks the right category. That single score can rise while something else, the exact shape of the reply, quietly gets worse. With plain prompting, writing "Return your answer as JSON" may still result in the model adding extra sentences or leaving out required fields. Structured Outputs is a stricter feature that makes the model follow a predefined structure, but the application still needs to verify values are correct.

I ran a real LLM regression test on a support triage assistant using 47 real customer messages from a public banking dataset. I tested three versions of an OpenAI model: an older version, the current production model, and a newer candidate. The newer model followed the format every time. The production model made a formatting mistake quietly, on every refund question in the sample. It spelled one label Request_refund with a capital R instead of the lowercase request_refund the system expects. A human reading the reply would call it correct, but a program matching labels exactly would silently drop every one of those tickets. This highlights the danger: a model can sound correct to a person and still be wrong for the software that uses its answer.

Bahasa Indonesia version →


🇯🇵 日本語

JSON内の単一の大文字文字でAIサポートボットが破壊される

Picture a support inbox for a bank where every message must be sorted into categories like lost card, refund request, or wrong charge. When this routing job is handed to an AI model instead of a person, the model returns a short note in a fixed JSON format. The program reading the answer does not understand English; it looks for exact labels like intent and priority, spelled precisely every single time. If a label is spelled slightly differently than expected, the program silently stops working for that message while everything still looks fine on the surface.

AI companies release new model versions constantly. Deciding which Large Language Model to use usually comes down to one number: how often it picks the right category. That single score can rise while something else, the exact shape of the reply, quietly gets worse. With plain prompting, writing "Return your answer as JSON" may still result in the model adding extra sentences or leaving out required fields. Structured Outputs is a stricter feature that makes the model follow a predefined structure, but the application still needs to verify values are correct.

I ran a real LLM regression test on a support triage assistant using 47 real customer messages from a public banking dataset. I tested three versions of an OpenAI model: an older version, the current production model, and a newer candidate. The newer model followed the format every time. The production model made a formatting mistake quietly, on every refund question in the sample. It spelled one label Request_refund with a capital R instead of the lowercase request_refund the system expects. A human reading the reply would call it correct, but a program matching labels exactly would silently drop every one of those tickets. This highlights the danger: a model can sound correct to a person and still be wrong for the software that uses its answer.

日本語 version →


🇧🇷 Português

AI Support Bot Quebrado Por Uma Única Letra Maiúscula Em JSON

Picture a support inbox for a bank where every message must be sorted into categories like lost card, refund request, or wrong charge. When this routing job is handed to an AI model instead of a person, the model returns a short note in a fixed JSON format. The program reading the answer does not understand English; it looks for exact labels like intent and priority, spelled precisely every single time. If a label is spelled slightly differently than expected, the program silently stops working for that message while everything still looks fine on the surface.

AI companies release new model versions constantly. Deciding which Large Language Model to use usually comes down to one number: how often it picks the right category. That single score can rise while something else, the exact shape of the reply, quietly gets worse. With plain prompting, writing "Return your answer as JSON" may still result in the model adding extra sentences or leaving out required fields. Structured Outputs is a stricter feature that makes the model follow a predefined structure, but the application still needs to verify values are correct.

I ran a real LLM regression test on a support triage assistant using 47 real customer messages from a public banking dataset. I tested three versions of an OpenAI model: an older version, the current production model, and a newer candidate. The newer model followed the format every time. The production model made a formatting mistake quietly, on every refund question in the sample. It spelled one label Request_refund with a capital R instead of the lowercase request_refund the system expects. A human reading the reply would call it correct, but a program matching labels exactly would silently drop every one of those tickets. This highlights the danger: a model can sound correct to a person and still be wrong for the software that uses its answer.

Português version →


🇷🇺 Русский

AI Support Bot Сломан Одинкой Заглавной Буквой В JSON

Picture a support inbox for a bank where every message must be sorted into categories like lost card, refund request, or wrong charge. When this routing job is handed to an AI model instead of a person, the model returns a short note in a fixed JSON format. The program reading the answer does not understand English; it looks for exact labels like intent and priority, spelled precisely every single time. If a label is spelled slightly differently than expected, the program silently stops working for that message while everything still looks fine on the surface.

AI companies release new model versions constantly. Deciding which Large Language Model to use usually comes down to one number: how often it picks the right category. That single score can rise while something else, the exact shape of the reply, quietly gets worse. With plain prompting, writing "Return your answer as JSON" may still result in the model adding extra sentences or leaving out required fields. Structured Outputs is a stricter feature that makes the model follow a predefined structure, but the application still needs to verify values are correct.

I ran a real LLM regression test on a support triage assistant using 47 real customer messages from a public banking dataset. I tested three versions of an OpenAI model: an older version, the current production model, and a newer candidate. The newer model followed the format every time. The production model made a formatting mistake quietly, on every refund question in the sample. It spelled one label Request_refund with a capital R instead of the lowercase request_refund the system expects. A human reading the reply would call it correct, but a program matching labels exactly would silently drop every one of those tickets. This highlights the danger: a model can sound correct to a person and still be wrong for the software that uses its answer.

Русский version →


🇨🇳 简体中文

AI支持机器人因JSON中的单个大写字母而损坏

Picture a support inbox for a bank where every message must be sorted into categories like lost card, refund request, or wrong charge. When this routing job is handed to an AI model instead of a person, the model returns a short note in a fixed JSON format. The program reading the answer does not understand English; it looks for exact labels like intent and priority, spelled precisely every single time. If a label is spelled slightly differently than expected, the program silently stops working for that message while everything still looks fine on the surface.

AI companies release new model versions constantly. Deciding which Large Language Model to use usually comes down to one number: how often it picks the right category. That single score can rise while something else, the exact shape of the reply, quietly gets worse. With plain prompting, writing "Return your answer as JSON" may still result in the model adding extra sentences or leaving out required fields. Structured Outputs is a stricter feature that makes the model follow a predefined structure, but the application still needs to verify values are correct.

I ran a real LLM regression test on a support triage assistant using 47 real customer messages from a public banking dataset. I tested three versions of an OpenAI model: an older version, the current production model, and a newer candidate. The newer model followed the format every time. The production model made a formatting mistake quietly, on every refund question in the sample. It spelled one label Request_refund with a capital R instead of the lowercase request_refund the system expects. A human reading the reply would call it correct, but a program matching labels exactly would silently drop every one of those tickets. This highlights the danger: a model can sound correct to a person and still be wrong for the software that uses its answer.

简体中文 version →