HeadlinesBriefing favicon HeadlinesBriefing.com

تNormalization فشلات غير قابل التفسير في تطوير الذكاء الاصطناعي

Hacker News •
×

في حلقة حديثة من President Curtis، ي strugglingPresident بفتح باب في منتين منبيتين. هذه الأبواب لا تعمل بسبب عوائق في طريقها: جسم في البداية، ثم ما يعادل Billionworth من الذهب. في الحالتين، استجابة للإحباط، ي mutterCharacter "الشيء الغبي يكره." هذا ليس نموذج معقول للأبواب! الأبواب لا يجب أن "تكره" بشكل غير قابل للشرح! وجدت هذه اللحظات مثيرة للسخرية بشكل مجنون¹ لكن عقلي الغبي يكره فقط.

Jev: صناعة更多 أبواب تكره. إنترنلت يزدحم ب Jev، نموذج ذكاء اصطناعي طورته Type Safe AI، الذي يُعيد قيمًا مُ typed مع estimates probability. الأشياء المهمة عن Jev، كما أعتقد:她是 سريعة و cheap، يمكن بناء عليها بسرعة،她是 سريعة، و她是 cheap.我不是 particularly جيد في فهم ما التكنولوجيا ستحصل على adoption. لا تزال لا أفهم² Slack. انتظار. لا تزال يجب أن تفعل hard جزء؟ مشكلتي قد تكون expecting أن products work. لا one يشترى هذا ي تشغيل evals. هم ببساطة handing opaque أسئلة إلى Jev و getting opaque استجابات.

بشكل كريم، هذا يسمح لهما by checking "مُ powered by IA" مربع و shipping قبل Friday، وعندما هذا breaks لوجيكًا لأسفل، يمكنهم always shrug وقول "آه، IA يخطئ." Error budgets؟ failure modes؟ test sets؟ كل هذه يمكن معالجتها لاحقًا. يمكن للمستخدم اكتشاف failure rate! قمت بالفعل shipping!

False Confidence "أه،" strawman responding إلى منشورتي يرد، "أنت لم ت considered حقيقة أن Jev يعطيك confidence scores!" ما ستفعل به؟ لكي تفعل شيء معقول مع confidence scores، يجب أن يكون لديك فهم لـ calibration من هذه confidence scores، و ALSO نموذج لـ تكاليف uncertainty.

On calibration side: نسخة Jev's topline إعلانات هي mainly عن how well他们在 various benchmarks، لكن not عن how calibrated هم confidence scores. There's cookbook عن استخدام confidence scores للوصول إلى شجرة classification لكنها fundamentally not عن how good هم confidence scores. At best، people use confidence scores في cargo cult manner. At worst، people use them ك excuse لـ why API call فشل. النموذل was only 73% confident! That means error budget هو 27%!

Accountability متى button breaks على موقع ويب، عندي نموذج عن what should have happened. Where contract got broken. DNS لي broken. Somebody شiped slop الذي به Java Script syntax errors فقط along مسارات معينة. handler threw لم expects threw. Maybe access debug HTTP status 500، لكن expect أن يكون هناك شخص whose job هو فهم why endpoint 500ing. Ownership well-defined albe³.

ل المستخدمين، experience هو تقريبًا "stupid thing sucks." Software بالفعل feels capricious؛ more failures only change rate frustration. Seems like not much loss ل remove احتمال following failure إلى concrete cause. Sometimes things just suck. This leads إلى normalization of inexplicability.

My fear not أن more things will fail when things are accelerated ب LLM-driven development. They will. They have. Such is part of price of building things in novel manner. My fear is that "sometimes it just sucks" will be more and more accepted endpoint of investigations. This sad because LLM-accelerated development can indeed help us solve some of these issues. There are plenty automated QA workflows not written because lack of engineering time. The very eval that would get you most of the way to replacing (or justifying use of) Jev can be a few prompts away.

Tragedy of software engineering today is that we are actively engineering systems where neither user nor builder seems to have interest في checking whether there's body behind door. We just...