Jev-Driven SRE Diagnosis: What Worked and What Failed
🇬🇧 English
In our first study, we experimented with Jev as a decision aid for an LLM agent. The agent diagnosed and repaired incidents; Jev helped rank the agent's proposed tests and reviewed the evidence before submission. That post ended with a more ambitious idea: giving Jev a broad view of the cluster and letting its fast, cheap judgments guide the investigation. In this post, we present a Jev-driven diagnosis pipeline without any LLM agent. The pipeline programmatically collects and organizes cluster evidence, then feeds it to Jev. Jev selects a likely root cause and supporting observations, and the pipeline uses them to assemble a diagnosis report. Across 21 SREGym-Lite faults, the Jev-driven pipeline passes 80 of 105 diagnoses (76.2%), with a median diagnosis time of 14.6 seconds.
Jev answers questions by choosing from a supplied set of options. To use it for diagnosis, we need to provide both the evidence and the possible answers. We added a programmatic collector to turn cluster state into those inputs. First, the collector reads Kubernetes objects, events, recent pod logs, and resource usage. It groups the observations by component, such as a Deployment, and summarizes signs of failure. Jev receives these summaries and chooses a likely source to inspect.
The pipeline then gathered more detail about nginx-thrift and gave Jev 26 evidence items, including these two: E6: The nginx-thrift Pod was OOMKilled. E10: The Pod has a 16Mi memory limit, although its template says 256Mi. gatekeeper-mutating-webhook-configuration matches this Pod. Among the follow-up questions, Jev answered: Question: Is nginx-thrift the origin, a victim, or unrelated? Jev: origin Question: What category names the cause? Jev: admission_or_namespace_policy Question: Which evidence item best shows the mechanism? Jev: E10
🇨🇳 简体中文
Jev 驱动的 SRE 诊断结果与失败案例
在我们的首次研究中,我们将 Jev 作为 LLM 智能体的决策辅助工具进行了实验。该智能体负责诊断和修复事件;Jev 帮助对智能体提出的测试进行排序,并在提交前审查证据。那篇博文最后提出了一个更具雄心的想法:为 Jev 提供集群的全局视图,并利用其快速、低成本的判断来引导调查。在本文中,我们展示了一个完全不依赖 LLM 智能体的 Jev 驱动诊断管道。该管道以程序化方式收集和整理集群证据,然后将其输入 Jev。Jev 选择可能的根本原因及支持性观察结果,管道则利用这些信息组装诊断报告。在 21 个 SREGym-Lite 故障中,Jev 驱动的管道在 105 次诊断中通过了 80 次(76.2%),中位诊断时间为 14.6 秒。
Jev 通过从提供的选项集中进行选择来回答问题。为了将其用于诊断,我们需要同时提供证据和可能的答案。我们添加了一个程序化收集器,将集群状态转换为这些输入。首先,收集器读取 Kubernetes 对象、事件、最近的 Pod 日志和资源使用情况。它按组件(例如 Deployment)对观察结果进行分组,并总结故障迹象。Jev 接收这些摘要并选择一个可能的来源进行检查。
随后,管道收集了关于 nginx-thrift 的更多细节,并为 Jev 提供了 26 条证据项,包括以下两条:E6:nginx-thrift Pod 被 OOMKilled。E10:该 Pod 的内存限制为 16Mi,尽管其模板中写的是 256Mi。gatekeeper-mutating-webhook-configuration 与该 Pod 匹配。在后续问题中,Jev 回答如下:问题:nginx-thrift 是起源、受害者还是无关?Jev:起源。问题:哪个类别命名了原因?Jev:admission_or_namespace_policy。问题:哪条证据项最能展示机制?Jev:E10。
FAQ:Jev 驱动的诊断管道的准确性和速度如何?
Jev 驱动的管道在 21 个 SREGym-Lite 故障中实现了 76.2% 的准确率(105 次诊断中通过 80 次),中位诊断时间为 14.6 秒。
FAQ Q:Jev 驱动的诊断管道的准确性和速度如何?
FAQ A:Jev 驱动的管道在 21 个 SREGym-Lite 故障中实现了 76.2% 的准确率(105 次诊断中通过 80 次),中位诊断时间为 14.6 秒。