HeadlinesBriefing HeadlinesBriefing.com

Jev-Driven SRE Diagnosis Results and Failures

Hacker News •
×

In our first study, we experimented with Jev as a decision aid for an LLM agent. The agent diagnosed and repaired incidents; Jev helped rank the agent's proposed tests and reviewed the evidence before submission. That post ended with a more ambitious idea: giving Jev a broad view of the cluster and letting its fast, cheap judgments guide the investigation. In this post, we present a Jev-driven diagnosis pipeline without any LLM agent. The pipeline programmatically collects and organizes cluster evidence, then feeds it to Jev. Jev selects a likely root cause and supporting observations, and the pipeline uses them to assemble a diagnosis report. Across 21 SREGym-Lite faults, the Jev-driven pipeline passes 80 of 105 diagnoses (76.2%), with a median diagnosis time of 14.6 seconds.

Jev answers questions by choosing from a supplied set of options. To use it for diagnosis, we need to provide both the evidence and the possible answers. We added a programmatic collector to turn cluster state into those inputs. First, the collector reads Kubernetes objects, events, recent pod logs, and resource usage. It groups the observations by component, such as a Deployment, and summarizes signs of failure. Jev receives these summaries and chooses a likely source to inspect.

The pipeline then gathered more detail about nginx-thrift and gave Jev 26 evidence items, including these two: E6: The nginx-thrift Pod was OOMKilled. E10: The Pod has a 16Mi memory limit, although its template says 256Mi. gatekeeper-mutating-webhook-configuration matches this Pod. Among the follow-up questions, Jev answered: Question: Is nginx-thrift the origin, a victim, or unrelated? Jev: origin Question: What category names the cause? Jev: admission_or_namespace_policy Question: Which evidence item best shows the mechanism? Jev: E10

Source: Hacker News · Summarized by HeadlinesBriefing