Pangram verdict · v3.3
We believe that this text is a mix of AI and human-written content.
AI likelihood · overall
AIArticle text · 346 words · 4 segments analyzed
In our first study, we experimented with Jev as a decision aid for an LLM agent. The agent diagnosed and repaired incidents; Jev helped rank the agent's proposed tests and reviewed the evidence before submission. That post ended with a more ambitious idea: giving Jev a broad view of the cluster and letting its fast, cheap judgments guide the investigation. In this post, we present a Jev-driven diagnosis pipeline without any LLM agent. The pipeline programmatically collects and organizes cluster evidence, then feeds it to Jev. Jev selects a likely root cause and supporting observations, and the pipeline uses them to assemble a diagnosis report.
Across 21 SREGym-Lite faults, the Jev-driven pipeline passes 80 of 105 diagnoses (76.2%), with a median diagnosis time of 14.6 seconds.
Jev answers questions by choosing from a supplied set of options. To use it for diagnosis, we need to provide both the evidence and the possible answers. We added a programmatic collector to turn cluster state into those inputs. First, the collector reads Kubernetes objects, events, recent pod logs, and resource usage. It groups the observations by component, such as a Deployment, and summarizes signs of failure. Jev receives these summaries and chooses a likely source to inspect. The collector then gathers more detail about that component and prepares numbered evidence items. Jev decides whether the component is the origin, a downstream victim, or unrelated, and selects the evidence that best supports its answer. The pipeline uses these choices to assemble and submit a diagnosis. If the evidence cannot support the hypothesis, it examines another candidate. This version investigates candidates one at a time. Jev chooses among supplied options throughout the process. It does not generate commands or write the final report. The pipeline implementation is available on GitHub.
Figure 1. The diagnosis pipeline alternates programmatic evidence collection with Jev's focused decisions. Let us look at SREGym-Lite's mutating_webhook_resource_limits_social_network fault in the Social Network application. In this fault, pods created for nginx-thrift kept running out of memory. Its Deployment template specified a 256Mi memory limit, but new Pods had only 16Mi.