This is an interesting experiment. Curious why you used LLM-as-a-Judge (gpt-6-astra). Did we have a ground truth of actual RCA done by a human to compare against?
> Ultimately, we believe that incorporating Jev-driven diagnosis into an SRE agent’s workflow is a significant step toward effectively combining System One and System Two models.
Why? What’s the end goal? Lower cost? Faster analysis?
I configured an agent to respond to pages via Slack. It has access to ClickStack (for telemetry), Kubernetes, and GitHub. It runs one of the Sonnet models. The cost is so low, responding to less than a dozen pages a day (mostly from a very sensitive error count alarm) that swapping got Jev makes zero sense. This is especially true if the results are less trustworthy.
It seems to me that Jev is designed for when many quick decisions, with low input context, need to be made. SRE work is probably best suited for very few critical decisions that need to be made, with high input context.
Why? What’s the end goal? Lower cost? Faster analysis?
I configured an agent to respond to pages via Slack. It has access to ClickStack (for telemetry), Kubernetes, and GitHub. It runs one of the Sonnet models. The cost is so low, responding to less than a dozen pages a day (mostly from a very sensitive error count alarm) that swapping got Jev makes zero sense. This is especially true if the results are less trustworthy.