Running the test takes around 5-6 hours per vendor depending on how much wrestle it requires (not every system is ideal for ai agents - including ourselves) but the actual work was to come up with the real world scenarios and what would be the expected RCA and remediation for this. Our own SRE team worked for 1.5 weeks for that.
Or if there are some features of an AI SRE tool that make it better than Claude + some MCPs, should those be captured in this same benchmark?