Skip to main content

Primary sources for this edition

  1. IBM Research, ITBench (ICML 2025; github.com/itbench-hub/ITBench) — the open benchmark across SRE, CISO, and FinOps scenarios; source of the ~11–14% SRE resolution figures and scenario methodology.
  2. Artificial Analysis × IBM, ITBench-AA (May 2026) — independent frontier-model evaluation on Kubernetes incident diagnosis; all evaluated models below 50% on the headline precision metric, with published turn, token, and cost curves.
  3. SREGym (2026, arXiv) — live-benchmark research documenting reward hacking against fault injectors, alert-clearing false successes in prior benchmarks, and the measured diagnosis→mitigation gap (~69–88% conditional success).
  4. Microsoft Research, AIOpsLab — the live-environment agent evaluation framework this edition’s replay guidance builds on; and Microsoft’s Azure SRE Agent engineering posts — the 100+ tools → few core tools consolidation lesson, plus GA figures (1,300+ internal agents; 35,000+ incidents mitigated; 20,000+ hours saved) cited as first-party claims.
  5. AWS DevOps Agent GA materials (March–April 2026) and adoption guidance — recommendation-only starts, single-service scoping; vendor-reported pilot outcomes labeled as such.
  6. Anthropic — published multi-agent research-system engineering (orchestrator-worker gains at higher token cost) and the Model Context Protocol specification.
  7. OpenTelemetry GenAI SIG — the GenAI semantic conventions for model, agent, and tool telemetry; agent-application and framework conventions in active development; vendor adoption notes from major observability platforms.
  8. OWASP Top 10 for LLM Applications (LLM01: Prompt Injection) and practitioner threat research including CrowdStrike’s injection-technique taxonomy (150+ techniques; 300k+ analyzed adversarial prompts) — the Chapter 5 threat model’s public backbone.
  9. Gartner — Market Guide for AI SRE (January 2026); the 40%+ agentic-project cancellation prediction (2027); guardian-agent and multi-agent inquiry data as cited in the Field Guide.
  10. Category signals — Resolve AI’s $125M Series A at a $1B valuation (February 2026, category-record round); PagerDuty SRE Agent (Spring 2026 release); Datadog Bits AI SRE and peers — cited as market evidence, not endorsements.

How to read the numbers

The Field Guide’s three rules apply unchanged — provenance stated, ranges over points, your baseline beats every benchmark — with one engineer’s addendum: benchmark scores measure models under a benchmark’s harness; your POC measures a platform under yours. Neither transfers to the other automatically, which is exactly why Appendix A exists. Where this edition quotes a number, its class (independent benchmark, peer-reviewed research, vendor first-party, analyst prediction, market event) is stated inline; anything we could not source to that standard was cut.