Skip to main content
Gartner’s canceled-40% is an aggregate; failures are specific. Twelve production failure modes, each with its detection signal and its antidote — then the build-vs-buy question answered the way engineers actually face it.

11.1 The catalog

Two meta-observations. First, every failure mode above is visible on a dashboard before it is visible in an outage — but only if Chapters 7 and 8 were built. Programs that skip evaluation and observability do not avoid these failures; they meet them without instruments. Second, the catalog is the interview: ask any platform vendor which of these twelve they have personally hit and what they shipped in response. Scar tissue is the only credential in a category this young.

11.2 Build vs. buy, the engineering edition

Capable teams can absolutely build an operations agent — a weekend with a frontier model, a ReAct loop, and kubectl produces a demo that will impress your leadership. The decision is not whether you can build the agent; it is whether you should own the harness. Tally what this book actually specified: the context subsystem with discovery, budgets, memory, and provenance (Ch. 2); the tool and skill layer with risk metadata and contract tests (Ch. 3); orchestration with typed handoffs and two-engine cost control (Ch. 4); the injection defense stack with sandboxing, egress control, and tokenization (Ch. 5); policy-as-code, approval engineering, execution safety, and a regulator-grade audit schema (Ch. 6); a golden set, judges, and regression gates (Ch. 7); decision-graph observability (Ch. 8); and the day-2 operating discipline (Ch. 9). The demo is a fifth of the system, and it is the fun fifth. The other four-fifths is undifferentiated heavy lifting for most organizations — and the exact surface where the twelve failure modes live. The honest decision rule: build if agent operations is your product, or your constraints are so unusual that no platform’s trust architecture fits — and staff it as a product team with a roadmap, not a side quest. Buy if your differentiation lies elsewhere, and spend your engineering on what no vendor can ship: your golden set, your context curation, your policy design, and your operating discipline. Either way, hold the same bar — Appendix B is written to audit a vendor and to scope an internal build with equal precision, and Appendix A will pressure-test whichever path you choose against your own incidents.
KEY TAKEAWAYFailures are specific, instrumented, and mostly self-inflicted at the harness layer. Whether you build or buy, you own the golden set, the policy, the context, and the discipline — and you should demand scar tissue, not slideware, from anyone who wants to own the rest.