Why do monitoring, CMDB, service catalogs, and AI agents so often disagree about the same system? Why Systems Fail is a practical, beginner-friendly guide to ontology engineering for modern operations. Across 48 chapters, it turns abstract ideas-identity, meaning, time, provenance, constraints, and evidence-into concrete techniques for observability, CMDB, knowledge graphs, and AI-agent root cause investigation.
You'll learn how to:
- distinguish terms, concepts, records, and real-world entities;
- model services, resources, deployments, dependencies, telemetry, alerts, and incidents;
- use RDF, RDFS, OWL 2, SPARQL 1.1, and SHACL together;
- resolve identities across multiple data sources without hiding conflicts;
- represent time, provenance, confidence, and data quality;
- connect OpenTelemetry, ServiceNow CSDM, Backstage, Palantir Ontology, and metadata catalogs to a coherent semantic model;
- design evidence-bounded GraphRAG and AI-agent investigation workflows;
- test, version, govern, secure, and evolve an ontology in production.
Each chapter includes a diagram, practical examples, exercises, and traceable sources. The book also provides a bilingual glossary, standards quick reference, an executable CMDB/observability case study, exercise answers, and a curated reading path.
Written for platform engineers, SREs, observability teams, CMDB practitioners, data architects, knowledge-graph builders, and technical leaders who need systems to agree on what their data means.