Software discipline

DevOps, Platform Engineering, and Reliability Guides

Improve delivery and production reliability through controlled deployments, reproducible infrastructure, observability, incident response, recovery, and ownership.

Planning this work

DevOps and reliability work should make a visible service safer to change and easier to recover. The useful system connects source control, testing, artifacts, infrastructure, configuration, database changes, deployment, monitoring, incident authority, communication, recovery, and evidence rather than treating automation as the outcome.

These guides cover planned delivery improvement and operations during an incident. They help teams identify the current constraint, build fast trustworthy feedback, preserve deployment and data-change options, instrument user journeys, coordinate responders, verify recovery, learn from failures, and leave cloud and operational assets under accountable ownership.

Practical guides in this discipline