July 19, 2026

Scaling CI-CD: The Governance Blueprint

Scaling CI-CD: The Governance Blueprint
Scaling CI-CD: The Governance Blueprint
M365 FM Podcast
Scaling CI-CD: The Governance Blueprint

CI/CD pipelines are designed to accelerate software delivery, but as organizations grow, speed alone is no longer enough. Without governance, automated pipelines can spread security risks, configuration drift, compliance violations, and inconsistent deployment practices faster than ever before.

In this episode, you'll learn why governance is a fundamental part of modern CI/CD rather than an afterthought. We explore how policies, security controls, and quality gates should be embedded directly into the software delivery process so every deployment follows the same trusted standards automatically.

The episode explains key concepts such as Policy as Code, Infrastructure as Code, automated testing, security scanning, compliance validation, approval workflows, and continuous monitoring. Instead of relying on manual reviews, successful organizations build pipelines that enforce governance consistently across every environment.

You'll also discover how platform engineering and standardized deployment templates help development teams move faster while maintaining security and operational consistency. The discussion covers common scaling challenges, including multiple development teams, cloud-native applications, infrastructure changes, and balancing developer autonomy with organizational control.

Finally, the episode shares practical best practices for building resilient CI/CD platforms that support rapid innovation without sacrificing quality, compliance, or reliability. Whether you're a developer, DevOps engineer, platform engineer, architect, or IT leader, this episode provides a clear introduction to designing governance-first CI/CD pipelines that scale with your organization.

In today's fast-paced software development landscape, effective governance is crucial for successful Continuous Integration and Continuous Deployment (CI-CD) practices. Over 82% of software companies have embraced CI/CD tools, highlighting their significance in modern workflows. However, scaling CI-CD governance presents challenges. Organizations often encounter risks such as third-party dependencies and permissions sprawl. These issues can hinder your team's ability to maintain security and efficiency. Understanding these challenges is essential for you to implement robust governance practices that support your development goals.

Key Takeaways

  • CI-CD governance ensures secure and efficient software development by integrating compliance and security into workflows.
  • Establish clear governance policies to define team responsibilities and enhance compliance in CI-CD processes.
  • Continuous monitoring helps identify issues early, allowing for proactive improvements in CI-CD performance.
  • Overcome cultural resistance by fostering collaboration and engaging teams in discussions about governance benefits.
  • Address tooling limitations by assessing current tools and designing a scalable roadmap for CI-CD governance.
  • Implement automated testing and validation to maintain code quality and streamline CI-CD processes.
  • Use key metrics like lead time and change failure rate to measure the effectiveness of your CI-CD governance.
  • Integrate AI tools to enhance governance capabilities, optimize testing, and improve security in CI-CD pipelines.

CI-CD Governance Framework

CI-CD Governance Framework

What is CI-CD Governance?

CI-CD governance refers to the set of policies, practices, and tools that ensure your continuous integration and continuous delivery processes operate securely and efficiently. It integrates compliance and security measures directly into your CI-CD pipelines. This approach allows you to manage risks associated with software development and deployment effectively. By embedding governance into your workflows, you create a structured environment where development teams can thrive while adhering to necessary regulations.

Importance of Governance

Effective governance plays a crucial role in the success of your CI-CD initiatives. Here are some key reasons why governance is essential:

Compliance and Security

  1. Continuous Compliance: Governance ensures that compliance is part of your CI-CD processes. This integration reduces the risk of security breaches and avoids manual audit preparation.
  2. Security Practices: Standardized security practices, such as using secret managers and automated security scans, help protect your applications. Immutable artifacts and audit trails provide transparency and accountability.
  3. Role-Based Access Control: Establishing secure boundaries for CI-CD workflows through role-based access controls and network restrictions enhances security. This approach allows you to manage who can access what, minimizing potential vulnerabilities.

According to leading technology companies, effective governance includes:

  • Continuous compliance and governance integrated into CI/CD pipelines.
  • Policies and infrastructure defined and managed as code, allowing for quick adaptation to regulatory changes.
  • Securely generated and stored immutable artifacts, audit trails, and logs.
  • Human systems designed to address vulnerabilities and foster a culture of security.

Best Practices

Implementing best practices in CI-CD governance can significantly enhance your pipeline's performance. Here are some strategies to consider:

  • Establish Clear Policies: Define clear governance policies that align with your organization's goals. This clarity helps teams understand their responsibilities and the importance of compliance.
  • Continuous Monitoring: Regularly monitor your CI-CD processes to identify areas for improvement. This practice ensures that you maintain high standards of security and compliance.
  • Iterative Improvements: Adopt an iterative approach to governance. Use feedback from audits and performance metrics to refine your processes continuously.

The impact of strong governance on CI-CD performance is evident in various frameworks. For instance, the ISO/IEC 27001:2022 framework demonstrates how governance enhances CI/CD security and compliance through defined policies and risk management. Similarly, the DORA framework highlights the importance of controlled change and ICT risk management in improving CI-CD performance.

By embedding governance into your CI-CD processes, you not only ensure compliance but also create a more resilient and efficient software delivery ecosystem.

Scaling CI-CD Challenges

Scaling CI-CD governance presents several challenges that can hinder your organization's progress. Understanding these obstacles is crucial for implementing effective governance practices.

Cultural Resistance

Cultural resistance is a significant barrier to adopting CI-CD governance. In fact, it accounts for 38 percent of the challenges organizations face. This resistance often stems from a reluctance to change established workflows and practices. Teams may feel comfortable with their current processes, making them hesitant to embrace new governance frameworks. To overcome this resistance, you must foster a culture that values collaboration and continuous improvement. Engaging your teams in discussions about the benefits of governance can help shift mindsets and encourage buy-in.

Tooling Limitations

Tooling limitations can also impede your efforts to scale CI-CD governance. Many organizations encounter various obstacles related to their toolchains. Here are some common challenges:

  • Expense and Effort: Maintaining complex toolchains often requires significant IT budget allocation, leading to high costs and time consumption.
  • Miscommunication and Misalignment: A lack of visibility and collaboration between development and operations teams can cause delays and increase errors.
  • Complexity of Tools: Diverse tools chosen by different teams create silos, complicating integration and standardization across the organization.

To address these tooling gaps, consider the following steps:

  1. Assessing the Current State: Understand your existing delivery performance, including identifying manual steps and bottlenecks.
  2. Designing a Scalable Roadmap: Incremental transformations are more successful. Start with one application and gradually expand.
  3. Building the Right Team and Culture: Collaboration among teams is essential for successful CI-CD adoption.

Additionally, the concept of "maintenance tax" highlights hidden costs in running CI-CD tools. Legacy systems and siloed processes can complicate governance scaling. Inconsistent tooling may lead to failed, insecure, or non-auditable releases.

Without proper governance, your CI-CD pipelines become vulnerable to various risks. These include unauthorized access and inadequate secret management. Such weaknesses can lead to significant security breaches, compromising both the reliability and security of the software you develop.

To summarize, addressing cultural resistance and tooling limitations is vital for scaling CI-CD governance effectively. By fostering a collaborative culture and assessing your tools, you can create a more resilient and efficient software delivery ecosystem.

Best Practices for CI-CD Governance

Establishing effective governance practices is essential for enhancing your CI-CD processes. By implementing clear policies and continuous monitoring, you can create a robust framework that supports compliance and security.

Establishing Clear Policies

Clear governance policies serve as the foundation for effective CI-CD practices. They help you define expectations and responsibilities for your development teams. Here are some recommended policies to consider:

  • Role-Based Access Control: Limit access to sensitive data and code to authorized personnel only.
  • Code Review and Approval Processes: Ensure that all code changes undergo thorough reviews to maintain quality and security.
  • Compliance and Security Monitoring: Implement systems to detect and prevent breaches proactively.
  • Automated Testing and Validation: Use automated testing to verify that code functions as expected before deployment.
  • Metrics and Reporting: Measure pipeline performance to identify areas for improvement.

By adopting these policies, you can streamline your CI-CD processes and enhance compliance rates. A clear governance model can shorten release cycles by eliminating delays caused by manual checks and unclear approvals. This efficiency leads to quicker feature delivery, boosting revenue and product momentum.

Evidence Explanation
Automated compliance reporting tools These tools integrate with CI/CD systems to track metrics like code review completion rates, deployment approval times, and security scan results, simplifying the reporting process and enhancing compliance.
Policy templates by centralized governance teams These templates set minimum requirements for all teams, promoting consistency and allowing flexibility, which supports rigorous audit trails and monitoring practices.
Structured governance layer Without it, rapid CI/CD releases can become chaotic, leading to duplicate efforts and fragmented evidence trails, which negatively impact compliance and efficiency.

Continuous Monitoring

Continuous monitoring is vital for maintaining the integrity of your CI-CD processes. It allows you to identify issues early and respond proactively. Here are some effective practices for continuous monitoring:

  • Troubleshoot CI/CD Issues Effectively: Quickly identify and resolve problems that arise during the CI-CD process.
  • Alert on Critical Scenarios: Set up alerts for critical issues that require immediate attention.
  • Monitor CI/CD Trends Over Time: Analyze trends to understand the performance of your CI-CD pipelines.
  • Create Dashboards for Quick Issue Identification: Use dashboards to visualize key metrics and facilitate rapid decision-making.
  • Establish Performance Baselines for Proactive Improvements: Define performance benchmarks to guide your continuous improvement efforts.

Metrics for Success

To measure the effectiveness of your CI-CD governance, you should track specific metrics. Here are some key metrics to consider:

Metric Description
Lead time to production Measures the time from when work begins until it is live in production, indicating the speed of the CI/CD process.
Number of bugs (change failure rate) Represents the percentage of deployments that fail in production, reflecting software quality and pipeline effectiveness.
Regression test duration Tracks the time taken to complete regression tests, which impacts the ability to deploy reliably and quickly.
Broken build time Measures how long a broken build remains unresolved, indicating the team's commitment to quality and responsiveness.
Production downtime during deployment Assesses the downtime experienced during deployments, crucial for applications requiring high availability.

By focusing on these metrics, you can gain insights into your CI-CD governance's performance and identify areas for improvement. Continuous monitoring and iterative improvements will help you refine your processes and enhance overall efficiency.

Incorporating templates and automation into your governance practices can further streamline your CI-CD processes. By using standardized templates, you can ensure consistency across teams and reduce the risk of errors. Automation can help you enforce policies and monitor compliance without adding significant overhead.

Tools for CI-CD Governance

Tools for CI-CD Governance

In the realm of CI-CD governance, selecting the right tools is essential for ensuring compliance and security. Various tools can help you streamline your processes and enhance your governance framework. Here are some of the most widely adopted tools in the software industry:

  1. Northflank
  2. GitHub Actions
  3. GitLab CI/CD
  4. CircleCI
  5. Jenkins
  6. Travis CI
  7. Azure DevOps
  8. Harness
  9. Bitbucket Pipelines
  10. TeamCity
  11. Argo CD
  12. Spinnaker
  13. AWS CodePipeline
  14. Google Cloud Build

These tools offer a range of features that can help you manage your CI-CD processes effectively.

Governance-Specific Tools

Governance-specific tools integrate seamlessly with popular CI-CD platforms. They enhance your ability to enforce policies and maintain compliance. Below is a table highlighting key features and governance considerations for some of these platforms:

CI/CD Platform Key Features Governance Considerations
GitHub Actions Broad developer ecosystem and integrations; Familiar workflows Governance and data residency may require extra controls; not Oracle-native
GitLab CI/CD Unified platform for repo, CI/CD, and security scanning; Strong DevSecOps capabilities Self-managed adds ops burden; SaaS may have governance constraints
Jenkins Maximum CI/CD customization; Extremely flexible and extensible Operational and security burden; plugin sprawl

These tools not only facilitate automated deployment but also help you maintain a structured governance model.

Integrating AI Governance

Integrating AI governance into your CI-CD processes can significantly enhance your governance capabilities. AI governance architecture operationalizes governance within the delivery lifecycle. This integration makes compliance enforcement a part of every change, ensuring that you maintain accountability without sacrificing agility.

Integrating AI into CI/CD effectively operationalizes AI TRiSM within the delivery lifecycle, making governance continuous, risk scoring automated, and compliance enforcement part of every change.

The benefits of integrating AI governance tools into your CI-CD pipelines include:

  • Faster Release Cycles: AI optimizes testing by running only relevant tests based on code changes.
  • Reduced Failure Rates: Predictive models help prevent risky deployments by analyzing code changes and deployment history.
  • Cost Optimization: AI recommends resource adjustments to improve cloud utilization.
  • Improved Security: Automated vulnerability detection enhances security in DevSecOps practices.
  • Proactive Incident Management: Predictive monitoring helps detect anomalies before they affect customers.

By embedding governance into the delivery pipeline rather than treating it as a separate layer, you can preserve agility while ensuring accountability.

Roadmap for Implementing Governance

Implementing CI-CD governance requires a structured approach. You can follow a roadmap that guides you through the process step by step. This roadmap will help you assess your current state, develop a governance framework, and implement it in phases.

Initial Assessment

Before you implement governance, conduct an initial assessment. This assessment helps you understand your current processes and identify areas for improvement. Here are the recommended steps for your initial assessment:

  1. Conduct a thorough audit of your existing software development lifecycle.
  2. Map out every step from code commit to production deployment.
  3. Identify manual handoffs and lengthy wait times.
  4. Assess environments for configuration drift.
  5. Evaluate team culture and skill sets.
  6. Analyze the current level of collaboration between Dev and Ops.

This assessment will provide you with valuable insights into your existing workflows and highlight areas that require attention.

Developing a Governance Framework

Once you complete your initial assessment, you can develop a governance framework tailored to your organization. Consider using established frameworks like:

  • Jenkins: A widely used automation server for CI/CD processes.
  • GitHub Actions: A tool that allows automation of workflows directly from GitHub.
  • dbt Cloud: A cloud-based service that supports CI/CD for data transformation.
  • SQLFluff: A tool for linting SQL code, gaining popularity for ensuring code quality.

These frameworks can help you establish clear policies and practices that align with your organization's goals.

Phased Implementation

Implementing governance should occur in phases. This gradual approach allows you to refine your processes and build credibility within your organization.

Pilot Programs

Start with pilot programs to test your governance model. Here’s how pilot programs contribute to successful scaling:

Step Description
Start Small Begin with a focused pilot in a single business unit or specific application type.
Validate Governance Use the pilot to validate your governance model and refine your enablement approach.
Build Credibility Achieve demonstrable success to build internal credibility before scaling.
Scale Intentionally Expand based on readiness, treating each expansion as an opportunity to refine the operating model.

Scaling Up

After successful pilot programs, you can scale up your governance implementation. However, be aware of common pitfalls during this phase:

Pitfall Description
Weak or missing automated testing Insufficient automated testing can leave code quality vulnerable to defects.
Slow or flaky pipelines Long or unreliable pipelines discourage frequent commits and slow down feedback loops.
Hard-coded secrets Storing sensitive information in code poses major security risks.
Environment drift Inconsistencies between environments can lead to unexpected issues during deployment.
No rollback or observability Lack of a rollback strategy can lead to extended downtime after failed deployments.
Misaligned teams Poor communication and unclear ownership can cause confusion and inefficiencies in the pipeline.

By following this roadmap, you can implement CI-CD governance effectively. This structured approach will help you create a resilient and efficient software delivery ecosystem.


Scaling CI-CD governance is essential for enhancing your software delivery processes. Organizations that succeed in this area often prioritize governance from the start. They implement change management strategies to ensure operational buy-in and standardize policies for effective oversight.

As you move forward, consider these key recommendations:

  • Embed governance into your CI/CD pipelines to ensure compliance and security.
  • Utilize automation to streamline governance processes and reduce labor intensity.
  • Establish continuous evaluation loops to adapt to evolving business needs.

The CI/CD tools market is projected to grow significantly, driven by the increasing adoption of DevOps practices and cloud-native applications. By embracing these trends, you can enhance productivity and reduce time-to-market, ultimately leading to improved software delivery performance.

Metric Description
Deployment Frequency Measures how often new releases are deployed to production.
Lead Time for Changes Time taken from code commit to deployment in production.
Change Failure Rate Percentage of changes that fail in production, indicating reliability.
Time to Restore Service Time taken to recover from a failure in production, reflecting operational resilience.

By following these practices, you can build a resilient and efficient software delivery ecosystem that meets your organization's needs.

FAQ

What is CI-CD governance?

CI-CD governance involves policies and practices that ensure your continuous integration and deployment processes are secure and efficient. It integrates compliance and security measures directly into your CI-CD workflows.

Why is CI-CD governance important?

Effective CI-CD governance helps you manage risks, maintain security, and ensure compliance. It creates a structured environment where development teams can thrive while adhering to necessary regulations.

How can I overcome cultural resistance to CI-CD governance?

Foster a culture of collaboration and continuous improvement. Engage your teams in discussions about the benefits of governance to encourage buy-in and shift mindsets.

What tools can help with CI-CD governance?

Popular tools include GitHub Actions, GitLab CI/CD, Jenkins, and CircleCI. These tools streamline processes and enhance your governance framework by integrating compliance and security measures.

How do I measure the success of CI-CD governance?

Track key metrics such as lead time to production, change failure rate, and production downtime during deployment. These metrics provide insights into your governance performance and areas for improvement.

What are some best practices for CI-CD governance?

Establish clear policies, implement continuous monitoring, and adopt iterative improvements. These practices enhance compliance and security while streamlining your CI-CD processes.

How can automation improve CI-CD governance?

Automation enforces policies and monitors compliance without adding significant overhead. It reduces manual errors and streamlines governance processes, allowing teams to focus on delivering high-quality software.

What is the role of AI in CI-CD governance?

AI enhances governance by automating compliance checks and risk assessments. It helps optimize testing, predict failures, and improve security, making governance a seamless part of your CI-CD pipeline.

🚀 Want to be part of m365.fm?

Then stop just listening… and start showing up.

👉 Connect with me on LinkedIn and let’s make something happen:

  • 🎙️ Be a podcast guest and share your story
  • 🎧 Host your own episode (yes, seriously)
  • 💡 Pitch topics the community actually wants to hear
  • 🌍 Build your personal brand in the Microsoft 365 space

This isn’t just a podcast — it’s a platform for people who take action.

🔥 Most people wait. The best ones don’t.

👉 Connect with me on LinkedIn and send me a message:
"I want in"

Let’s build something awesome 👊

1
00:00:00,000 --> 00:00:03,720
Your automation is breaking, it isn't happening obviously, it isn't happening all at once.

2
00:00:03,720 --> 00:00:07,840
But it is happening in the way that matters most, silently, across dozens of teams.

3
00:00:07,840 --> 00:00:10,040
Each one solving the same problem differently.

4
00:00:10,040 --> 00:00:13,800
You have seen this before, a team automates a process and it works beautifully, it works

5
00:00:13,800 --> 00:00:16,040
so well that 50 other teams wanted to.

6
00:00:16,040 --> 00:00:19,280
They ask for the code, you copy the Yamo, they modify it just a little bit.

7
00:00:19,280 --> 00:00:22,080
Now you have 50 variants of the same pipeline.

8
00:00:22,080 --> 00:00:25,680
Each one drifts further from the original, each one becomes harder to maintain, this is the

9
00:00:25,680 --> 00:00:27,680
moment you realize a hard truth.

10
00:00:27,680 --> 00:00:29,360
Automation without structure isn't efficiency.

11
00:00:29,360 --> 00:00:32,000
It is just technical debt moving at a higher velocity.

12
00:00:32,000 --> 00:00:36,440
The assumption was simple, automate everything and you will go faster, more pipelines, more

13
00:00:36,440 --> 00:00:38,880
deployments, more control.

14
00:00:38,880 --> 00:00:43,040
But the reality is different, pipelines brawl, Yamal hell inconsistent controls across

15
00:00:43,040 --> 00:00:46,760
hundreds of teams, you see security gaps where one team handles secrets completely differently

16
00:00:46,760 --> 00:00:47,760
from the next.

17
00:00:47,760 --> 00:00:51,720
The support load becomes so heavy that your platform engineers spend all day answering

18
00:00:51,720 --> 00:00:55,840
the same questions, they hear the same thing from slightly different teams.

19
00:00:55,840 --> 00:00:58,120
How do I deploy this the way you want me to?

20
00:00:58,120 --> 00:01:02,400
What separates world class delivery from fragile automation isn't the tool you choose, it's

21
00:01:02,400 --> 00:01:03,720
the structure you build.

22
00:01:03,720 --> 00:01:07,600
This episode is about moving from individual pipelines to a governed platform that actually

23
00:01:07,600 --> 00:01:08,760
scales.

24
00:01:08,760 --> 00:01:10,760
Why enterprise automation fails at scale?

25
00:01:10,760 --> 00:01:12,200
The pattern is predictable.

26
00:01:12,200 --> 00:01:16,440
It is so predictable that you can watch it happen in slow motion, a team automates a process

27
00:01:16,440 --> 00:01:17,440
and it works.

28
00:01:17,440 --> 00:01:21,320
The process is now repeatable, deployments that used to take hours now take minutes.

29
00:01:21,320 --> 00:01:25,080
The team is happy, everyone starts asking why can't we all do this so the approach gets

30
00:01:25,080 --> 00:01:26,800
shared, one team adopts it.

31
00:01:26,800 --> 00:01:30,600
And another, by the time the 10th team gets their hands on it, the original pipeline is

32
00:01:30,600 --> 00:01:34,080
unrecognizable, each team customizes it for their own context.

33
00:01:34,080 --> 00:01:37,720
The first team adds a security scan, the second team removes it because they think it is too

34
00:01:37,720 --> 00:01:39,520
slow for their use case.

35
00:01:39,520 --> 00:01:43,320
The third team pins it to a specific version while the fourth team leaves it open because

36
00:01:43,320 --> 00:01:44,760
they want the latest updates.

37
00:01:44,760 --> 00:01:48,600
The fifth team adds approval gates, the sixth team bypasses them for hot fixes, you now

38
00:01:48,600 --> 00:01:52,680
have 50 pipelines, they look the same, but they aren't, they are variants on a theme.

39
00:01:52,680 --> 00:01:56,720
Each one has its own logic, its own exceptions, its own way of thinking about what a deployment

40
00:01:56,720 --> 00:01:57,720
even means.

41
00:01:57,720 --> 00:01:59,200
Toolsprull comes next.

42
00:01:59,200 --> 00:02:03,480
Some teams use Azure DevOps because the organization standardized on it, others use GitHub

43
00:02:03,480 --> 00:02:05,720
actions because their code already lives there.

44
00:02:05,720 --> 00:02:09,160
A third group uses Jenkins because it was running in the data center before anyone thought

45
00:02:09,160 --> 00:02:10,160
about the cloud.

46
00:02:10,160 --> 00:02:14,240
A fourth team uses GitLab because they migrated once and don't want to move back.

47
00:02:14,240 --> 00:02:18,880
You have the same deployment problem solved five different ways across five different platforms.

48
00:02:18,880 --> 00:02:22,160
Security gaps appear because there is no consistent approach to anything.

49
00:02:22,160 --> 00:02:24,360
One team stores secrets and Azure Key Vault.

50
00:02:24,360 --> 00:02:25,760
Another uses GitHub secrets.

51
00:02:25,760 --> 00:02:30,560
A third team hard codes them in an Azure storage account for testing and never removes them.

52
00:02:30,560 --> 00:02:33,080
Approval gates exist in some pipelines and not in others.

53
00:02:33,080 --> 00:02:36,480
Some teams have automated security scans while others skip them because they conflict with

54
00:02:36,480 --> 00:02:37,880
their build process.

55
00:02:37,880 --> 00:02:41,160
You can't enforce the standard because there is no standard to enforce it.

56
00:02:41,160 --> 00:02:45,240
There is just precedent, fragmented, conflicting, precedent.

57
00:02:45,240 --> 00:02:46,840
The support load explodes.

58
00:02:46,840 --> 00:02:51,600
The platform team spends all their time answering variations of the same questions.

59
00:02:51,600 --> 00:02:53,280
How do I set up approval gates?

60
00:02:53,280 --> 00:02:55,120
What's the right way to handle secrets?

61
00:02:55,120 --> 00:02:57,520
So I pin this action to a commit asher.

62
00:02:57,520 --> 00:02:58,840
Each answer is contextual.

63
00:02:58,840 --> 00:03:00,520
Each context is slightly different.

64
00:03:00,520 --> 00:03:02,600
The cost stays invisible until you actually measure it.

65
00:03:02,600 --> 00:03:05,400
You see duplicated effort across teams solving the same problem.

66
00:03:05,400 --> 00:03:09,800
You see slower deployments because inconsistent pipelines mean inconsistent debugging.

67
00:03:09,800 --> 00:03:12,840
You see higher incidence rates because controls are applied randomly.

68
00:03:12,840 --> 00:03:18,000
An organization with 50 teams and 50 custom pipelines costs more to maintain than an organization

69
00:03:18,000 --> 00:03:20,440
with 50 teams and one governed platform.

70
00:03:20,440 --> 00:03:21,760
The root cause isn't laziness.

71
00:03:21,760 --> 00:03:22,760
It isn't bad tooling.

72
00:03:22,760 --> 00:03:25,400
It is not that your teams don't care about best practices.

73
00:03:25,400 --> 00:03:27,680
The root cause is a missing operating model.

74
00:03:27,680 --> 00:03:29,840
There is no clear decision about what governance means.

75
00:03:29,840 --> 00:03:31,520
There is no structure to enforce it.

76
00:03:31,520 --> 00:03:34,280
There are no templates that make the right choice the easy choice.

77
00:03:34,280 --> 00:03:36,800
Without structure, teams optimize locally.

78
00:03:36,800 --> 00:03:39,000
Each choice makes sense in isolation.

79
00:03:39,000 --> 00:03:41,640
But together, they create fragility.

80
00:03:41,640 --> 00:03:43,120
The cognitive load problem.

81
00:03:43,120 --> 00:03:46,040
Fragmentation creates a problem you won't see on a dashboard.

82
00:03:46,040 --> 00:03:47,600
It lives in the heads of your engineers.

83
00:03:47,600 --> 00:03:51,680
Every person touching a pipeline has to manage a constellation of decisions.

84
00:03:51,680 --> 00:03:53,960
They have to figure out the version control strategy.

85
00:03:53,960 --> 00:03:55,760
They have to decide on artifact naming.

86
00:03:55,760 --> 00:03:58,760
They have to navigate environment promotion and approval gates.

87
00:03:58,760 --> 00:04:02,000
They have to handle secrets and credentials in a governed platform.

88
00:04:02,000 --> 00:04:03,000
These are constants.

89
00:04:03,000 --> 00:04:04,440
They are part of the operating model.

90
00:04:04,440 --> 00:04:05,800
Everyone knows how they work.

91
00:04:05,800 --> 00:04:08,200
But in a sprawled environment, they are variables.

92
00:04:08,200 --> 00:04:11,800
Each team invents their own answer and each answer is slightly different.

93
00:04:11,800 --> 00:04:15,520
When an engineer moves from team A to team B, they can't assume anything.

94
00:04:15,520 --> 00:04:18,040
The Yammer looks familiar, but the logic is different.

95
00:04:18,040 --> 00:04:19,440
The secret management is different.

96
00:04:19,440 --> 00:04:20,760
The approval rules are different.

97
00:04:20,760 --> 00:04:22,480
This is extraneous cognitive load.

98
00:04:22,480 --> 00:04:25,280
It is mental effort spent on infrastructure instead of the product.

99
00:04:25,280 --> 00:04:27,360
You aren't thinking about the feature you're building.

100
00:04:27,360 --> 00:04:29,840
You're thinking about the mechanics of how to deploy it.

101
00:04:29,840 --> 00:04:34,200
You are forced to learn new conventions because the last team's rules don't apply here.

102
00:04:34,200 --> 00:04:38,480
The data shows that when teams spend more than 20% of their time fighting infrastructure,

103
00:04:38,480 --> 00:04:40,040
they stop shipping features.

104
00:04:40,040 --> 00:04:41,920
They are just managing complexity.

105
00:04:41,920 --> 00:04:45,200
And when you're managing complexity, you aren't solving hard problems.

106
00:04:45,200 --> 00:04:48,080
You're just context switching between different deployment models.

107
00:04:48,080 --> 00:04:49,720
The platform team feels this first.

108
00:04:49,720 --> 00:04:52,920
They end up answering the same questions from 50 different contexts.

109
00:04:52,920 --> 00:04:54,400
They get asked about approval gates.

110
00:04:54,400 --> 00:04:55,840
They get asked about security scans.

111
00:04:55,840 --> 00:04:57,560
They get asked about naming conventions.

112
00:04:57,560 --> 00:05:01,200
But because every team is doing something different, the answers are always contextual.

113
00:05:01,200 --> 00:05:04,440
The answer depends on what that specific team already built.

114
00:05:04,440 --> 00:05:09,360
So the platform team spends hours researching and writing documentation that only applies

115
00:05:09,360 --> 00:05:10,520
to one group.

116
00:05:10,520 --> 00:05:14,480
By the time Friday rolls around, another team asks the same question in a slightly different

117
00:05:14,480 --> 00:05:15,480
way.

118
00:05:15,480 --> 00:05:16,480
And the documentation doesn't fit.

119
00:05:16,480 --> 00:05:18,360
The platform team stops building the platform.

120
00:05:18,360 --> 00:05:21,720
They become a support desk for the fragments that escaped into the wild.

121
00:05:21,720 --> 00:05:23,440
This isn't a documentation problem.

122
00:05:23,440 --> 00:05:27,680
You could write a thousand-page guide on best practices and it wouldn't change the reality

123
00:05:27,680 --> 00:05:28,680
on the ground.

124
00:05:28,680 --> 00:05:31,120
The documentation would just be a map of the chaos.

125
00:05:31,120 --> 00:05:32,360
It's an architecture problem.

126
00:05:32,360 --> 00:05:33,960
The system is asking too much of people.

127
00:05:33,960 --> 00:05:37,040
It expects every engineer to understand every variation.

128
00:05:37,040 --> 00:05:41,040
It expects every new hire to learn local quirks instead of one universal standard.

129
00:05:41,040 --> 00:05:44,840
When cognitive load is high, delivery slows down, not because the tools are slow, but because

130
00:05:44,840 --> 00:05:46,800
making decisions takes longer.

131
00:05:46,800 --> 00:05:48,800
And then, burn out follows.

132
00:05:48,800 --> 00:05:50,680
The solution isn't more automation.

133
00:05:50,680 --> 00:05:54,000
More automation in a fragmented system just makes the fragments move faster.

134
00:05:54,000 --> 00:05:57,200
The solution is fewer decisions, standardize the choices.

135
00:05:57,200 --> 00:05:58,720
Embed them into the platform.

136
00:05:58,720 --> 00:06:00,480
Make the right way, the default way.

137
00:06:00,480 --> 00:06:04,640
This allows teams to focus on their product instead of reinventing the deployment model

138
00:06:04,640 --> 00:06:05,640
for the hundreds of times.

139
00:06:05,640 --> 00:06:08,840
And this is where platform governance comes in.

140
00:06:08,840 --> 00:06:10,080
The cost of decentralization.

141
00:06:10,080 --> 00:06:12,280
The cognitive load problem has a financial shadow.

142
00:06:12,280 --> 00:06:18,880
It's a cost problem that executives understand immediately, and it's usually large enough

143
00:06:18,880 --> 00:06:21,200
to justify a complete reorganization.

144
00:06:21,200 --> 00:06:23,480
Decentralized CI/CD looks like autonomy.

145
00:06:23,480 --> 00:06:24,480
Teams feel free.

146
00:06:24,480 --> 00:06:25,480
They choose their own tools.

147
00:06:25,480 --> 00:06:27,400
They optimize for their own context.

148
00:06:27,400 --> 00:06:30,920
It feels like you're enabling speed, but in reality, it's expensive.

149
00:06:30,920 --> 00:06:32,440
It starts with tooling costs.

150
00:06:32,440 --> 00:06:35,360
When every team runs its own stack, every team needs its own licenses.

151
00:06:35,360 --> 00:06:37,440
You're paying for Azure DevOps pipelines.

152
00:06:37,440 --> 00:06:39,320
You're paying for GitHub actions runner hours.

153
00:06:39,320 --> 00:06:41,920
You're paying for Jenkins plugins and artifact storage.

154
00:06:41,920 --> 00:06:46,520
Essentialized platform negotiates volume pricing and spreads that cost across the whole company.

155
00:06:46,520 --> 00:06:48,000
Decentralized teams don't do that.

156
00:06:48,000 --> 00:06:52,160
They negotiate separate contracts or they use free tiers that break when they try to scale.

157
00:06:52,160 --> 00:06:56,880
50 teams with 50 tool stacks is never cheaper than 50 teams sharing one platform.

158
00:06:56,880 --> 00:06:58,640
It is significantly more expensive.

159
00:06:58,640 --> 00:07:00,000
Infrastructure costs make it worse.

160
00:07:00,000 --> 00:07:02,320
Essentialized platform scales horizontally.

161
00:07:02,320 --> 00:07:05,720
You provision enough runners to handle the peak load of the entire group.

162
00:07:05,720 --> 00:07:09,160
You use auto scaling and caching to stop doing the same work twice.

163
00:07:09,160 --> 00:07:10,880
You optimize for utilization.

164
00:07:10,880 --> 00:07:14,360
Decentralized teams provision for their own individual peaks.

165
00:07:14,360 --> 00:07:16,680
Team A needs five runners at 3pm every day.

166
00:07:16,680 --> 00:07:19,800
So they keep five runners sitting idle for the other 23 hours.

167
00:07:19,800 --> 00:07:22,600
Team B provisions, six runners because they expect to grow.

168
00:07:22,600 --> 00:07:25,200
Team C provisions 10 because they aren't sure what they need.

169
00:07:25,200 --> 00:07:29,840
Across 50 teams, you might have 250 runners doing the work of 50.

170
00:07:29,840 --> 00:07:32,600
That is a five times multiplier on your infrastructure bill.

171
00:07:32,600 --> 00:07:36,360
And because these are small, isolated clusters, they can't use advanced optimization.

172
00:07:36,360 --> 00:07:40,640
Essentialized platform can shift load between teams, but a decentralized landscape is stuck.

173
00:07:40,640 --> 00:07:42,400
The people costs are the highest of all.

174
00:07:42,400 --> 00:07:44,880
In this model, every team needs DevOps expertise.

175
00:07:44,880 --> 00:07:48,480
You need someone who understands how to build pipelines and troubleshoot failures.

176
00:07:48,480 --> 00:07:50,840
You need someone to manage secrets and service connections.

177
00:07:50,840 --> 00:07:54,560
At scale, that's the equivalent of one full time DevOps engineer per team.

178
00:07:54,560 --> 00:07:58,120
50 teams means you need 50 people with those specific responsibilities.

179
00:07:58,120 --> 00:08:00,960
Essentialized platform only needs one platform team.

180
00:08:00,960 --> 00:08:04,640
A group of five to eight people can serve 100 product teams.

181
00:08:04,640 --> 00:08:06,480
They aren't taking away all the responsibilities.

182
00:08:06,480 --> 00:08:08,440
They are just removing the routine decisions.

183
00:08:08,440 --> 00:08:12,840
The product teams still own their pipelines, but they don't have to be infrastructure experts.

184
00:08:12,840 --> 00:08:13,840
The math is brutal.

185
00:08:13,840 --> 00:08:17,920
One platform team supporting 50 groups costs less than 50 groups hiring 50 engineers.

186
00:08:17,920 --> 00:08:19,520
Then you have the governance costs.

187
00:08:19,520 --> 00:08:22,320
Every team has to implement its own compliance checks.

188
00:08:22,320 --> 00:08:24,920
Ordered trails are scattered across different tools.

189
00:08:24,920 --> 00:08:27,400
Security standards drift because there is no way to enforce them.

190
00:08:27,400 --> 00:08:30,840
The compliance team has to audit every single pipeline separately.

191
00:08:30,840 --> 00:08:33,640
The security team can't roll out a new requirement to everyone at once.

192
00:08:33,640 --> 00:08:36,360
They have to go door to door team by team.

193
00:08:36,360 --> 00:08:38,320
Centralized governance means you do one audit.

194
00:08:38,320 --> 00:08:40,680
You do one implementation of security standards.

195
00:08:40,680 --> 00:08:43,280
You make one change and it propagates to everyone.

196
00:08:43,280 --> 00:08:44,800
The research is clear on this.

197
00:08:44,800 --> 00:08:50,520
Decentralized models cost 30 to 50 percent more per unit of delivery than centralized platforms.

198
00:08:50,520 --> 00:08:54,240
That includes the tools, the infrastructure, the people and the governance.

199
00:08:54,240 --> 00:08:56,720
Most enterprises don't measure this until it's too late.

200
00:08:56,720 --> 00:08:59,640
By then, they've already built the chaos.

201
00:08:59,640 --> 00:09:03,840
The risk of ungoverned pipelines cost and cognitive load are invisible problems until

202
00:09:03,840 --> 00:09:05,880
they blow up, but risk is different.

203
00:09:05,880 --> 00:09:06,880
Work blows up first.

204
00:09:06,880 --> 00:09:10,320
When teams own their own CI/CD, governance becomes optional.

205
00:09:10,320 --> 00:09:14,200
Nobody explicitly says they want to skip the rules, but without an enforcement mechanism

206
00:09:14,200 --> 00:09:17,840
or a clear standard to follow, governance just becomes a suggestion.

207
00:09:17,840 --> 00:09:20,560
Suggestions never survive contact with a deadline.

208
00:09:20,560 --> 00:09:23,320
Security risks show up in very specific, provable ways.

209
00:09:23,320 --> 00:09:28,000
If you use GitHub actions, you are pulling code from the GitHub marketplace, which means

210
00:09:28,000 --> 00:09:32,680
you are running someone else's code inside your pipeline with full access to your secrets

211
00:09:32,680 --> 00:09:34,000
and repositories.

212
00:09:34,000 --> 00:09:38,240
This practice says you should pin actions to specific commit SHAs rather than version

213
00:09:38,240 --> 00:09:40,760
tags because a tag is just a pointer that can move.

214
00:09:40,760 --> 00:09:44,320
The person maintaining that action can push a new commit and move the tag at any time,

215
00:09:44,320 --> 00:09:47,480
and suddenly your pipeline is running different code than it was yesterday.

216
00:09:47,480 --> 00:09:50,960
When governance is decentralized, some teams know this and others don't.

217
00:09:50,960 --> 00:09:55,320
Some teams understand the risk, but skip the step because tracking commit charges annoying,

218
00:09:55,320 --> 00:09:58,600
while others might only do it for their most critical actions.

219
00:09:58,600 --> 00:10:01,560
This turns your organizational security profile into a lottery.

220
00:10:01,560 --> 00:10:06,240
When a third party action eventually gets compromised, your total exposure depends entirely on which

221
00:10:06,240 --> 00:10:09,320
specific teams pinned their code and which ones didn't.

222
00:10:09,320 --> 00:10:13,040
You can't audit this globally, so you have to go to every single team one by one to find

223
00:10:13,040 --> 00:10:14,040
the holes.

224
00:10:14,040 --> 00:10:16,040
Compliance gaps always show up in the audit trail.

225
00:10:16,040 --> 00:10:19,680
In a governed platform, every single deployment is logged with details on who approved it,

226
00:10:19,680 --> 00:10:21,560
when it happened, and why it was necessary.

227
00:10:21,560 --> 00:10:24,440
Everything is tied to a work item and is fully searchable.

228
00:10:24,440 --> 00:10:28,400
In a decentralized landscape, approval gates exist in some pipelines, but are missing in

229
00:10:28,400 --> 00:10:32,160
others, leaving logs scattered across multiple disconnected systems.

230
00:10:32,160 --> 00:10:36,080
If a regulator asks you to show who approved a specific change to production, you won't

231
00:10:36,080 --> 00:10:39,640
be able to answer without spending weeks digging through fragments.

232
00:10:39,640 --> 00:10:41,920
Data protection usually fails silently.

233
00:10:41,920 --> 00:10:46,080
You find secrets stored in environment variables instead of a dedicated manager, or passwords

234
00:10:46,080 --> 00:10:49,040
committed to repositories for testing that never got removed.

235
00:10:49,040 --> 00:10:52,520
You see database credentials and build logs and API keys in manifests.

236
00:10:52,520 --> 00:10:56,280
In a governed platform, the team enforces secrets handling at the template level, so teams

237
00:10:56,280 --> 00:10:59,600
can't accidentally commit a secret because the structure doesn't allow it.

238
00:10:59,600 --> 00:11:03,840
In a decentralized system, secrets handling depends entirely on team discipline and maintaining

239
00:11:03,840 --> 00:11:06,000
that level of discipline is expensive.

240
00:11:06,000 --> 00:11:09,320
Deployment risk increases because there are no consistent rollback procedures across the

241
00:11:09,320 --> 00:11:10,160
board.

242
00:11:10,160 --> 00:11:14,120
Some teams export their artifacts before they deploy, while others don't, and some have

243
00:11:14,120 --> 00:11:16,360
a documented plan while others just wing it.

244
00:11:16,360 --> 00:11:20,320
You might have one team running Canary releases, while the team next to them goes straight

245
00:11:20,320 --> 00:11:22,040
to 100% production.

246
00:11:22,040 --> 00:11:25,480
When something breaks, the time it takes to recover isn't consistent because it depends

247
00:11:25,480 --> 00:11:28,840
entirely on what that specific team decided to do that day.

248
00:11:28,840 --> 00:11:33,400
Incident response slows down because every team has built their own unique debugging environment.

249
00:11:33,400 --> 00:11:37,600
They use different monitoring tools, different log systems, and different runbooks.

250
00:11:37,600 --> 00:11:40,960
When a deployment causes an incident, the response team has to learn the specific teams tooling

251
00:11:40,960 --> 00:11:43,480
before they can even start troubleshooting the actual problem.

252
00:11:43,480 --> 00:11:47,040
An incident that should take 30 minutes to resolve ends up taking three hours because two

253
00:11:47,040 --> 00:11:50,240
and a half of those hours were spent just learning the infrastructure.

254
00:11:50,240 --> 00:11:52,040
The real cost comes when a breach happens.

255
00:11:52,040 --> 00:11:53,960
It isn't a matter of if but when.

256
00:11:53,960 --> 00:11:57,400
And data leaks or credentials are exposed, your organization has to prove that you had

257
00:11:57,400 --> 00:11:59,960
controls in place and were following best practices.

258
00:11:59,960 --> 00:12:04,040
If you have consistent approvals, audit trails, and security scanning enforced at the platform

259
00:12:04,040 --> 00:12:06,120
level, you can actually prove your case.

260
00:12:06,120 --> 00:12:10,320
You can show regulators and lawyers exactly what controls were active and defend your actions.

261
00:12:10,320 --> 00:12:12,000
If you don't have that, you are vulnerable.

262
00:12:12,000 --> 00:12:15,920
You aren't just vulnerable to the breach itself, but to the massive legal liability that

263
00:12:15,920 --> 00:12:16,920
follows it.

264
00:12:16,920 --> 00:12:20,040
This is where platform governance changes the entire game.

265
00:12:20,040 --> 00:12:22,280
What platform governance actually means.

266
00:12:22,280 --> 00:12:24,440
The problem with governance is the word itself.

267
00:12:24,440 --> 00:12:28,280
When people hear governance, they immediately think of control, restrictions, and someone

268
00:12:28,280 --> 00:12:30,240
in an office saying no to their progress.

269
00:12:30,240 --> 00:12:31,960
That isn't what platform governance is.

270
00:12:31,960 --> 00:12:34,400
Platform governance is actually the opposite of control.

271
00:12:34,400 --> 00:12:37,440
It is freedom engineered directly into your infrastructure.

272
00:12:37,440 --> 00:12:38,720
Old governance was reactive.

273
00:12:38,720 --> 00:12:42,800
You set a policy, someone broke it, and then you audited them and wrote a stronger policy.

274
00:12:42,800 --> 00:12:46,880
Someone found a loophole, so you added an exception and then handling those exceptions became

275
00:12:46,880 --> 00:12:48,760
a second layer of governance.

276
00:12:48,760 --> 00:12:53,560
Eventually, the whole system became an obstacle that people just tried to navigate around.

277
00:12:53,560 --> 00:12:54,880
Platform governance is proactive.

278
00:12:54,880 --> 00:12:58,760
Instead of forbidding bad things, it makes bad things impossible by design rather than by

279
00:12:58,760 --> 00:12:59,760
restriction.

280
00:12:59,760 --> 00:13:03,800
A governed platform provides specific things out of the box, like standardized pipeline templates

281
00:13:03,800 --> 00:13:05,800
that every team uses as their starting point.

282
00:13:05,800 --> 00:13:09,200
They use them because the template is easier than building from scratch, not because they

283
00:13:09,200 --> 00:13:10,200
are forced to.

284
00:13:10,200 --> 00:13:14,640
These templates have security checks baked in, automated approval gates between environments,

285
00:13:14,640 --> 00:13:17,760
and consistent secrets handling that happens at the platform level.

286
00:13:17,760 --> 00:13:20,080
The key is that these features are baked in.

287
00:13:20,080 --> 00:13:23,120
They aren't bolted on, they aren't optional, and they aren't just a suggestion.

288
00:13:23,120 --> 00:13:26,040
When you build a template, you embed the controls inside the code.

289
00:13:26,040 --> 00:13:30,360
A team takes that template and adds their own logic to fit their specific context, but

290
00:13:30,360 --> 00:13:34,200
the core steps like the build and the security scanning are not customizable.

291
00:13:34,200 --> 00:13:38,200
They are part of the fundamental structure, and they run automatically every time.

292
00:13:38,200 --> 00:13:41,680
A team can't skip a security scan because they forgot to add it to their workflow.

293
00:13:41,680 --> 00:13:45,400
The scan is already there as part of what it means to use the template, and it runs

294
00:13:45,400 --> 00:13:47,680
whether they explicitly call for it or not.

295
00:13:47,680 --> 00:13:49,920
This creates a very clear ownership boundary.

296
00:13:49,920 --> 00:13:54,120
The platform team owns the templates and the policies, which defines what must happen.

297
00:13:54,120 --> 00:13:58,320
The product team's own the logic inside those templates, which defines how it happens

298
00:13:58,320 --> 00:13:59,800
for their specific app.

299
00:13:59,800 --> 00:14:03,640
They can customize and extend the work, but they cannot remove the core structures.

300
00:14:03,640 --> 00:14:05,480
This boundary is the entire model.

301
00:14:05,480 --> 00:14:07,800
It isn't a constraint, it's a liberation.

302
00:14:07,800 --> 00:14:11,640
Product teams can stop thinking about governance entirely because they don't have to worry if

303
00:14:11,640 --> 00:14:15,160
they set up approval gates or secrets handling correctly.

304
00:14:15,160 --> 00:14:19,440
Those decisions were made once at the platform level and solved for everyone at the same time.

305
00:14:19,440 --> 00:14:22,400
Governance scales through automation instead of human review.

306
00:14:22,400 --> 00:14:26,760
A policy engine enforces the rules that the platform team writes once in code, and those

307
00:14:26,760 --> 00:14:29,520
rules apply to hundreds of teams instantly.

308
00:14:29,520 --> 00:14:33,880
When a rule needs an update, like a new security scan or a changed approval requirement,

309
00:14:33,880 --> 00:14:36,920
the team updates the template and the change propagates everywhere.

310
00:14:36,920 --> 00:14:40,480
In a decentralized model, the compliance team has to send an email to 50 different teams

311
00:14:40,480 --> 00:14:42,080
to explain a new requirement.

312
00:14:42,080 --> 00:14:46,040
They spend months answering questions, helping with implementation and following up with teams

313
00:14:46,040 --> 00:14:47,200
that ignored the email.

314
00:14:47,200 --> 00:14:51,080
They spend half a year rolling out a change that takes five minutes in a centralized model.

315
00:14:51,080 --> 00:14:52,560
The inversion is now complete.

316
00:14:52,560 --> 00:14:57,400
In a fragmented system, governance is expensive, slow, and usually imperfect.

317
00:14:57,400 --> 00:15:01,040
In a governed platform, it is cheap, fast, and consistent.

318
00:15:01,040 --> 00:15:03,640
It isn't overhead anymore, it's the foundation.

319
00:15:03,640 --> 00:15:06,840
The big insight here is that governance only works when it feels like freedom.

320
00:15:06,840 --> 00:15:10,520
When a team uses a template because it removes work from their plate and stops them from

321
00:15:10,520 --> 00:15:13,280
having to think about deployment mechanics, that is freedom.

322
00:15:13,280 --> 00:15:17,280
They are finally free to focus on their actual product instead of their pipeline.

323
00:15:17,280 --> 00:15:20,960
When a security control is embedded so deeply that the team doesn't even have to think about

324
00:15:20,960 --> 00:15:22,600
it, that isn't control.

325
00:15:22,600 --> 00:15:25,000
That is safety that doesn't require constant vigilance.

326
00:15:25,000 --> 00:15:29,440
When you can deploy a new security requirement to 200 teams in minutes through a template change,

327
00:15:29,440 --> 00:15:30,600
that isn't bureaucracy.

328
00:15:30,600 --> 00:15:32,560
That is leverage.

329
00:15:32,560 --> 00:15:34,640
This is the true meaning of platform governance.

330
00:15:34,640 --> 00:15:36,360
It doesn't mean we're watching you.

331
00:15:36,360 --> 00:15:39,440
It means we've removed the hard parts for you.

332
00:15:39,440 --> 00:15:42,880
The controls and standards are all there, but they are invisible because they are just

333
00:15:42,880 --> 00:15:43,960
how the platform works.

334
00:15:43,960 --> 00:15:47,840
The structure that makes this possible is the ring-based deployment model.

335
00:15:47,840 --> 00:15:49,560
The ring model, why it matters.

336
00:15:49,560 --> 00:15:51,320
A ring is fundamentally simple.

337
00:15:51,320 --> 00:15:55,520
It is a cohort of users or environments that receives a change in sequence, with validation

338
00:15:55,520 --> 00:15:58,400
between every single step, not everyone at once.

339
00:15:58,400 --> 00:16:01,840
Sequence, validation, then the next group, ring-0 is your internal team.

340
00:16:01,840 --> 00:16:05,520
The people actually building the system, they have a high tolerance for risk because they

341
00:16:05,520 --> 00:16:07,520
understand the architecture deeply.

342
00:16:07,520 --> 00:16:10,360
And they know how to troubleshoot or roll back if things go sideways.

343
00:16:10,360 --> 00:16:14,480
They get rapid feedback because they are actively monitoring and testing the changes they

344
00:16:14,480 --> 00:16:15,960
just made.

345
00:16:15,960 --> 00:16:18,120
Ring one is your early adopters and pilot users.

346
00:16:18,120 --> 00:16:19,760
This is real business context.

347
00:16:19,760 --> 00:16:22,320
It is not a lab or a controlled test environment.

348
00:16:22,320 --> 00:16:24,280
It is actual people doing actual work.

349
00:16:24,280 --> 00:16:27,560
The feedback here is structured because you are watching specific metrics instead of just

350
00:16:27,560 --> 00:16:28,960
hoping nothing breaks.

351
00:16:28,960 --> 00:16:32,320
When a business unit volunteers to pilot a new feature, they are using it in production

352
00:16:32,320 --> 00:16:37,360
with real data, real workflows, and real load-bound ring two and beyond is broad production.

353
00:16:37,360 --> 00:16:38,880
The full population.

354
00:16:38,880 --> 00:16:42,920
At this stage, the tolerance for risk is zero because the blast radius is enormous.

355
00:16:42,920 --> 00:16:46,600
If something breaks here, thousands of people know about it immediately.

356
00:16:46,600 --> 00:16:49,320
Each ring has entry criteria, tests pass.

357
00:16:49,320 --> 00:16:51,360
Policies are compliant, the deployment succeeded.

358
00:16:51,360 --> 00:16:54,960
The previous ring did not explode, and each ring has exit criteria.

359
00:16:54,960 --> 00:16:57,880
Metrics stay within bounds, no high severity incidents occur.

360
00:16:57,880 --> 00:17:02,080
User satisfaction is acceptable, and usage looks exactly like what you expected if you meet

361
00:17:02,080 --> 00:17:04,080
the criteria you move to the next ring.

362
00:17:04,080 --> 00:17:05,520
If you don't, you pause.

363
00:17:05,520 --> 00:17:09,600
You investigate and fix the issue in the current ring before you even think about expanding.

364
00:17:09,600 --> 00:17:11,120
The ring model isn't new.

365
00:17:11,120 --> 00:17:15,160
Microsoft has been using this for decades, literally decades, Windows deployment rings,

366
00:17:15,160 --> 00:17:18,640
Defender Signature updates, the Microsoft 365 apps roll out.

367
00:17:18,640 --> 00:17:23,160
When a new version of Windows is built, it doesn't ship to a billion devices on day one.

368
00:17:23,160 --> 00:17:27,400
It goes to the Windows inside a ring first, then to a small percentage of commercial devices,

369
00:17:27,400 --> 00:17:29,440
then a larger percentage, and finally to everyone.

370
00:17:29,440 --> 00:17:30,520
Every step is measured.

371
00:17:30,520 --> 00:17:32,200
Every step has exit criteria.

372
00:17:32,200 --> 00:17:36,480
Google uses rings for infrastructure. Amazon uses them for service deployments, matter uses

373
00:17:36,480 --> 00:17:39,800
them for app updates, the pattern is universal because it works.

374
00:17:39,800 --> 00:17:41,280
But here is the shift.

375
00:17:41,280 --> 00:17:44,120
Cloud native platforms make rings much easier to implement and measure than they used to

376
00:17:44,120 --> 00:17:45,120
be.

377
00:17:45,120 --> 00:17:48,880
Before the cloud, running a separate environment for ring zero meant buying and provisioning

378
00:17:48,880 --> 00:17:53,120
separate hardware, ring one meant another set of servers, and ring two meant your production

379
00:17:53,120 --> 00:17:54,120
servers.

380
00:17:54,120 --> 00:17:56,480
That involved capital expense and long lead times.

381
00:17:56,480 --> 00:17:59,880
You had to commit to the infrastructure before you even knew if the ring model made sense

382
00:17:59,880 --> 00:18:01,040
for your team.

383
00:18:01,040 --> 00:18:05,560
When it makes rings cheap, you spin up environments on demand, you scale them independently,

384
00:18:05,560 --> 00:18:07,000
and you tear them down when you are finished.

385
00:18:07,000 --> 00:18:10,520
The friction of creating separate environments has completely disappeared.

386
00:18:10,520 --> 00:18:12,280
The measurement improved too.

387
00:18:12,280 --> 00:18:16,120
Before the cloud, checking if ring zero was healthy, meant logging into a server and checking

388
00:18:16,120 --> 00:18:17,440
logs manually.

389
00:18:17,440 --> 00:18:21,760
Ring one meant checking another set of logs by hand across different tools.

390
00:18:21,760 --> 00:18:25,240
By the time you had enough data to make a decision, weeks had already passed.

391
00:18:25,240 --> 00:18:27,880
Cloud platforms integrate logging and metrics natively.

392
00:18:27,880 --> 00:18:31,360
You query ring zero performance in seconds and compare it to ring one immediately.

393
00:18:31,360 --> 00:18:34,120
You get dashboards that show the progression automatically.

394
00:18:34,120 --> 00:18:37,560
The decision to promote to the next ring is data driven, not a gut feeling.

395
00:18:37,560 --> 00:18:40,000
Rings reduce the blast radius, that is the core inside.

396
00:18:40,000 --> 00:18:42,800
If something breaks in ring zero, 50 internal users know about it.

397
00:18:42,800 --> 00:18:46,560
They understand the context and can provide feedback, so the damage is bounded.

398
00:18:46,560 --> 00:18:48,520
You fix it and then move to ring one.

399
00:18:48,520 --> 00:18:51,520
If something breaks there, maybe a hundred users know about it.

400
00:18:51,520 --> 00:18:53,600
It is still bounded and still manageable.

401
00:18:53,600 --> 00:18:57,800
But if you shipped directly to production without rings, if you went from the team tested

402
00:18:57,800 --> 00:19:00,240
it, to everyone uses it.

403
00:19:00,240 --> 00:19:01,240
And something broke.

404
00:19:01,240 --> 00:19:03,720
10,000 users would know about it immediately.

405
00:19:03,720 --> 00:19:06,800
An incident that should take two hours to fix would take two days.

406
00:19:06,800 --> 00:19:08,800
You would be operating in crisis mode.

407
00:19:08,800 --> 00:19:12,200
Every deployment would feel risky, so teams would become conservative and your deployment

408
00:19:12,200 --> 00:19:13,760
frequency would drop.

409
00:19:13,760 --> 00:19:17,680
Rings change the psychology of deployment because the blast radius is smaller at each step.

410
00:19:17,680 --> 00:19:18,840
You can actually take risks.

411
00:19:18,840 --> 00:19:21,800
You can deploy more frequently and try new things.

412
00:19:21,800 --> 00:19:22,800
Most things work.

413
00:19:22,800 --> 00:19:26,160
And the few things that don't work get caught in ring zero or ring one before they ever

414
00:19:26,160 --> 00:19:27,920
touch the whole organization.

415
00:19:27,920 --> 00:19:30,040
But rings only work if you measure them properly.

416
00:19:30,040 --> 00:19:31,600
Defining ring success metrics.

417
00:19:31,600 --> 00:19:35,720
The instinct when building a ring strategy is to measure the technical surface.

418
00:19:35,720 --> 00:19:38,040
Error rates, latency, availability.

419
00:19:38,040 --> 00:19:39,320
Did the system stay up?

420
00:19:39,320 --> 00:19:40,320
Did it perform?

421
00:19:40,320 --> 00:19:42,600
These metrics tell you if the infrastructure works.

422
00:19:42,600 --> 00:19:44,280
But they don't tell you if it is safe to expand.

423
00:19:44,280 --> 00:19:47,880
You can have a system with zero errors and perfect availability in ring zero.

424
00:19:47,880 --> 00:19:52,440
But that is an internal team of 50 people in a known environment with predictable load.

425
00:19:52,440 --> 00:19:56,480
That same system might completely fail under the load profile of ring one or it might work

426
00:19:56,480 --> 00:19:57,480
perfectly.

427
00:19:57,480 --> 00:19:59,840
You simply don't know because the context is different.

428
00:19:59,840 --> 00:20:02,200
Ring one has a thousand users instead of 50.

429
00:20:02,200 --> 00:20:04,200
Real business data instead of test data.

430
00:20:04,200 --> 00:20:05,960
And peak loads at different times of day.

431
00:20:05,960 --> 00:20:07,920
The metrics that matter aren't absolute.

432
00:20:07,920 --> 00:20:09,120
They are comparative.

433
00:20:09,120 --> 00:20:13,520
You need metrics that compare ring and against ring and minus one in the same time period.

434
00:20:13,520 --> 00:20:16,560
With the same business context, normalized for load differences.

435
00:20:16,560 --> 00:20:18,320
The core metrics look like this.

436
00:20:18,320 --> 00:20:19,560
First, deployment success rate.

437
00:20:19,560 --> 00:20:22,160
How often does the deployment complete without an error?

438
00:20:22,160 --> 00:20:25,600
In ring zero, you might tolerate failures because the team can troubleshoot.

439
00:20:25,600 --> 00:20:27,640
In ring one, failures are less acceptable.

440
00:20:27,640 --> 00:20:29,560
But less acceptable isn't a number.

441
00:20:29,560 --> 00:20:30,560
You have to define it.

442
00:20:30,560 --> 00:20:33,880
Maybe ring one requires 95% successful deployments.

443
00:20:33,880 --> 00:20:35,920
While ring two requires 98%.

444
00:20:35,920 --> 00:20:38,760
Incident count and severity matter more than raw error rates.

445
00:20:38,760 --> 00:20:42,320
Ring zero might generate five incidents a day in normal operation.

446
00:20:42,320 --> 00:20:46,000
While ring one might generate one, ring two should generate fewer than one on average.

447
00:20:46,000 --> 00:20:47,440
But the incident needs context.

448
00:20:47,440 --> 00:20:50,000
Is it a page load latency spike or a data corruption bug?

449
00:20:50,000 --> 00:20:54,320
The latency spike in ring one might be acceptable, but a data corruption bug is an automatic

450
00:20:54,320 --> 00:20:57,480
rollback regardless of which ring you are in.

451
00:20:57,480 --> 00:20:59,080
Next is key workflow latency.

452
00:20:59,080 --> 00:21:02,440
Measure the P95 latency for critical user parts, not the average.

453
00:21:02,440 --> 00:21:07,440
The P95, because the user who hits the 95th percentile is the one who is about to call support.

454
00:21:07,440 --> 00:21:09,800
In ring zero, that P95 might be three seconds.

455
00:21:09,800 --> 00:21:11,520
In ring one, it should be similar.

456
00:21:11,520 --> 00:21:15,040
If the latency in ring one is noticeably higher, something changed.

457
00:21:15,040 --> 00:21:16,040
Maybe it is the load.

458
00:21:16,040 --> 00:21:17,760
Or maybe it is a performance regression.

459
00:21:17,760 --> 00:21:20,600
You need to understand which one it is before moving to ring two.

460
00:21:20,600 --> 00:21:24,600
Then you have error rates by component, not a single error rate for the whole system.

461
00:21:24,600 --> 00:21:25,960
But component level data.

462
00:21:25,960 --> 00:21:29,000
The art service might be throwing errors while the payment service is fine.

463
00:21:29,000 --> 00:21:32,280
You need granularity because a ring one decision depends on context.

464
00:21:32,280 --> 00:21:35,800
An error in the notification service might not block you from moving to ring two.

465
00:21:35,800 --> 00:21:38,520
But an error in the payment service absolutely does.

466
00:21:38,520 --> 00:21:40,080
Business metrics matter just as much.

467
00:21:40,080 --> 00:21:42,720
Look at the time it takes to complete a critical process.

468
00:21:42,720 --> 00:21:45,120
In ring zero, your users are the development team.

469
00:21:45,120 --> 00:21:47,400
So they don't represent real usage patterns.

470
00:21:47,400 --> 00:21:50,840
In ring one, you have actual users doing actual work.

471
00:21:50,840 --> 00:21:53,600
Measure how long it takes them to complete a key workflow.

472
00:21:53,600 --> 00:21:57,880
If you are rolling out a new checkout flow, measure the checkout success in ring one and

473
00:21:57,880 --> 00:22:00,120
compare it to the baseline in production.

474
00:22:00,120 --> 00:22:03,880
If the ring one conversion is significantly lower, you have a problem that technical metrics

475
00:22:03,880 --> 00:22:05,600
alone will never catch.

476
00:22:05,600 --> 00:22:07,720
Don't forget user satisfaction in the pilot ring.

477
00:22:07,720 --> 00:22:09,080
This isn't just about NPS.

478
00:22:09,080 --> 00:22:10,560
It is about direct feedback.

479
00:22:10,560 --> 00:22:14,400
Our users complaining, not just about bugs, but about usability.

480
00:22:14,400 --> 00:22:17,880
They will tell you if the feature actually solves the problem they thought it would solve.

481
00:22:17,880 --> 00:22:21,560
A feature can be technically flawless and still miss the mark on the thing that matters

482
00:22:21,560 --> 00:22:22,720
most to the business.

483
00:22:22,720 --> 00:22:24,880
Finally, look at governance metrics.

484
00:22:24,880 --> 00:22:30,520
Policy violations, unapproved connector usage, access control drift, and audit log completeness.

485
00:22:30,520 --> 00:22:32,920
These aren't success metrics in the traditional sense.

486
00:22:32,920 --> 00:22:35,480
They are health metrics for the governance layer itself.

487
00:22:35,480 --> 00:22:36,800
Are the controls working?

488
00:22:36,800 --> 00:22:38,280
Are teams staying within policy?

489
00:22:38,280 --> 00:22:40,480
Are we capturing complete audit trails?

490
00:22:40,480 --> 00:22:42,800
The ring decision framework is simple in concept.

491
00:22:42,800 --> 00:22:44,280
But it is complex in execution.

492
00:22:44,280 --> 00:22:48,120
If all metrics are green and no high severity incidents occurred, you promote to the next

493
00:22:48,120 --> 00:22:49,120
ring.

494
00:22:49,120 --> 00:22:51,880
But you must define green explicitly, not subjectively.

495
00:22:51,880 --> 00:22:56,680
If your threshold for latency is that P95 must not increase by more than 10%.

496
00:22:56,680 --> 00:22:59,120
Then measure it, compare it, and decide.

497
00:22:59,120 --> 00:23:00,880
If metrics degrade, you pause.

498
00:23:00,880 --> 00:23:01,880
Do not rationalize.

499
00:23:01,880 --> 00:23:04,080
Do not hope it improves at scale, pause.

500
00:23:04,080 --> 00:23:05,360
Roll back if you have to.

501
00:23:05,360 --> 00:23:08,160
Fix the issue in the current ring and then re-evaluate.

502
00:23:08,160 --> 00:23:09,880
The whole point of using rings is bounded risk.

503
00:23:09,880 --> 00:23:13,920
If you skip ring one because you are optimistic about ring two, you have eliminated the only

504
00:23:13,920 --> 00:23:15,440
control you actually have.

505
00:23:15,440 --> 00:23:18,640
This requires a specific team structure to implement and maintain.

506
00:23:18,640 --> 00:23:20,280
The platform team operating model.

507
00:23:20,280 --> 00:23:22,240
A platform team is not a DevOps team.

508
00:23:22,240 --> 00:23:25,560
That distinction matters because it changes everything about how you staff the team, how

509
00:23:25,560 --> 00:23:27,360
you structure it, and what you measure.

510
00:23:27,360 --> 00:23:29,960
A DevOps team traditionally owns the infrastructure.

511
00:23:29,960 --> 00:23:31,160
They manage servers.

512
00:23:31,160 --> 00:23:32,640
They deploy applications.

513
00:23:32,640 --> 00:23:34,440
They troubleshoot production issues.

514
00:23:34,440 --> 00:23:35,440
They are responsive.

515
00:23:35,440 --> 00:23:37,760
When something breaks, they jump in and fix it.

516
00:23:37,760 --> 00:23:39,000
They are on call.

517
00:23:39,000 --> 00:23:40,400
They are reactive.

518
00:23:40,400 --> 00:23:43,440
A platform team owns the delivery system, not the deliverables.

519
00:23:43,440 --> 00:23:48,400
The system itself, the templates, the policies, the automation that enables hundreds of teams

520
00:23:48,400 --> 00:23:53,560
to deploy safely without requiring platform engineers to be involved in every deployment.

521
00:23:53,560 --> 00:23:54,560
They are product focused.

522
00:23:54,560 --> 00:23:55,560
They are proactive.

523
00:23:55,560 --> 00:23:57,720
Their job is to make the hard parts invisible.

524
00:23:57,720 --> 00:23:59,360
This requires different roles.

525
00:23:59,360 --> 00:24:02,680
A platform product owner who owns the roadmap and the outcomes.

526
00:24:02,680 --> 00:24:06,360
Not just build this feature, but what is the biggest friction point for product teams

527
00:24:06,360 --> 00:24:08,200
right now and how do we remove it?

528
00:24:08,200 --> 00:24:09,360
They run user research.

529
00:24:09,360 --> 00:24:10,360
They talk to teams.

530
00:24:10,360 --> 00:24:11,360
They measure satisfaction.

531
00:24:11,360 --> 00:24:15,600
They make decisions about what to build based on data, not on what the platform team feels

532
00:24:15,600 --> 00:24:16,600
like building.

533
00:24:16,600 --> 00:24:19,600
Platform engineers who actually build the templates and the automation.

534
00:24:19,600 --> 00:24:21,840
They are infrastructure focused, but product minded.

535
00:24:21,840 --> 00:24:25,720
They understand yaml and policies and service connections, but they also understand that

536
00:24:25,720 --> 00:24:30,000
their code is consumed by hundreds of developers who do not have time to learn infrastructure

537
00:24:30,000 --> 00:24:31,000
deeply.

538
00:24:31,000 --> 00:24:36,120
Every line of template code has to earn its existence by making something easier or safer

539
00:24:36,120 --> 00:24:37,760
for the people who use it.

540
00:24:37,760 --> 00:24:42,080
A security and compliance lead who embeds guardrails into the platform from day one.

541
00:24:42,080 --> 00:24:43,800
Not security as an afterthought.

542
00:24:43,800 --> 00:24:46,360
Not we will add a security scan later.

543
00:24:46,360 --> 00:24:48,440
Security as a first class design constraint.

544
00:24:48,440 --> 00:24:49,600
What controls do we need?

545
00:24:49,600 --> 00:24:51,080
What policies do we need to enforce?

546
00:24:51,080 --> 00:24:55,520
How do we encode them into templates so that teams cannot accidentally violate them?

547
00:24:55,520 --> 00:24:59,200
This role often bridges internal security and the platform team.

548
00:24:59,200 --> 00:25:02,320
They translate security requirements into platform capabilities.

549
00:25:02,320 --> 00:25:06,400
An adoption lead who trains teams, writes documentation and runs champions programs.

550
00:25:06,400 --> 00:25:10,400
The best platform in the world is worthless if nobody uses it or worse if people use it wrong

551
00:25:10,400 --> 00:25:12,280
because they did not understand it.

552
00:25:12,280 --> 00:25:15,960
This role is often overlooked and when it is adoption stalls.

553
00:25:15,960 --> 00:25:18,160
Training and documentation are not nice to have.

554
00:25:18,160 --> 00:25:19,480
They are core to the platform.

555
00:25:19,480 --> 00:25:22,840
The platform team treats product teams as customers.

556
00:25:22,840 --> 00:25:25,520
Actual customers, not subordinates, not teams they oversee.

557
00:25:25,520 --> 00:25:26,520
Customers.

558
00:25:26,520 --> 00:25:29,720
This means they have to earn adoption the same way a SaaS company earns adoption.

559
00:25:29,720 --> 00:25:31,080
Make the product better.

560
00:25:31,080 --> 00:25:33,840
Respond to feedback, measure satisfaction, iterate.

561
00:25:33,840 --> 00:25:37,680
If teams do not want to use the platform that is a product problem, not a disciplined problem,

562
00:25:37,680 --> 00:25:39,560
decision rights are explicitly clear.

563
00:25:39,560 --> 00:25:42,600
The platform team owns the templates and the policies.

564
00:25:42,600 --> 00:25:44,160
That is non-negotiable.

565
00:25:44,160 --> 00:25:47,120
Product teams own their pipeline logic within those templates.

566
00:25:47,120 --> 00:25:48,120
They customize.

567
00:25:48,120 --> 00:25:49,120
They extend.

568
00:25:49,120 --> 00:25:51,080
They make decisions about their own delivery.

569
00:25:51,080 --> 00:25:52,840
But they cannot remove the core controls.

570
00:25:52,840 --> 00:25:56,160
This clarity prevents the kind of negotiation that becomes exhausting.

571
00:25:56,160 --> 00:25:57,760
Can we skip the security scan?

572
00:25:57,760 --> 00:25:59,240
No, it is part of the template.

573
00:25:59,240 --> 00:26:01,320
But can we add this custom build step?

574
00:26:01,320 --> 00:26:02,800
Yes, go for it.

575
00:26:02,800 --> 00:26:05,400
The platform team does not approve every deployment.

576
00:26:05,400 --> 00:26:06,400
That is the entire model.

577
00:26:06,400 --> 00:26:07,400
They are not a gate.

578
00:26:07,400 --> 00:26:08,400
They are not a bottleneck.

579
00:26:08,400 --> 00:26:11,880
They provide tools and guardrails that make safe deployments the default.

580
00:26:11,880 --> 00:26:15,320
When a team deploy is, the platform has already enforced the security checks.

581
00:26:15,320 --> 00:26:16,640
The approval gates are already there.

582
00:26:16,640 --> 00:26:18,320
The audit trail is already being captured.

583
00:26:18,320 --> 00:26:20,280
The team does not wait for the platform team.

584
00:26:20,280 --> 00:26:21,840
The platform team is not even involved.

585
00:26:21,840 --> 00:26:25,120
This model scales through templates and automation, not through headcount.

586
00:26:25,120 --> 00:26:29,720
A platform team of five to eight engineers can serve 50 to 100 product teams.

587
00:26:29,720 --> 00:26:33,560
Not because they work 80 hour weeks, but because they work once and the work propagates to everyone

588
00:26:33,560 --> 00:26:34,960
through automation.

589
00:26:34,960 --> 00:26:35,960
Update a template once.

590
00:26:35,960 --> 00:26:38,200
200 pipelines update automatically.

591
00:26:38,200 --> 00:26:39,280
Deploy a new policy once.

592
00:26:39,280 --> 00:26:40,560
It applies to every team.

593
00:26:40,560 --> 00:26:43,760
Headcount does not grow linearly with the number of teams served.

594
00:26:43,760 --> 00:26:45,320
That is the leverage point.

595
00:26:45,320 --> 00:26:49,200
Success for the platform team is measured differently than for traditional engineering teams.

596
00:26:49,200 --> 00:26:51,280
Not by features shipped or bugs fixed.

597
00:26:51,280 --> 00:26:55,960
By outcomes, time to first deployment for new teams, adoption of golden paths, deployment

598
00:26:55,960 --> 00:26:59,880
frequency across the organization, developer satisfaction with the platform.

599
00:26:59,880 --> 00:27:03,680
These metrics tell you if the platform is actually working, if teams are faster, if teams

600
00:27:03,680 --> 00:27:07,320
are happier, if they are adopting what you have built, the templates themselves are where

601
00:27:07,320 --> 00:27:09,680
the magic happens.

602
00:27:09,680 --> 00:27:10,920
Golden path pipelines.

603
00:27:10,920 --> 00:27:12,520
The template is an abstraction.

604
00:27:12,520 --> 00:27:16,080
What actually gets built, what developers interact with is the golden path.

605
00:27:16,080 --> 00:27:20,480
A golden path is an opinionated, well supported way to build and deploy a service.

606
00:27:20,480 --> 00:27:25,160
The keyword is opinionated, not blank, not here is a yaml file fill in what you need

607
00:27:25,160 --> 00:27:26,160
to open it up.

608
00:27:26,160 --> 00:27:28,800
Here is how we have decided to do this.

609
00:27:28,800 --> 00:27:32,040
And it handles 80% of your use case out of the box.

610
00:27:32,040 --> 00:27:34,800
This is the difference between a template library and a golden path.

611
00:27:34,800 --> 00:27:37,080
A template library is a collection of components.

612
00:27:37,080 --> 00:27:40,080
You pick which ones you want, you assemble them, you are still making decisions.

613
00:27:40,080 --> 00:27:42,520
You are still solving the same cognitive load problem.

614
00:27:42,520 --> 00:27:44,480
A golden path is a decision made for you.

615
00:27:44,480 --> 00:27:48,520
Or more precisely, a decision the organization made that you inherit, you do not assemble,

616
00:27:48,520 --> 00:27:49,520
you extend.

617
00:27:49,520 --> 00:27:51,040
The core structure is already there.

618
00:27:51,040 --> 00:27:53,360
The path includes standard build steps.

619
00:27:53,360 --> 00:27:58,400
File, test, scan for security vulnerabilities, scan dependencies, build an artifact.

620
00:27:58,400 --> 00:28:02,360
These are not optional, they are part of what it means to follow this golden path.

621
00:28:02,360 --> 00:28:06,800
Standard deployment stages, dev test staging production, same across every service.

622
00:28:06,800 --> 00:28:10,400
Standard approval gates, who can approve promotion between stages.

623
00:28:10,400 --> 00:28:12,040
Standard rollback procedures.

624
00:28:12,040 --> 00:28:14,880
How to undo a deployment if something goes wrong.

625
00:28:14,880 --> 00:28:17,360
When you follow a golden path, you get all of this.

626
00:28:17,360 --> 00:28:20,240
You do not have to decide, should we run tests?

627
00:28:20,240 --> 00:28:24,560
The path already decided, you do not have to decide what security scans do we need.

628
00:28:24,560 --> 00:28:29,160
The path already decided, you do not have to decide who can approve production deployments.

629
00:28:29,160 --> 00:28:31,000
The path already decided.

630
00:28:31,000 --> 00:28:33,720
What you do decide is the logic specific to your service.

631
00:28:33,720 --> 00:28:37,920
Your build commands, your test suite, your deployment configuration, your health checks,

632
00:28:37,920 --> 00:28:41,680
you customize within the path, you extend the path, but you do not rewrite the path,

633
00:28:41,680 --> 00:28:44,880
teams can modify golden paths, they can request exceptions.

634
00:28:44,880 --> 00:28:47,800
We need a different build step because our tech stack is unusual.

635
00:28:47,800 --> 00:28:50,520
Just take a little bit of the amount to enter Tokyo's.

636
00:28:50,520 --> 00:28:53,320
Okay, the platform team reviews the request.

637
00:28:53,320 --> 00:28:55,840
If it is genuinely unique, you get an exception.

638
00:28:55,840 --> 00:28:56,920
But exceptions are tracked.

639
00:28:56,920 --> 00:29:01,280
If three teams request the same exception, the platform team reconsider the path.

640
00:29:01,280 --> 00:29:04,840
Maybe the path was not as universal as they thought, maybe it needs to evolve.

641
00:29:04,840 --> 00:29:06,160
This creates a feedback loop.

642
00:29:06,160 --> 00:29:07,840
The golden path is not sacred.

643
00:29:07,840 --> 00:29:09,480
It is not handed down from on high.

644
00:29:09,480 --> 00:29:13,000
It is a starting point that teams can shape through usage and feedback.

645
00:29:13,000 --> 00:29:14,280
The research on this is clear.

646
00:29:14,280 --> 00:29:18,960
70 to 80% of services can be deployed using just two to three golden path templates.

647
00:29:18,960 --> 00:29:21,400
Not because all services are identical, they are not.

648
00:29:21,400 --> 00:29:26,960
But because the deployment mechanics, how you build tests, scan, approve and deploy are remarkably consistent.

649
00:29:26,960 --> 00:29:30,240
The variation is in the application logic, not in the delivery pipeline.

650
00:29:30,240 --> 00:29:33,080
The remaining 20 to 30% need customization.

651
00:29:33,080 --> 00:29:34,640
Maybe they use an unusual tech stack.

652
00:29:34,640 --> 00:29:36,960
Maybe they have specialized performance requirements.

653
00:29:36,960 --> 00:29:39,560
Maybe they are data pipelines instead of web services.

654
00:29:39,560 --> 00:29:42,800
But even those services start from a baseline, they do not start from scratch.

655
00:29:42,800 --> 00:29:47,760
They start from a golden path designed for a different use case and they modify it rather than inventing from blank.

656
00:29:47,760 --> 00:29:51,640
Starting from a known good baseline rather than from zero is a dramatic difference.

657
00:29:51,640 --> 00:29:55,880
A team that starts with a blank YAML file spends weeks figuring out the right structure.

658
00:29:55,880 --> 00:29:58,880
The team that starts with a golden path spends a day adapting it.

659
00:29:58,880 --> 00:30:01,400
The difference compounds across the organization.

660
00:30:01,400 --> 00:30:03,520
Golden paths reduce decision fatigue.

661
00:30:03,520 --> 00:30:07,480
New teams do not spend weeks figuring out how do we do CI/CD here?

662
00:30:07,480 --> 00:30:08,520
They follow the path.

663
00:30:08,520 --> 00:30:10,120
They get to work on their application.

664
00:30:10,120 --> 00:30:11,280
The decisions are made.

665
00:30:11,280 --> 00:30:12,320
The structure is there.

666
00:30:12,320 --> 00:30:18,120
This is where the model moves from abstract governance into concrete reality that teams actually use.

667
00:30:18,120 --> 00:30:19,680
Template-driven governance.

668
00:30:19,680 --> 00:30:21,600
Templates aren't just a way to deliver code.

669
00:30:21,600 --> 00:30:22,600
They are the governance.

670
00:30:22,600 --> 00:30:26,280
That's the shift in thinking that changes how you actually build a platform.

671
00:30:26,280 --> 00:30:29,080
When you bake a control into a template, it becomes invisible.

672
00:30:29,080 --> 00:30:33,120
It isn't a rule that a manager has to enforce or a policy that an auditor has to check.

673
00:30:33,120 --> 00:30:34,480
It's just how the system works.

674
00:30:34,480 --> 00:30:35,480
It's structural.

675
00:30:35,480 --> 00:30:39,160
It is part of the very fabric of what happens every time someone uses that template.

676
00:30:39,160 --> 00:30:40,440
Think about security scans.

677
00:30:40,440 --> 00:30:44,200
You need SAS scanning for your source code and dependency scanning to catch compromised

678
00:30:44,200 --> 00:30:45,200
libraries.

679
00:30:45,200 --> 00:30:47,800
If you're building images, you need container scanning too.

680
00:30:47,800 --> 00:30:51,760
In this model, the template runs these scans automatically during the build stage.

681
00:30:51,760 --> 00:30:54,280
A developer doesn't sit there and decide whether to run them.

682
00:30:54,280 --> 00:30:55,960
They are part of the template.

683
00:30:55,960 --> 00:30:56,960
They run.

684
00:30:56,960 --> 00:30:57,960
The results come back.

685
00:30:57,960 --> 00:31:00,240
The build either passes or it fails based on what the scan found.

686
00:31:00,240 --> 00:31:04,600
This is a massive departure from a policy that says, "Team should run security scans."

687
00:31:04,600 --> 00:31:06,240
A policy is just a suggestion.

688
00:31:06,240 --> 00:31:07,240
People forget.

689
00:31:07,240 --> 00:31:09,360
They skip steps when a deadline is screaming at them.

690
00:31:09,360 --> 00:31:13,340
They don't always understand why the scan matters or they just do it wrong, but an embedded

691
00:31:13,340 --> 00:31:14,740
control isn't optional.

692
00:31:14,740 --> 00:31:15,840
It runs every single time.

693
00:31:15,840 --> 00:31:17,840
There is no negotiation and no exception.

694
00:31:17,840 --> 00:31:22,200
A team can't ship code that fails a scan because they'd have to physically remove the scan

695
00:31:22,200 --> 00:31:24,320
step from the template to bypass it.

696
00:31:24,320 --> 00:31:25,320
And they can't do that.

697
00:31:25,320 --> 00:31:26,320
The template is locked.

698
00:31:26,320 --> 00:31:27,560
They don't have right access.

699
00:31:27,560 --> 00:31:31,280
Even if they did, that change would trigger a code review where the platform team would

700
00:31:31,280 --> 00:31:32,360
see it and stop it.

701
00:31:32,360 --> 00:31:35,000
The same logic applies to production approvals.

702
00:31:35,000 --> 00:31:40,200
Starting from dev to test might be automatic, but moving from staging to production requires

703
00:31:40,200 --> 00:31:42,160
two approvals from specific roles.

704
00:31:42,160 --> 00:31:44,240
The template defines exactly who can sign off.

705
00:31:44,240 --> 00:31:46,800
These aren't manual checks that someone might forget to do.

706
00:31:46,800 --> 00:31:50,880
The pipelines simply won't move to the next stage without those digital signatures.

707
00:31:50,880 --> 00:31:52,520
The approvals are structural.

708
00:31:52,520 --> 00:31:54,160
Naming conventions work the same way.

709
00:31:54,160 --> 00:31:57,400
You don't want to hunt for artifacts or guess environment names.

710
00:31:57,400 --> 00:32:00,720
These rules aren't buried in a PDF somewhere hoping people follow them.

711
00:32:00,720 --> 00:32:02,160
They are enforced by the code.

712
00:32:02,160 --> 00:32:04,720
The template generates the names based on your rules.

713
00:32:04,720 --> 00:32:07,520
If an environment isn't named correctly, the pipeline fails.

714
00:32:07,520 --> 00:32:09,320
It's that simple.

715
00:32:09,320 --> 00:32:10,680
Promotion is also standardized.

716
00:32:10,680 --> 00:32:15,400
A service doesn't just jump from a developer's laptop straight into production.

717
00:32:15,400 --> 00:32:17,240
It has to flow through a sequence.

718
00:32:17,240 --> 00:32:20,160
Build, test, staging, and then production.

719
00:32:20,160 --> 00:32:21,640
The template enforces this path.

720
00:32:21,640 --> 00:32:24,280
You cannot deploy to production without going through staging first.

721
00:32:24,280 --> 00:32:25,960
The gates are built into the logic.

722
00:32:25,960 --> 00:32:29,960
Because the template is versioned, the platform team can add a new security scan or change

723
00:32:29,960 --> 00:32:32,680
an approval step and tag it as a new version.

724
00:32:32,680 --> 00:32:36,080
Making pipelines stay on their old version so nothing breaks unexpectedly.

725
00:32:36,080 --> 00:32:38,040
New pipelines start with the current version.

726
00:32:38,040 --> 00:32:41,040
This means the team doesn't wake up to a broken build because of a change they didn't see

727
00:32:41,040 --> 00:32:42,200
coming.

728
00:32:42,200 --> 00:32:43,560
Everything lives in source control.

729
00:32:43,560 --> 00:32:47,200
If someone wants to update the template, they submit a pull request to the platform team.

730
00:32:47,200 --> 00:32:50,200
It gets reviewed, it gets questioned, the history is visible to everyone.

731
00:32:50,200 --> 00:32:52,760
You can see exactly who changed what and why they did it.

732
00:32:52,760 --> 00:32:54,440
The real power here is leverage.

733
00:32:54,440 --> 00:32:57,920
The platform team updates the template once and that change can reach hundreds of teams

734
00:32:57,920 --> 00:32:59,080
automatically.

735
00:32:59,080 --> 00:33:02,560
There is no manual rollout and no please update your pipeline's emails.

736
00:33:02,560 --> 00:33:05,120
The next time a team triggers a build, they get the fix.

737
00:33:05,120 --> 00:33:07,520
If they're pinned to an old version, they get a warning.

738
00:33:07,520 --> 00:33:08,920
They can upgrade when they're ready.

739
00:33:08,920 --> 00:33:11,800
The teams actually like this because it removes the mental load.

740
00:33:11,800 --> 00:33:15,080
They don't have to think about approval gates or figure out which security scans are

741
00:33:15,080 --> 00:33:16,160
required this week.

742
00:33:16,160 --> 00:33:17,960
They don't have to worry about naming conventions.

743
00:33:17,960 --> 00:33:21,640
The template handles the how so they can focus on the what.

744
00:33:21,640 --> 00:33:23,720
It's governance that doesn't feel like governance.

745
00:33:23,720 --> 00:33:25,000
It doesn't require discipline.

746
00:33:25,000 --> 00:33:26,280
It requires nothing.

747
00:33:26,280 --> 00:33:27,840
It just works.

748
00:33:27,840 --> 00:33:31,240
But how this looks in practice depends on the tools you're using.

749
00:33:31,240 --> 00:33:33,240
Azure DevOps ring implementation.

750
00:33:33,240 --> 00:33:37,440
When you move from the theory into actual Azure DevOps configuration, these rings map directly

751
00:33:37,440 --> 00:33:40,040
to environments and multi-stage YAML pipelines.

752
00:33:40,040 --> 00:33:43,960
This is where the abstract ideas become the actual infrastructure your teams use every day.

753
00:33:43,960 --> 00:33:47,200
In Azure DevOps, you set up environments as specific objects.

754
00:33:47,200 --> 00:33:49,840
Ring zero becomes an environment called dev canary.

755
00:33:49,840 --> 00:33:51,200
Ring one is staging pilot.

756
00:33:51,200 --> 00:33:52,640
Ring two is production.

757
00:33:52,640 --> 00:33:57,000
Each one is a target for your deployment and each one has its own set of guardrails.

758
00:33:57,000 --> 00:34:00,160
The approval requirements get stricter as you move closer to the customer.

759
00:34:00,160 --> 00:34:02,800
For dev canary, you might have auto approvals.

760
00:34:02,800 --> 00:34:06,400
No one needs to sign off on the deployment to an internal canary environment because speed

761
00:34:06,400 --> 00:34:07,400
is the priority there.

762
00:34:07,400 --> 00:34:09,000
The team is iterating and testing.

763
00:34:09,000 --> 00:34:10,280
They're moving fast.

764
00:34:10,280 --> 00:34:13,680
Extra overhead would just slow them down so the environment lets them push code without

765
00:34:13,680 --> 00:34:14,920
a human in the loop.

766
00:34:14,920 --> 00:34:16,280
Staging pilot is different.

767
00:34:16,280 --> 00:34:20,240
It requires one approval, a designated person from the platform team, or a senior lead

768
00:34:20,240 --> 00:34:21,400
reviews the request.

769
00:34:21,400 --> 00:34:24,920
They make sure the change looks right and verify that the work items are linked.

770
00:34:24,920 --> 00:34:27,680
Once they hit approve, the deployment moves forward.

771
00:34:27,680 --> 00:34:28,960
Production is the final gate.

772
00:34:28,960 --> 00:34:30,640
And it requires two approvals.

773
00:34:30,640 --> 00:34:32,960
This is where you enforce the separation of duties.

774
00:34:32,960 --> 00:34:36,680
The person who wrote the code cannot be the only person who approves it for production.

775
00:34:36,680 --> 00:34:38,120
You need a second set of eyes.

776
00:34:38,120 --> 00:34:42,800
These two approvals act as a deliberate speed bump to stop accidental deployments or unreviewed

777
00:34:42,800 --> 00:34:45,080
changes from hitting your users.

778
00:34:45,080 --> 00:34:48,520
Beyond just people clicking buttons, each environment has automated checks.

779
00:34:48,520 --> 00:34:52,080
These are policies that look at the data before the deployment can proceed.

780
00:34:52,080 --> 00:34:56,460
As your monitor checks can query your observability data to see if the error rate is too high or if

781
00:34:56,460 --> 00:34:57,920
latency is spiking.

782
00:34:57,920 --> 00:35:00,800
If the service looks unhealthy and staging, the check fails.

783
00:35:00,800 --> 00:35:01,800
The gate stays closed.

784
00:35:01,800 --> 00:35:04,200
You don't go to production if staging is on fire.

785
00:35:04,200 --> 00:35:06,440
Policy as code checks also verify compliance.

786
00:35:06,440 --> 00:35:10,040
They look to see if the required security scans are present in the artifacts and ensure

787
00:35:10,040 --> 00:35:11,760
the solution meets company standards.

788
00:35:11,760 --> 00:35:15,240
These checks stop any deployment that would violate your governance rules before it even

789
00:35:15,240 --> 00:35:16,240
starts.

790
00:35:16,240 --> 00:35:18,040
Tracability is also baked in.

791
00:35:18,040 --> 00:35:21,360
Every deployment must be linked to a work item, whether that's a feature request or a bug

792
00:35:21,360 --> 00:35:22,360
fix.

793
00:35:22,360 --> 00:35:25,040
That work item is the source of truth for why the change exists.

794
00:35:25,040 --> 00:35:29,760
If someone asks six months from now why a specific version was deployed, the audit trail is right

795
00:35:29,760 --> 00:35:30,760
there.

796
00:35:30,760 --> 00:35:31,760
You don't have to guess.

797
00:35:31,760 --> 00:35:34,040
The Yannel pipeline itself handles the sequencing.

798
00:35:34,040 --> 00:35:37,160
You have stages for build, test and each of the rings.

799
00:35:37,160 --> 00:35:40,320
Ring one only runs if ring zero was successful.

800
00:35:40,320 --> 00:35:42,480
Ring two only runs if ring one is healthy.

801
00:35:42,480 --> 00:35:43,480
These aren't suggestions.

802
00:35:43,480 --> 00:35:45,480
They are logic statements in the code.

803
00:35:45,480 --> 00:35:48,240
If the condition isn't met, the stage does not run.

804
00:35:48,240 --> 00:35:49,240
Period.

805
00:35:49,240 --> 00:35:53,480
Service connections are scoped to each environment and the permissions get tighter as you move

806
00:35:53,480 --> 00:35:54,680
toward production.

807
00:35:54,680 --> 00:35:58,960
The dev canary connection might have broad permissions to create or delete resources because

808
00:35:58,960 --> 00:36:00,720
the team needs to experiment.

809
00:36:00,720 --> 00:36:04,240
But staging pilot has fewer permissions and production has the bare minimum.

810
00:36:04,240 --> 00:36:07,880
It can deploy artifacts and update config but it can't just create new infrastructure

811
00:36:07,880 --> 00:36:08,880
on its own.

812
00:36:08,880 --> 00:36:10,880
The setup prevents massive accidents.

813
00:36:10,880 --> 00:36:14,840
A team playing around in dev canary can't accidentally delete a production database because

814
00:36:14,840 --> 00:36:18,040
their service connection literally doesn't have the permission to touch it.

815
00:36:18,040 --> 00:36:20,520
The separation is enforced at the identity level.

816
00:36:20,520 --> 00:36:23,920
Even if someone tried to hard code a production reference into their script, the system would

817
00:36:23,920 --> 00:36:24,920
block it.

818
00:36:24,920 --> 00:36:28,320
Finally the audit logs capture every single move you know who approved the release, when

819
00:36:28,320 --> 00:36:31,120
they did it and which commit hash was involved.

820
00:36:31,120 --> 00:36:33,360
Everything is connected and everything is searchable.

821
00:36:33,360 --> 00:36:37,120
If you're dealing with an incident or an audit, you can reconstruct the entire history

822
00:36:37,120 --> 00:36:38,880
of the system in minutes.

823
00:36:38,880 --> 00:36:42,080
GitHub actions does the same thing but the mechanics look a little different.

824
00:36:42,080 --> 00:36:43,760
GitHub actions ring implementation.

825
00:36:43,760 --> 00:36:46,080
GitHub takes a different path to reach the same goal.

826
00:36:46,080 --> 00:36:49,880
As your dev ops gives you environments as a built-in feature but GitHub is different,

827
00:36:49,880 --> 00:36:53,200
you have to build your rings yourself by combining environments, protection rules

828
00:36:53,200 --> 00:36:54,840
and reusable workflows.

829
00:36:54,840 --> 00:36:58,880
The outcome is the same but the mechanics, they're distinct.

830
00:36:58,880 --> 00:37:02,280
Rings and GitHub map directly to the environments in your repository settings.

831
00:37:02,280 --> 00:37:03,440
You create three of them.

832
00:37:03,440 --> 00:37:06,280
Ring zero, ring one and ring two.

833
00:37:06,280 --> 00:37:09,800
Each one is a deployment target with its own set of protection rules.

834
00:37:09,800 --> 00:37:11,800
Ring zero might have almost no protection.

835
00:37:11,800 --> 00:37:15,080
Ring one requires specific people to review and approve the promotion.

836
00:37:15,080 --> 00:37:16,760
But ring two, that's where things change.

837
00:37:16,760 --> 00:37:19,960
You might require multiple reviewers plus a mandatory wait timer.

838
00:37:19,960 --> 00:37:22,520
Even if someone is standing there ready to click yes.

839
00:37:22,520 --> 00:37:26,160
The deployment to production cannot happen immediately after staging is approved.

840
00:37:26,160 --> 00:37:28,400
Branch restrictions add another layer of control.

841
00:37:28,400 --> 00:37:31,480
You can set it up so ring two deployments only happen from your main branch or specific

842
00:37:31,480 --> 00:37:32,480
release branches.

843
00:37:32,480 --> 00:37:35,640
This means you can't deploy to production from a feature branch because the system prevents

844
00:37:35,640 --> 00:37:36,640
it structurally.

845
00:37:36,640 --> 00:37:39,200
The real power starts with reusable workflows.

846
00:37:39,200 --> 00:37:43,840
Instead of every team writing their own deployment logic, the platform team maintains a single

847
00:37:43,840 --> 00:37:45,600
deploy to ring workflow.

848
00:37:45,600 --> 00:37:46,600
It's simple.

849
00:37:46,600 --> 00:37:50,000
The input is just the service name, the version and the target ring.

850
00:37:50,000 --> 00:37:51,640
The workflow handles the rest.

851
00:37:51,640 --> 00:37:52,640
It checks out the code.

852
00:37:52,640 --> 00:37:53,880
It validates the deployment.

853
00:37:53,880 --> 00:37:56,680
It runs pre-flight checks and records the results.

854
00:37:56,680 --> 00:37:59,360
Everything is standardized and everything is consistent.

855
00:37:59,360 --> 00:38:02,720
The orchestration workflow that the product teams actually use is straightforward.

856
00:38:02,720 --> 00:38:06,640
Build, test, deploy to ring zero, then you wait for the automatic checks to pass.

857
00:38:06,640 --> 00:38:09,640
If they do, you move to ring one and wait for human approval.

858
00:38:09,640 --> 00:38:11,360
If that succeeds, you move to ring two.

859
00:38:11,360 --> 00:38:15,000
The workflow defines the sequence once and every team follows it.

860
00:38:15,000 --> 00:38:17,400
Environment secrets are scope to specific rings.

861
00:38:17,400 --> 00:38:21,600
Your database credentials for ring zero are different from ring two and your API keys

862
00:38:21,600 --> 00:38:22,600
are different as well.

863
00:38:22,600 --> 00:38:26,800
The workflow can just ask for a secret called database URL, which one it actually gets depends

864
00:38:26,800 --> 00:38:28,240
on where the job is running.

865
00:38:28,240 --> 00:38:32,800
Ring zero jobs get the ring zero database, while ring two jobs get the production database.

866
00:38:32,800 --> 00:38:37,120
This stops credentials from leaking and forces a clean separation without making teams manage

867
00:38:37,120 --> 00:38:39,040
a dozen different secret names.

868
00:38:39,040 --> 00:38:42,960
You can also make the workflow talk to external systems before it promotes code.

869
00:38:42,960 --> 00:38:47,280
It can query your monitoring API to check error rates in ring zero over the last hour.

870
00:38:47,280 --> 00:38:50,760
If those rates are low, the promotion to ring one is approved automatically.

871
00:38:50,760 --> 00:38:53,120
If errors are up, the promotion is blocked.

872
00:38:53,120 --> 00:38:58,200
You've moved from a human staring at a dashboard to an automated system using objective data.

873
00:38:58,200 --> 00:39:02,080
The decision is faster, it's consistent, and it's all documented in the logs.

874
00:39:02,080 --> 00:39:03,680
But this flexibility has a trade-off.

875
00:39:03,680 --> 00:39:06,160
GitHub actions doesn't enforce these guardrails for you.

876
00:39:06,160 --> 00:39:07,160
You have to build them.

877
00:39:07,160 --> 00:39:10,080
You have to write the workflow and configure the protections yourself.

878
00:39:10,080 --> 00:39:14,120
If you don't, there is nothing stopping a team from deploying straight to production on

879
00:39:14,120 --> 00:39:15,120
a Friday afternoon.

880
00:39:15,120 --> 00:39:17,720
That's why the platform team is so important here.

881
00:39:17,720 --> 00:39:20,720
They write the reusable workflow and they maintain it.

882
00:39:20,720 --> 00:39:22,720
The product teams use it without changing it.

883
00:39:22,720 --> 00:39:26,360
They can't skip security checks because they don't have the permission to modify the code.

884
00:39:26,360 --> 00:39:29,440
The platform team owns it and every change goes through their review.

885
00:39:29,440 --> 00:39:32,560
This open structure also lets you build more complex logic.

886
00:39:32,560 --> 00:39:34,920
That monitoring query can be as advanced as you want.

887
00:39:34,920 --> 00:39:38,840
You could compare current metrics against historical baselines or even use machine learning

888
00:39:38,840 --> 00:39:41,720
to predict issues as your DevOps can do this too.

889
00:39:41,720 --> 00:39:44,920
But GitHub makes it feel more natural to extend.

890
00:39:44,920 --> 00:39:47,960
The platform you choose doesn't change the core operating model.

891
00:39:47,960 --> 00:39:49,840
You have to standardize the mechanics.

892
00:39:49,840 --> 00:39:52,240
So the outcomes and automate the promotion.

893
00:39:52,240 --> 00:39:53,240
The tools are different.

894
00:39:53,240 --> 00:39:55,040
The principle is the same.

895
00:39:55,040 --> 00:39:59,160
Canary releases, the advanced pattern, ring deployments are staged, rollouts to specific

896
00:39:59,160 --> 00:40:00,480
user cohorts.

897
00:40:00,480 --> 00:40:02,200
Canary releases take that same model.

898
00:40:02,200 --> 00:40:04,840
But they apply it to traffic percentages instead of user groups.

899
00:40:04,840 --> 00:40:07,720
The distinction is subtle, but the consequences are massive.

900
00:40:07,720 --> 00:40:10,160
In a ring deployment, you define your populations.

901
00:40:10,160 --> 00:40:11,640
Ring zero is your internal team.

902
00:40:11,640 --> 00:40:13,360
Ring one is a pilot business unit.

903
00:40:13,360 --> 00:40:14,440
Ring two is everyone else.

904
00:40:14,440 --> 00:40:15,680
You move in a sequence.

905
00:40:15,680 --> 00:40:17,000
You monitor the health of ring zero.

906
00:40:17,000 --> 00:40:19,480
And if the data looks good, you deploy to ring one.

907
00:40:19,480 --> 00:40:20,600
Then you monitor ring one.

908
00:40:20,600 --> 00:40:23,440
And if that stays stable, you finally open the gates for ring two.

909
00:40:23,440 --> 00:40:25,560
A canary release in words that unit.

910
00:40:25,560 --> 00:40:30,840
Instead of saying, this specific group gets the new version, you say 1% of all traffic gets

911
00:40:30,840 --> 00:40:32,400
the new version.

912
00:40:32,400 --> 00:40:35,840
Everyone stays in the same environment, the same infrastructure, the same database, the

913
00:40:35,840 --> 00:40:37,120
same backend.

914
00:40:37,120 --> 00:40:42,600
But the traffic router directs 1% of requests to the new version, while 99% stay on the stable

915
00:40:42,600 --> 00:40:43,600
version.

916
00:40:43,600 --> 00:40:47,400
You can do this through a service mesh, feature flags or load balancer logic.

917
00:40:47,400 --> 00:40:49,440
The how matters less than the what.

918
00:40:49,440 --> 00:40:54,360
You monitor that 1% with extreme focus error rates, latency, business KPIs.

919
00:40:54,360 --> 00:40:56,400
This data comes from real production traffic.

920
00:40:56,400 --> 00:40:57,440
It isn't a simulation.

921
00:40:57,440 --> 00:41:02,200
These are real users, real load patterns and real edge cases that your test environment

922
00:41:02,200 --> 00:41:03,200
never saw.

923
00:41:03,200 --> 00:41:06,800
If the canary looks healthy after a few hours or days, you turn the dial.

924
00:41:06,800 --> 00:41:11,120
5%, 10%, 25%, eventually you hit 100.

925
00:41:11,120 --> 00:41:14,360
But if the metrics degraded any point, the traffic router snaps back instantly.

926
00:41:14,360 --> 00:41:16,480
1% goes back to the stable version.

927
00:41:16,480 --> 00:41:18,160
Everyone else stays on the stable version.

928
00:41:18,160 --> 00:41:19,440
The incident is bounded.

929
00:41:19,440 --> 00:41:21,960
The power here is that you aren't making a binary choice.

930
00:41:21,960 --> 00:41:25,320
It isn't everyone stays or everyone goes.

931
00:41:25,320 --> 00:41:28,280
You are incrementally proving that the new version is safe.

932
00:41:28,280 --> 00:41:30,120
Each percentage increases evidence.

933
00:41:30,120 --> 00:41:33,360
The more traffic flows without an incident, the more confident you become.

934
00:41:33,360 --> 00:41:36,680
The metrics that matter in a canary are different because the risk is different.

935
00:41:36,680 --> 00:41:38,640
You aren't comparing ring zero to ring one.

936
00:41:38,640 --> 00:41:42,040
You are comparing the new version directly against the stable version.

937
00:41:42,040 --> 00:41:43,320
Error rate is your first signal.

938
00:41:43,320 --> 00:41:47,560
If the new version's error rate climbs even 5% higher than the stable version, something

939
00:41:47,560 --> 00:41:48,080
is wrong.

940
00:41:48,080 --> 00:41:50,760
You roll back, you investigate, you fix it.

941
00:41:50,760 --> 00:41:53,120
Latency is just as important, but it's more nuanced.

942
00:41:53,120 --> 00:41:56,400
Some slowdown might be okay if it's in a path that doesn't matter.

943
00:41:56,400 --> 00:41:59,640
If a reporting query runs 5% slower, nobody cares.

944
00:41:59,640 --> 00:42:03,080
But if your payment flow latency jumps 10%, your customers will notice.

945
00:42:03,080 --> 00:42:04,800
The threshold depends on the change.

946
00:42:04,800 --> 00:42:07,640
A refactor in the request path should have zero latency regression.

947
00:42:07,640 --> 00:42:12,000
A new feature that adds processing might be tolerated up to 10% for basic flows.

948
00:42:12,000 --> 00:42:13,840
But it should be zero for critical ones.

949
00:42:13,840 --> 00:42:16,200
Business KPIs matter more here than in a ring.

950
00:42:16,200 --> 00:42:19,000
The canary is testing against your production baseline.

951
00:42:19,000 --> 00:42:23,920
If you ship a new checkout flow and the conversion rate drops 3% in the canary, that is your data.

952
00:42:23,920 --> 00:42:26,640
Maybe the flow is better, but users need time to learn it.

953
00:42:26,640 --> 00:42:28,800
Or maybe the flow is just confusing.

954
00:42:28,800 --> 00:42:31,080
Canary metrics tell you which one it is.

955
00:42:31,080 --> 00:42:34,080
Implementing this requires infrastructure most enterprises just don't have yet.

956
00:42:34,080 --> 00:42:35,320
You need traffic splitting.

957
00:42:35,320 --> 00:42:39,080
A service mesh like Istio can do this, a platform like LaunchDarkly can do this.

958
00:42:39,080 --> 00:42:41,240
A sophisticated load balancer can do this.

959
00:42:41,240 --> 00:42:42,760
But you can't do it without them.

960
00:42:42,760 --> 00:42:46,080
You have to root traffic algorithmically, not manually.

961
00:42:46,080 --> 00:42:48,320
You also need automated health evaluation.

962
00:42:48,320 --> 00:42:50,400
The monitoring system shouldn't just collect data.

963
00:42:50,400 --> 00:42:51,400
It should evaluate it.

964
00:42:51,400 --> 00:42:54,160
The system compares the canary against the stable version.

965
00:42:54,160 --> 00:42:58,640
If it sees degradation beyond your limit, it reverses the traffic split automatically.

966
00:42:58,640 --> 00:42:59,880
No human intervention.

967
00:42:59,880 --> 00:43:01,080
The rollback is instant.

968
00:43:01,080 --> 00:43:04,760
The blast radius was already bounded by that tiny percentage of traffic.

969
00:43:04,760 --> 00:43:09,560
Research shows that canary deployments reduce incident blast radius by 80 to 90%.

970
00:43:09,560 --> 00:43:10,840
Compare that to a big bang release.

971
00:43:10,840 --> 00:43:13,120
In a big bang release, you ship to everyone at once.

972
00:43:13,120 --> 00:43:14,120
Something breaks.

973
00:43:14,120 --> 00:43:15,120
Everyone is affected.

974
00:43:15,120 --> 00:43:18,960
Thousands of thousands of users are stuck and the organization goes into crisis mode.

975
00:43:18,960 --> 00:43:22,000
In a canary release, you ship to 1%.

976
00:43:22,000 --> 00:43:23,000
Something breaks.

977
00:43:23,000 --> 00:43:24,240
1% of users are affected.

978
00:43:24,240 --> 00:43:27,720
You notice immediately you roll back and the incident is over before most people even

979
00:43:27,720 --> 00:43:28,720
knew it started.

980
00:43:28,720 --> 00:43:30,720
Canary is work best for high-risk changes.

981
00:43:30,720 --> 00:43:34,760
A new algorithm, a massive refactor, a feature touching a critical path.

982
00:43:34,760 --> 00:43:36,560
These are the moments where the risk is real.

983
00:43:36,560 --> 00:43:41,160
For low-risk changes like config updates or cosmetic tweaks, a ring deployment is simpler.

984
00:43:41,160 --> 00:43:42,160
And it's usually enough.

985
00:43:42,160 --> 00:43:45,560
The difference between a ring and a canary is the difference between regional risk and

986
00:43:45,560 --> 00:43:47,040
proportional risk.

987
00:43:47,040 --> 00:43:49,200
Rings compartmentalized by population.

988
00:43:49,200 --> 00:43:51,040
Canaries compartmentalized by percentage.

989
00:43:51,040 --> 00:43:52,880
Both reduce the blast radius.

990
00:43:52,880 --> 00:43:55,280
Canaries are just more granular.

991
00:43:55,280 --> 00:43:56,720
Observability has a governance requirement.

992
00:43:56,720 --> 00:44:00,680
The canary pattern I just described, the splitting, the evaluation, the rollback.

993
00:44:00,680 --> 00:44:03,280
It can't exist without a specific foundation.

994
00:44:03,280 --> 00:44:04,920
Most enterprises treat this as optional.

995
00:44:04,920 --> 00:44:06,560
It's called observability.

996
00:44:06,560 --> 00:44:08,680
This is not the same thing as monitoring.

997
00:44:08,680 --> 00:44:11,680
Monitoring is what you have when you collect data into dashboards.

998
00:44:11,680 --> 00:44:15,560
You have alerts, you get a notification when error rates spike.

999
00:44:15,560 --> 00:44:17,440
Someone logs in and looks at a chart.

1000
00:44:17,440 --> 00:44:18,640
Monitoring is reactive visibility.

1001
00:44:18,640 --> 00:44:21,320
You know something broke because an alert fired.

1002
00:44:21,320 --> 00:44:22,800
Observability is different.

1003
00:44:22,800 --> 00:44:26,400
It's the ability to ask questions about your system that you didn't anticipate when you

1004
00:44:26,400 --> 00:44:28,000
built it.

1005
00:44:28,000 --> 00:44:31,520
Distributed tracing lets you follow one single request through your entire stack.

1006
00:44:31,520 --> 00:44:32,520
Which service is touched it?

1007
00:44:32,520 --> 00:44:33,520
Where did it spend time?

1008
00:44:33,520 --> 00:44:34,520
Where did it fail?

1009
00:44:34,520 --> 00:44:36,520
You can answer these questions in real time.

1010
00:44:36,520 --> 00:44:40,480
Even if you didn't set up specific instruments for them in advance, metrics collection captures

1011
00:44:40,480 --> 00:44:41,480
the signals.

1012
00:44:41,480 --> 00:44:44,640
Distributes, latency percentiles, resource usage.

1013
00:44:44,640 --> 00:44:48,480
Structured logging lets you search millions of entries to find the five that actually matter.

1014
00:44:48,480 --> 00:44:52,040
This difference is vital because governance at scale depends on measurement.

1015
00:44:52,040 --> 00:44:54,840
You can't even run ring deployments without observability.

1016
00:44:54,840 --> 00:44:57,440
Ring zero is healthy, but how do you actually know?

1017
00:44:57,440 --> 00:45:00,960
You need to query the error rate for ring zero over the last four hours and get an answer

1018
00:45:00,960 --> 00:45:01,960
in seconds.

1019
00:45:01,960 --> 00:45:04,440
Then you query ring one for the same window and compare them.

1020
00:45:04,440 --> 00:45:06,240
That comparison is your decision point.

1021
00:45:06,240 --> 00:45:09,920
If ring zero has zero point five percent errors and ring one has one percent, something

1022
00:45:09,920 --> 00:45:10,920
changed.

1023
00:45:10,920 --> 00:45:14,240
You have to understand why before you promote anything to ring two.

1024
00:45:14,240 --> 00:45:15,880
Canary releases demand this in real time.

1025
00:45:15,880 --> 00:45:18,360
You've sent one percent of traffic to a new version.

1026
00:45:18,360 --> 00:45:21,840
Your system needs to evaluate that one percent against the ninety nine percent baseline

1027
00:45:21,840 --> 00:45:22,840
constantly.

1028
00:45:22,840 --> 00:45:27,320
Not once an hour, not once a day, near real time, second by second.

1029
00:45:27,320 --> 00:45:31,000
If the canary's error rate jumps from zero point three percent to two percent in five

1030
00:45:31,000 --> 00:45:33,080
minutes, the system has to catch it.

1031
00:45:33,080 --> 00:45:36,120
It has to evaluate the threshold and trigger a rollback in seconds.

1032
00:45:36,120 --> 00:45:39,120
You can't do that if you're waiting for someone to log in and check a dashboard.

1033
00:45:39,120 --> 00:45:40,880
The detection has to be automated.

1034
00:45:40,880 --> 00:45:43,040
The evaluation has to be programmatic.

1035
00:45:43,040 --> 00:45:44,400
Most enterprises have monitoring.

1036
00:45:44,400 --> 00:45:48,240
They have the dashboards and the alerts, but they don't have the infrastructure to ask arbitrary

1037
00:45:48,240 --> 00:45:49,720
questions at high speed.

1038
00:45:49,720 --> 00:45:52,040
They have snapshots of a moment in time.

1039
00:45:52,040 --> 00:45:54,880
Observability requires continuous, queryable streams of data.

1040
00:45:54,880 --> 00:45:59,320
It lets you reconstruct behavior and correlate signals across the whole stack.

1041
00:45:59,320 --> 00:46:01,000
Building this is a platform investment.

1042
00:46:01,000 --> 00:46:03,360
It isn't something individual product teams should own.

1043
00:46:03,360 --> 00:46:07,360
The platform team sets up centralized logging, so data from hundreds of services flows into

1044
00:46:07,360 --> 00:46:08,360
one place.

1045
00:46:08,360 --> 00:46:10,040
They build the metrics infrastructure.

1046
00:46:10,040 --> 00:46:13,320
They implement the tracing, so transactions are visible end to end.

1047
00:46:13,320 --> 00:46:18,440
They create the dashboards per ring, so you can see exactly what each population is experiencing.

1048
00:46:18,440 --> 00:46:22,240
This infrastructure is expensive, the licensing, the storage and the compute costs add up.

1049
00:46:22,240 --> 00:46:25,320
But you pay that cost once and spread it across hundreds of teams.

1050
00:46:25,320 --> 00:46:28,720
If every team tried to build their own stack, the cost would explode.

1051
00:46:28,720 --> 00:46:32,640
And if you have no observability at all, your rings and canaries become unreliable.

1052
00:46:32,640 --> 00:46:35,360
You end up promoting changes based on income data.

1053
00:46:35,360 --> 00:46:39,360
You start trusting that things worked just because nobody complained, rather than because

1054
00:46:39,360 --> 00:46:41,160
you actually verified it.

1055
00:46:41,160 --> 00:46:43,200
Observability is what makes governance trustworthy.

1056
00:46:43,200 --> 00:46:44,200
It is the foundation.

1057
00:46:44,200 --> 00:46:48,120
And once it's in place, you can finally implement the last piece of the model.

1058
00:46:48,120 --> 00:46:49,720
Automated promotion gates.

1059
00:46:49,720 --> 00:46:50,800
Automated promotion gates.

1060
00:46:50,800 --> 00:46:52,640
A promotion gate is a policy.

1061
00:46:52,640 --> 00:46:56,440
It automatically decides if it's safe to move your code from one ring to the next.

1062
00:46:56,440 --> 00:46:59,640
This isn't a human decision point disguised as automation.

1063
00:46:59,640 --> 00:47:00,960
It's an actual algorithmic decision.

1064
00:47:00,960 --> 00:47:04,440
The gate reads the metrics your observability tools have collected.

1065
00:47:04,440 --> 00:47:07,560
It compares them against the thresholds you've defined.

1066
00:47:07,560 --> 00:47:10,920
Then it either opens allowing the promotion to move forward or it closes.

1067
00:47:10,920 --> 00:47:12,560
If it closes, the promotion is blocked.

1068
00:47:12,560 --> 00:47:14,320
The change stays exactly where it is.

1069
00:47:14,320 --> 00:47:15,840
In practice, it looks like this.

1070
00:47:15,840 --> 00:47:17,200
You define a policy.

1071
00:47:17,200 --> 00:47:22,000
If the error rate in ring one is under 1% and latency has increased less than 5% compared

1072
00:47:22,000 --> 00:47:26,840
to the baseline in ring zero, then automatically promote to ring two.

1073
00:47:26,840 --> 00:47:29,320
The gate evaluates this logic continuously.

1074
00:47:29,320 --> 00:47:30,320
Ring one is running.

1075
00:47:30,320 --> 00:47:31,840
Observability is gathering data.

1076
00:47:31,840 --> 00:47:34,440
The automated gate queries those metrics in real time.

1077
00:47:34,440 --> 00:47:40,480
If the error rate is 0.8%, and the latency increase is only 3%, both conditions are met.

1078
00:47:40,480 --> 00:47:41,480
The gate opens.

1079
00:47:41,480 --> 00:47:44,720
The deployment proceeds to ring two without any human involvement at all.

1080
00:47:44,720 --> 00:47:49,160
But if that same gate detects the error rate creeping toward 2%, the logic reverses.

1081
00:47:49,160 --> 00:47:50,600
The condition is no longer met.

1082
00:47:50,600 --> 00:47:51,600
The gate closes.

1083
00:47:51,600 --> 00:47:52,920
Deployment to ring two is blocked.

1084
00:47:52,920 --> 00:47:54,400
The change stays in ring one.

1085
00:47:54,400 --> 00:47:56,560
The team gets a notification that something isn't right.

1086
00:47:56,560 --> 00:47:57,560
You investigate.

1087
00:47:57,560 --> 00:47:58,560
You fix the issue.

1088
00:47:58,560 --> 00:48:02,560
Once the fix is live and the metrics improve, the gate conditions are satisfied again,

1089
00:48:02,560 --> 00:48:04,400
and the promotion proceeds automatically.

1090
00:48:04,400 --> 00:48:06,840
Defining these gates requires precision.

1091
00:48:06,840 --> 00:48:09,560
Policy as code means you have to write the logic explicitly.

1092
00:48:09,560 --> 00:48:14,080
You use YAML or J-SUN to specify exactly which metrics matter, what the acceptable thresholds

1093
00:48:14,080 --> 00:48:16,560
are, and how long the evaluation window should be.

1094
00:48:16,560 --> 00:48:18,240
You aren't writing vague guidelines.

1095
00:48:18,240 --> 00:48:21,280
You're writing logic that a system can execute.

1096
00:48:21,280 --> 00:48:23,520
Error rate should be low is a guideline.

1097
00:48:23,520 --> 00:48:28,320
Error rate must be less than 1% evaluated over the last hour, with at least 1000 requests

1098
00:48:28,320 --> 00:48:30,040
sampled is a policy.

1099
00:48:30,040 --> 00:48:32,120
This precision serves two main purposes.

1100
00:48:32,120 --> 00:48:34,120
First, it makes the process reproducible.

1101
00:48:34,120 --> 00:48:38,480
The same metrics checked against the same thresholds will always produce the same outcome.

1102
00:48:38,480 --> 00:48:39,800
There is no ambiguity.

1103
00:48:39,800 --> 00:48:43,440
There is no, well, it looked okay to me from a tired engineer.

1104
00:48:43,440 --> 00:48:45,360
The gate either passes or it fails.

1105
00:48:45,360 --> 00:48:47,200
Second, it makes the process auditable.

1106
00:48:47,200 --> 00:48:50,880
You can look back and ask exactly why a change moved from ring one to ring two.

1107
00:48:50,880 --> 00:48:52,680
The answer is right there in the code.

1108
00:48:52,680 --> 00:48:53,680
These were the conditions.

1109
00:48:53,680 --> 00:48:55,560
These were the metric values at the time.

1110
00:48:55,560 --> 00:48:57,680
The conditions were met, so the promotion happened.

1111
00:48:57,680 --> 00:48:58,920
The decision is documented.

1112
00:48:58,920 --> 00:49:00,480
The logic is transparent.

1113
00:49:00,480 --> 00:49:03,920
The logic gates don't remove human judgment from deployment decisions.

1114
00:49:03,920 --> 00:49:06,680
They remove human judgment from routine approvals.

1115
00:49:06,680 --> 00:49:11,120
If you've defined a gate that says "promote" when the error rate is below 1%, then once

1116
00:49:11,120 --> 00:49:13,080
that happens, the gate should just open.

1117
00:49:13,080 --> 00:49:16,840
You don't need a person to log in, check a dashboard and manually click a button.

1118
00:49:16,840 --> 00:49:19,320
That is work that shouldn't require human attention.

1119
00:49:19,320 --> 00:49:21,040
Complex decisions still go to humans.

1120
00:49:21,040 --> 00:49:24,600
If the metrics are technically in bounds, but you're worried about roll out timing because

1121
00:49:24,600 --> 00:49:26,960
of a holiday weekend, that's a human call.

1122
00:49:26,960 --> 00:49:28,680
The gate can't make it.

1123
00:49:28,680 --> 00:49:30,440
Automated gates handle the routine.

1124
00:49:30,440 --> 00:49:31,720
Humans handle the exceptions.

1125
00:49:31,720 --> 00:49:33,800
The result is velocity without sacrifice.

1126
00:49:33,800 --> 00:49:36,800
In traditional governance, promotions require manual approval.

1127
00:49:36,800 --> 00:49:38,280
Someone has to review the data.

1128
00:49:38,280 --> 00:49:39,360
Someone has to make the call.

1129
00:49:39,360 --> 00:49:41,080
That takes time.

1130
00:49:41,080 --> 00:49:44,280
Even in organizations where people are fast, there is still latency.

1131
00:49:44,280 --> 00:49:46,200
An engineer finishes work in ring zero.

1132
00:49:46,200 --> 00:49:47,680
They submit for promotion to ring one.

1133
00:49:47,680 --> 00:49:50,240
They wait an hour or maybe a day for an approval.

1134
00:49:50,240 --> 00:49:52,120
Then they wait all over again for ring two.

1135
00:49:52,120 --> 00:49:55,760
With automated gates, the moment the conditions are satisfied, the promotion happens.

1136
00:49:55,760 --> 00:49:56,760
There is no waiting.

1137
00:49:56,760 --> 00:49:58,240
There is no human bottleneck.

1138
00:49:58,240 --> 00:50:02,520
The deployment flows to the next ring immediately, but this is the critical part.

1139
00:50:02,520 --> 00:50:05,080
The deployment only flows if the metrics justify it.

1140
00:50:05,080 --> 00:50:06,280
The safety criteria are met.

1141
00:50:06,280 --> 00:50:08,080
The blast radius is controlled.

1142
00:50:08,080 --> 00:50:09,280
Automation doesn't mean being reckless.

1143
00:50:09,280 --> 00:50:13,120
It means reducing friction for safe decisions while keeping safety non-negotiable.

1144
00:50:13,120 --> 00:50:14,440
This is where the model shifts.

1145
00:50:14,440 --> 00:50:19,040
It stops feeling like an orchestrated release and starts feeling like continuous delivery.

1146
00:50:19,040 --> 00:50:21,000
The difference between CD and CD.

1147
00:50:21,000 --> 00:50:24,360
There are two phrases that sound identical but mean completely different things.

1148
00:50:24,360 --> 00:50:28,080
The distinction between them determines whether your organization is delivering fast or

1149
00:50:28,080 --> 00:50:29,560
just delivering.

1150
00:50:29,560 --> 00:50:33,240
Continuous delivery is an automated pipeline that can release to production at any time,

1151
00:50:33,240 --> 00:50:35,640
but it requires a human to actually make the final decision.

1152
00:50:35,640 --> 00:50:36,640
The pipeline is ready.

1153
00:50:36,640 --> 00:50:37,640
The code is tested.

1154
00:50:37,640 --> 00:50:39,200
The deployment is validated.

1155
00:50:39,200 --> 00:50:40,200
Everything is prepared.

1156
00:50:40,200 --> 00:50:45,120
But the final step, shipping it to users, requires a person to click a button and say yes.

1157
00:50:45,120 --> 00:50:49,160
Continuous deployment is an automated pipeline that releases to production without any human

1158
00:50:49,160 --> 00:50:51,400
approval as long as the metrics are green.

1159
00:50:51,400 --> 00:50:55,200
The moment the data indicates safety, the deployment flows to production automatically.

1160
00:50:55,200 --> 00:50:56,560
There is no human decision point.

1161
00:50:56,560 --> 00:50:58,040
There is no approval gate.

1162
00:50:58,040 --> 00:51:00,280
The decision is entirely algorithmic.

1163
00:51:00,280 --> 00:51:04,160
Most enterprises should aim for continuous delivery, not continuous deployment.

1164
00:51:04,160 --> 00:51:08,680
This distinction matters because it shapes how you think about governance, speed and risk.

1165
00:51:08,680 --> 00:51:12,200
Continuous delivery gives you the speed benefits without removing human judgment from the

1166
00:51:12,200 --> 00:51:14,120
biggest release decisions.

1167
00:51:14,120 --> 00:51:16,520
Your deployments happen in minutes instead of days.

1168
00:51:16,520 --> 00:51:19,960
The moment ring one metrics are healthy, the change is ready for ring two.

1169
00:51:19,960 --> 00:51:21,360
But ring two is production.

1170
00:51:21,360 --> 00:51:23,840
These are real users and real business impacts.

1171
00:51:23,840 --> 00:51:25,560
A human decides when to pull the trigger.

1172
00:51:25,560 --> 00:51:28,160
Maybe they wait until the morning when the business is less critical.

1173
00:51:28,160 --> 00:51:30,600
Maybe they coordinate with marketing for a specific launch.

1174
00:51:30,600 --> 00:51:33,800
They might want to ensure backup systems are running or they might just want one last

1175
00:51:33,800 --> 00:51:35,360
moment to review the change.

1176
00:51:35,360 --> 00:51:37,040
This human judgment is valuable.

1177
00:51:37,040 --> 00:51:39,400
Not every decision can be turned into a metric.

1178
00:51:39,400 --> 00:51:41,200
Business context matters.

1179
00:51:41,200 --> 00:51:42,720
Competitive context matters.

1180
00:51:42,720 --> 00:51:45,240
The existence of unknown unknowns matters.

1181
00:51:45,240 --> 00:51:49,280
An engineer who has been doing this for 15 years has a type of patent recognition that

1182
00:51:49,280 --> 00:51:50,960
no algorithm can capture.

1183
00:51:50,960 --> 00:51:52,600
That judgment is worth keeping.

1184
00:51:52,600 --> 00:51:54,080
Continuous delivery preserves it.

1185
00:51:54,080 --> 00:51:56,840
This deployment removes that human approval gate entirely.

1186
00:51:56,840 --> 00:52:00,080
If the error rate is below the threshold and latency is acceptable, the deployment just

1187
00:52:00,080 --> 00:52:01,080
happens.

1188
00:52:01,080 --> 00:52:02,160
There is no waiting for an approval.

1189
00:52:02,160 --> 00:52:06,360
The promotion gate evaluates the metrics, sees they are green, and the change reaches production

1190
00:52:06,360 --> 00:52:07,800
without anyone stepping in.

1191
00:52:07,800 --> 00:52:11,660
This only works if your observability is perfect and your rollback procedures are bullet

1192
00:52:11,660 --> 00:52:13,640
proof, you have to measure everything.

1193
00:52:13,640 --> 00:52:16,840
You need the ability to detect problems within seconds, not minutes.

1194
00:52:16,840 --> 00:52:20,720
You need absolute confidence that a rollback will work every single time.

1195
00:52:20,720 --> 00:52:22,120
Most organizations aren't there yet.

1196
00:52:22,120 --> 00:52:23,800
Their observability is incomplete.

1197
00:52:23,800 --> 00:52:26,240
The rollback procedures have failure modes.

1198
00:52:26,240 --> 00:52:29,120
They have dependencies that make rollbacks complicated.

1199
00:52:29,120 --> 00:52:32,160
For those organizations, continuous deployment is dangerous.

1200
00:52:32,160 --> 00:52:34,920
Ring deployments enable continuous delivery naturally.

1201
00:52:34,920 --> 00:52:37,040
Ring zero and ring one are fully automated.

1202
00:52:37,040 --> 00:52:38,600
No human approval is required.

1203
00:52:38,600 --> 00:52:42,920
The metrics are evaluated, the gates open, and the deployment proceeds without any friction.

1204
00:52:42,920 --> 00:52:44,160
The momentum builds.

1205
00:52:44,160 --> 00:52:45,200
But ring two is different.

1206
00:52:45,200 --> 00:52:46,200
Ring two is production.

1207
00:52:46,200 --> 00:52:47,800
Ring two requires a human click.

1208
00:52:47,800 --> 00:52:49,920
The human approval for ring two is lightweight.

1209
00:52:49,920 --> 00:52:51,200
The metrics are already green.

1210
00:52:51,200 --> 00:52:53,600
The change has already been validated in ring zero.

1211
00:52:53,600 --> 00:52:56,960
It has been running in ring one for hours, or maybe even days.

1212
00:52:56,960 --> 00:52:59,240
All the health indicators point to safety.

1213
00:52:59,240 --> 00:53:01,920
The approval isn't a deep, comprehensive review.

1214
00:53:01,920 --> 00:53:03,840
It's not, should we ship this?

1215
00:53:03,840 --> 00:53:05,680
It's, are you ready to ship this?

1216
00:53:05,680 --> 00:53:07,680
The decision is essentially already made.

1217
00:53:07,680 --> 00:53:10,760
The approval is just a gate that acknowledges responsibility.

1218
00:53:10,760 --> 00:53:11,760
Someone is signing off.

1219
00:53:11,760 --> 00:53:13,080
Someone is saying, I reviewed the metrics.

1220
00:53:13,080 --> 00:53:15,880
I saw the change and I'm confident enough to release it.

1221
00:53:15,880 --> 00:53:18,960
This model speeds up your deployment velocity while keeping the human element where

1222
00:53:18,960 --> 00:53:20,320
it actually matters.

1223
00:53:20,320 --> 00:53:21,800
Teams get speed.

1224
00:53:21,800 --> 00:53:23,400
Organizations get judgment.

1225
00:53:23,400 --> 00:53:26,840
The release doesn't wait for bureaucracy, but it doesn't proceed blindly either.

1226
00:53:26,840 --> 00:53:29,160
This becomes even more important as you scale.

1227
00:53:29,160 --> 00:53:33,080
The next level of complexity happens when you move beyond traditional services and into

1228
00:53:33,080 --> 00:53:38,400
the hybrid world where low-code platforms and traditional development start to converge.

1229
00:53:38,400 --> 00:53:40,680
Power platform, ALM and ring deployments.

1230
00:53:40,680 --> 00:53:42,800
The governance principles we just covered.

1231
00:53:42,800 --> 00:53:45,920
Staged rollouts, metrics and template controls apply to everything.

1232
00:53:45,920 --> 00:53:49,280
But when you move to power platform, the mechanics change because the platform itself

1233
00:53:49,280 --> 00:53:50,400
is built differently.

1234
00:53:50,400 --> 00:53:51,400
This is low-code.

1235
00:53:51,400 --> 00:53:53,160
You aren't deploying compiled binaries.

1236
00:53:53,160 --> 00:53:54,240
You're deploying solutions.

1237
00:53:54,240 --> 00:53:58,360
Your targets are environments connected to dataverse and your developers range from seasoned

1238
00:53:58,360 --> 00:54:01,880
pros to business users who have never written a line of code.

1239
00:54:01,880 --> 00:54:04,360
Your governance model has to account for that reality.

1240
00:54:04,360 --> 00:54:09,000
A power platform ring deployment looks like a traditional CIS or CD ring, but it operates

1241
00:54:09,000 --> 00:54:10,280
on different assumptions.

1242
00:54:10,280 --> 00:54:14,280
You still progress from dev to test, then UAT and finally production.

1243
00:54:14,280 --> 00:54:18,600
Each stage is a distinct power platform environment with its own database.

1244
00:54:18,600 --> 00:54:20,720
The flow isn't code moving through a pipeline.

1245
00:54:20,720 --> 00:54:25,200
Its solutions, packaged containers of apps, flows, tables and configurations, moving from

1246
00:54:25,200 --> 00:54:26,760
one environment to the next.

1247
00:54:26,760 --> 00:54:29,960
The distinction between unmanaged and managed solutions is critical here.

1248
00:54:29,960 --> 00:54:34,200
In dev, solutions are unmanaged, which means developers work directly in the environment

1249
00:54:34,200 --> 00:54:36,120
to build apps and modify tables.

1250
00:54:36,120 --> 00:54:40,040
The solution acts as a container, but you can change components directly without using

1251
00:54:40,040 --> 00:54:41,440
the solution interface.

1252
00:54:41,440 --> 00:54:46,000
This flexibility is intentional because dev is an exploration space where you need speed.

1253
00:54:46,000 --> 00:54:50,000
Once you move to test, UAT or production, solutions become managed.

1254
00:54:50,000 --> 00:54:51,640
A managed solution is read only.

1255
00:54:51,640 --> 00:54:55,080
If something needs to change, you don't touch it in the target environment.

1256
00:54:55,080 --> 00:54:57,960
You go back to dev, update the solution and re-import it.

1257
00:54:57,960 --> 00:55:02,840
This constraint is intentional because it enforces discipline and ensures every change is reviewed

1258
00:55:02,840 --> 00:55:04,960
before it ever reaches production.

1259
00:55:04,960 --> 00:55:08,320
Governance and power platform works differently because the control points are different.

1260
00:55:08,320 --> 00:55:11,560
Instead of branch policies, you use data loss prevention policies.

1261
00:55:11,560 --> 00:55:14,440
DLP policies control which connectors can talk to each other.

1262
00:55:14,440 --> 00:55:18,280
You can decide that business connectors, the one-stouching sensitive data, can't be used

1263
00:55:18,280 --> 00:55:20,520
in the same flow as non-business connectors.

1264
00:55:20,520 --> 00:55:24,280
This prevents data from leaking because a flow literally cannot pull customer data and

1265
00:55:24,280 --> 00:55:26,880
send it to a public service if the policy forbids it.

1266
00:55:26,880 --> 00:55:29,520
Security roles define access at the dataverse level.

1267
00:55:29,520 --> 00:55:33,160
You control who sees which tables and which environments a maker can access.

1268
00:55:33,160 --> 00:55:36,520
A business user might have permission to build in a departmental environment, but stay

1269
00:55:36,520 --> 00:55:37,840
locked out of production.

1270
00:55:37,840 --> 00:55:41,360
A data steward might see sensitive tables, but have no right to edit them.

1271
00:55:41,360 --> 00:55:44,240
This role structure maps directly to your ring progression.

1272
00:55:44,240 --> 00:55:47,400
Dev roles are permissive while production roles are restrictive.

1273
00:55:47,400 --> 00:55:50,280
Environment variables replace hard-coded configurations.

1274
00:55:50,280 --> 00:55:54,680
Instead of embedding a specific API URL into a flow, you use a variable that changes based

1275
00:55:54,680 --> 00:55:55,680
on the ring.

1276
00:55:55,680 --> 00:56:00,440
In test, it points to the test API, and in production, it points to the production API.

1277
00:56:00,440 --> 00:56:04,400
The logic of the flow stays exactly the same, but the configuration adapts to wherever

1278
00:56:04,400 --> 00:56:05,640
it's running.

1279
00:56:05,640 --> 00:56:08,680
Power platform pipelines now provide ring promotion natively.

1280
00:56:08,680 --> 00:56:12,720
You configure a pipeline with stages from dev to production and the system shows you exactly

1281
00:56:12,720 --> 00:56:13,720
what's changing.

1282
00:56:13,720 --> 00:56:18,120
You move the move and the managed solution is exported and imported automatically.

1283
00:56:18,120 --> 00:56:21,600
Connections are reestablished to the right endpoints and configuration is updated without manual

1284
00:56:21,600 --> 00:56:22,600
intervention.

1285
00:56:22,600 --> 00:56:26,960
For complex organizations, these solutions integrate with Azure DevOps or GitHub.

1286
00:56:26,960 --> 00:56:30,960
You export the solution unpacked into source control for versioning and review and then

1287
00:56:30,960 --> 00:56:32,600
package it back up for deployment.

1288
00:56:32,600 --> 00:56:36,720
This hybrid approach gives you the audit trails of traditional DevOps combined with the

1289
00:56:36,720 --> 00:56:39,080
speed of a low-code authoring experience.

1290
00:56:39,080 --> 00:56:43,680
The fundamental shift is that power platform governance must account for citizen developers.

1291
00:56:43,680 --> 00:56:47,640
These users aren't writing YAML or managing infrastructure, they are building business logic

1292
00:56:47,640 --> 00:56:49,480
in a visual interface.

1293
00:56:49,480 --> 00:56:53,920
Governance cannot require DevOps expertise, so it has to be embedded in the platform itself.

1294
00:56:53,920 --> 00:56:58,800
DLP policies protect data without the developer needing to understand security architecture.

1295
00:56:58,800 --> 00:57:00,680
This is why a center of excellence is essential.

1296
00:57:00,680 --> 00:57:03,320
This isn't a traditional team of infrastructure engineers.

1297
00:57:03,320 --> 00:57:07,200
It's a cross-functional group that manages environment strategy, defines DLP policies

1298
00:57:07,200 --> 00:57:08,680
and creates templates.

1299
00:57:08,680 --> 00:57:12,920
The COE bridges the gap between business and technology to ensure that governance actually

1300
00:57:12,920 --> 00:57:15,200
enables innovation instead of blocking it.

1301
00:57:15,200 --> 00:57:19,000
Governance for low-code and citizen development, citizen developers are now building business

1302
00:57:19,000 --> 00:57:21,720
critical apps without any IT involvement.

1303
00:57:21,720 --> 00:57:25,920
A business analyst learns power apps and creates a tool that changes how their team

1304
00:57:25,920 --> 00:57:27,400
processes requests.

1305
00:57:27,400 --> 00:57:28,880
It works, so they build another.

1306
00:57:28,880 --> 00:57:31,280
Then they build a flow that connects to a backend service.

1307
00:57:31,280 --> 00:57:35,880
Suddenly, a business expert is touching systems that actually matter to the company.

1308
00:57:35,880 --> 00:57:36,880
This is happening everywhere.

1309
00:57:36,880 --> 00:57:39,480
It's the promise of low-code democratizing development.

1310
00:57:39,480 --> 00:57:42,760
You let the people who understand the business build the tools they need.

1311
00:57:42,760 --> 00:57:46,640
But the risk is that these experts are building systems without the discipline that professional

1312
00:57:46,640 --> 00:57:48,640
teams have spent years learning.

1313
00:57:48,640 --> 00:57:51,600
Governance for low-code is different because you can't code review a power app.

1314
00:57:51,600 --> 00:57:54,360
There is no pull request or text-based logic to inspect.

1315
00:57:54,360 --> 00:57:58,360
The maker is clicking buttons and configuring flows in a visual designer.

1316
00:57:58,360 --> 00:58:01,920
Since you can't put controls on the code, you have to put them on the environment.

1317
00:58:01,920 --> 00:58:05,360
You control which environments a developer can use and which connectors are available to

1318
00:58:05,360 --> 00:58:06,360
them.

1319
00:58:06,360 --> 00:58:10,280
These constraints happen at the platform level, making them invisible to the maker.

1320
00:58:10,280 --> 00:58:11,680
They don't see a list of rules.

1321
00:58:11,680 --> 00:58:14,320
They just see what is possible within their workspace.

1322
00:58:14,320 --> 00:58:16,400
DLP policies are your primary tool here.

1323
00:58:16,400 --> 00:58:20,400
By defining which connectors can work together, you prevent the most dangerous mistakes.

1324
00:58:20,400 --> 00:58:24,360
A business connector touching sensitive data cannot coexist with a personal cloud storage

1325
00:58:24,360 --> 00:58:26,120
connector in the same flow.

1326
00:58:26,120 --> 00:58:30,520
This makes it structurally impossible for a well-meaning employee to accidentally leak customer

1327
00:58:30,520 --> 00:58:31,520
data.

1328
00:58:31,520 --> 00:58:33,880
Your environment strategy separates risk through topology.

1329
00:58:33,880 --> 00:58:38,200
A sandbox environment is for experimentation with low governance and high freedom.

1330
00:58:38,200 --> 00:58:42,200
A departmental environment has more rules ensuring a developer and finance can't accidentally

1331
00:58:42,200 --> 00:58:44,080
connect to HR systems.

1332
00:58:44,080 --> 00:58:48,200
Enterprise environments run the mission-critical apps with maximum governance and strict change

1333
00:58:48,200 --> 00:58:49,200
control.

1334
00:58:49,200 --> 00:58:52,400
A citizen developer doesn't need to read policy documents to navigate this.

1335
00:58:52,400 --> 00:58:54,960
They navigate it through where they are allowed to create.

1336
00:58:54,960 --> 00:58:56,840
You don't tell them they can't use a connector.

1337
00:58:56,840 --> 00:58:59,160
You just make it unavailable in their environment.

1338
00:58:59,160 --> 00:59:01,520
You don't tell them to follow an approval process.

1339
00:59:01,520 --> 00:59:04,840
You put them in an environment that triggers approvals automatically.

1340
00:59:04,840 --> 00:59:07,760
Solution-based ALM adds another layer of enforcement.

1341
00:59:07,760 --> 00:59:11,480
This must be packaged, versed, and moved through rings from dev to production.

1342
00:59:11,480 --> 00:59:15,240
The citizen developer doesn't manage these mechanics manually because the center of excellence

1343
00:59:15,240 --> 00:59:16,840
has automated the process.

1344
00:59:16,840 --> 00:59:20,760
The developer creates the app and the structure handles the export and promotion.

1345
00:59:20,760 --> 00:59:24,360
The COE also provides templates so nobody has to start from a blank screen.

1346
00:59:24,360 --> 00:59:28,680
They provide reference apps, reusable approval flows, and dataverse tables that are already

1347
00:59:28,680 --> 00:59:30,280
compliant with company rules.

1348
00:59:30,280 --> 00:59:34,440
These templates set developers up for success by giving them a known good baseline to build

1349
00:59:34,440 --> 00:59:35,440
on.

1350
00:59:35,440 --> 00:59:38,360
The philosophy here is to enable rather than block.

1351
00:59:38,360 --> 00:59:42,520
Instead of saying no custom connectors, you show them the review process for requesting one.

1352
00:59:42,520 --> 00:59:46,560
Instead of saying "Ask permission", you tell them to build in the sandbox and show them

1353
00:59:46,560 --> 00:59:48,160
how promotion works.

1354
00:59:48,160 --> 00:59:51,320
Citizen developers follow governance when safe behavior is the default.

1355
00:59:51,320 --> 00:59:54,480
When the platform itself is the enforcement mechanism, you don't have to worry about people

1356
00:59:54,480 --> 00:59:55,760
remembering the rules.

1357
00:59:55,760 --> 00:59:57,360
The structure creates the discipline for you.

1358
00:59:57,360 --> 01:00:00,880
If this changed how you think about low code, follow me, Miracopieta's on LinkedIn,

1359
01:00:00,880 --> 01:00:02,280
and if you want more of this.

1360
01:00:02,280 --> 01:00:06,160
Give a review, it helps more people find it, share this with your team.

1361
01:00:06,160 --> 01:00:09,760
Especially if you're dealing with these governance challenges right now.

1362
01:00:09,760 --> 01:00:11,120
Scaling across the portfolio.

1363
01:00:11,120 --> 01:00:14,680
The governance model we've been discussing works for one platform and one tech stack.

1364
01:00:14,680 --> 01:00:16,040
It's a solid foundation.

1365
01:00:16,040 --> 01:00:19,280
But in a large enterprise, you aren't just managing one thing.

1366
01:00:19,280 --> 01:00:22,240
You're likely managing dozens of different systems at the exact same time.

1367
01:00:22,240 --> 01:00:25,240
You have a monolithic app that's been running for 15 years.

1368
01:00:25,240 --> 01:00:29,280
You have a microservices architecture built on Kubernetes just last year.

1369
01:00:29,280 --> 01:00:33,080
There are power platform solutions handling business logic, Python pipelines feeding your

1370
01:00:33,080 --> 01:00:36,120
data and machine learning models retraining every month.

1371
01:00:36,120 --> 01:00:39,320
On top of that, you have terraform managing your infrastructure.

1372
01:00:39,320 --> 01:00:41,400
Every single one of these needs to ship reliably.

1373
01:00:41,400 --> 01:00:45,800
They all need governance and they all have to live together in the same organization.

1374
01:00:45,800 --> 01:00:49,000
The common mistake is building a separate governance model for every team.

1375
01:00:49,000 --> 01:00:52,680
The Python developers use one process while the Kubernetes team uses another.

1376
01:00:52,680 --> 01:00:56,840
Power Platform has its own center of excellence and infrastructure as code follows a different

1377
01:00:56,840 --> 01:00:58,640
approval workflow entirely.

1378
01:00:58,640 --> 01:01:00,520
Within two years, you have total chaos.

1379
01:01:00,520 --> 01:01:04,960
You end up with different metrics, different definitions of ready and different ways to escalate

1380
01:01:04,960 --> 01:01:06,160
when things break.

1381
01:01:06,160 --> 01:01:09,880
When an incident finally hits multiple systems, it becomes a coordination nightmare because nobody

1382
01:01:09,880 --> 01:01:11,320
agreed on how to respond.

1383
01:01:11,320 --> 01:01:15,640
The solution is to unify your principles while letting the implementation stay flexible.

1384
01:01:15,640 --> 01:01:20,240
Every system, whether it's a cloud native service, a data pipeline or an AI model, shares

1385
01:01:20,240 --> 01:01:21,680
the same core needs.

1386
01:01:21,680 --> 01:01:24,560
They all need to prove a change is safe before it goes live.

1387
01:01:24,560 --> 01:01:26,960
They all need to limit the damage when something fails.

1388
01:01:26,960 --> 01:01:28,840
They all need audit trails for compliance.

1389
01:01:28,840 --> 01:01:33,440
The principle is universal, but the mechanics have to change because the systems are different.

1390
01:01:33,440 --> 01:01:37,320
This is where your platform team shifts from one central group into a federation.

1391
01:01:37,320 --> 01:01:39,920
You have a core platform team that owns the shared essentials.

1392
01:01:39,920 --> 01:01:45,000
They handle the logging, the identity management, the secret storage and the policy frameworks.

1393
01:01:45,000 --> 01:01:47,560
These are the common denominators that apply to everyone.

1394
01:01:47,560 --> 01:01:49,800
On top of that core, you build domain aligned teams.

1395
01:01:49,800 --> 01:01:53,720
The data platform team owns Python pipeline governance and data quality metrics.

1396
01:01:53,720 --> 01:01:58,760
The cloud native team manages Kubernetes, container images and service mesh policies.

1397
01:01:58,760 --> 01:02:01,720
The power platform team handles environments and DLP rules.

1398
01:02:01,720 --> 01:02:05,120
The AI team focuses on model validation and retraining workflows.

1399
01:02:05,120 --> 01:02:08,320
Each domain team owns its own rings and its own templates.

1400
01:02:08,320 --> 01:02:12,760
For the data team, ring zero might be a staging cluster, while ring one is a limited production

1401
01:02:12,760 --> 01:02:13,760
scope.

1402
01:02:13,760 --> 01:02:17,160
Their metrics are specific to data, like freshness and lineage integrity.

1403
01:02:17,160 --> 01:02:20,440
Their templates enforce things like SQL, Linting and data contracts.

1404
01:02:20,440 --> 01:02:22,920
The cloud native team defines rings for Kubernetes.

1405
01:02:22,920 --> 01:02:27,440
Ring zero is a test cluster and ring one is a canary deployment getting 5% of the traffic.

1406
01:02:27,440 --> 01:02:31,200
Their metrics focus on service health, like error rates and latency.

1407
01:02:31,200 --> 01:02:35,360
Power platform rings are environment based, moving from dev to test to production.

1408
01:02:35,360 --> 01:02:38,400
Their metrics track app performance and flow reliability.

1409
01:02:38,400 --> 01:02:41,560
Even though the tools change, everyone is operating under the same rules.

1410
01:02:41,560 --> 01:02:42,800
Every system uses rings.

1411
01:02:42,800 --> 01:02:44,280
Every system measures outcomes.

1412
01:02:44,280 --> 01:02:46,680
Every system has approval gates that match the risk.

1413
01:02:46,680 --> 01:02:48,640
The real complexity happens at the boundaries.

1414
01:02:48,640 --> 01:02:51,160
Imagine a power app that pulls data from a pipeline.

1415
01:02:51,160 --> 01:02:55,360
The data team updates the pipeline, the format changes and the power app breaks.

1416
01:02:55,360 --> 01:02:57,360
Usually neither team thinks they own the problem.

1417
01:02:57,360 --> 01:02:59,120
But in this model, the ownership is clear.

1418
01:02:59,120 --> 01:03:02,840
The data team owns the pipeline and the power platform team owns the app.

1419
01:03:02,840 --> 01:03:05,080
When a cross-domain issue hits, they share the response.

1420
01:03:05,080 --> 01:03:08,160
They run a shared incident command and a shared review afterward.

1421
01:03:08,160 --> 01:03:12,360
The goal is to bake those lessons back into the governance model so the same failure can't

1422
01:03:12,360 --> 01:03:13,360
happen twice.

1423
01:03:13,360 --> 01:03:16,400
Finally, the platform team as a whole looks at the big picture.

1424
01:03:16,400 --> 01:03:20,080
They track deployment frequency and recovery times across every domain.

1425
01:03:20,080 --> 01:03:23,560
They measure how satisfied developers are with the governance itself.

1426
01:03:23,560 --> 01:03:28,160
These high-level metrics show if the federation is actually working or if one domain is starting

1427
01:03:28,160 --> 01:03:29,400
to slow everyone else down.

1428
01:03:29,400 --> 01:03:32,200
This is how you scale to hundreds of teams and thousands of apps.

1429
01:03:32,200 --> 01:03:34,280
It's not about one team doing all the work.

1430
01:03:34,280 --> 01:03:36,280
It's about creating a structure that is repeatable.

1431
01:03:36,280 --> 01:03:39,440
The principles stay the same even as new domains plug in.

1432
01:03:39,440 --> 01:03:41,560
Growth doesn't mean you have to reinvent the rules.

1433
01:03:41,560 --> 01:03:45,800
It just means applying the same model to a new context.

1434
01:03:45,800 --> 01:03:49,880
Deployment platform success.

1435
01:03:49,880 --> 01:03:53,720
A federated platform only survives if you can prove it's actually working.

1436
01:03:53,720 --> 01:03:58,120
You have to measure whether the structure produces better results than the messy alternative.

1437
01:03:58,120 --> 01:03:59,800
Dora metrics are the starting point.

1438
01:03:59,800 --> 01:04:03,800
Deployment frequency shows if you're shipping more often and lead time tells you how fast

1439
01:04:03,800 --> 01:04:05,320
ideas reach the user.

1440
01:04:05,320 --> 01:04:09,440
Change failure rate, track stability, while recovery time shows how quickly you fix things

1441
01:04:09,440 --> 01:04:10,440
when they break.

1442
01:04:10,440 --> 01:04:13,800
These four metrics are the industry standard because they work for everything from legacy

1443
01:04:13,800 --> 01:04:15,800
to home.

1444
01:04:15,800 --> 01:04:19,800
You can improve your Dora scores by deleting all your governance.

1445
01:04:19,800 --> 01:04:21,800
You'd ship everything instantly.

1446
01:04:21,800 --> 01:04:24,800
Your frequency would skyrocket and lead time would drop.

1447
01:04:24,800 --> 01:04:27,800
But your failure rate would explode and you'd spend all your time in chaos mode.

1448
01:04:27,800 --> 01:04:29,800
The metrics look better on paper.

1449
01:04:29,800 --> 01:04:31,800
But the organization is actually worse.

1450
01:04:31,800 --> 01:04:36,800
To get the full picture you need platform specific metrics to see if you're helping or just getting in the way.

1451
01:04:36,800 --> 01:04:38,800
Start by looking at the time to first deployment.

1452
01:04:38,800 --> 01:04:42,800
When a new team joins, how long does it take them to get their first version into production?

1453
01:04:42,800 --> 01:04:45,800
In a world without templates, this takes weeks.

1454
01:04:45,800 --> 01:04:50,800
Engineers have to figure out branching, write pipelines from scratch and stumble over edge cases.

1455
01:04:50,800 --> 01:04:53,800
In a govern platform, they just pick a golden path template and deploy.

1456
01:04:53,800 --> 01:04:56,800
If it takes days instead of weeks, your platform is an accelerator.

1457
01:04:56,800 --> 01:04:59,800
If it still takes a month, you're adding friction.

1458
01:04:59,800 --> 01:05:01,800
You also need to track the adoption of your golden paths.

1459
01:05:01,800 --> 01:05:03,800
Don't force people to use them.

1460
01:05:03,800 --> 01:05:05,800
Just watch if they choose them because they're better.

1461
01:05:05,800 --> 01:05:09,800
If 70% of teams adopt a template within a year, you've solved a real problem.

1462
01:05:09,800 --> 01:05:14,800
If adoption stays at 30%, your template is either too restrictive or it doesn't fit how people actually work.

1463
01:05:14,800 --> 01:05:17,800
The low adoption is a signal that you need to iterate.

1464
01:05:17,800 --> 01:05:19,800
Next, look at the self-service completion rate.

1465
01:05:19,800 --> 01:05:22,800
This measures if teams can get what they need without waiting on a ticket.

1466
01:05:22,800 --> 01:05:28,800
In a good model, developers provision their own environments and onboard their own services through a portal.

1467
01:05:28,800 --> 01:05:32,800
If 85% of these tasks happen without a human involved, your automation is winning.

1468
01:05:32,800 --> 01:05:37,800
If it's below 60%, your team is still a bottleneck for routine work, then there's the human element.

1469
01:05:37,800 --> 01:05:39,800
PlatformNPS asks a simple question.

1470
01:05:39,800 --> 01:05:42,800
Would you recommend this platform to a team mate?

1471
01:05:42,800 --> 01:05:44,800
If the score is high, the platform isn't enabler.

1472
01:05:44,800 --> 01:05:47,800
If it's low, the team see you as an obstacle.

1473
01:05:47,800 --> 01:05:52,800
You should also track how much time developers spend on infrastructure versus actual product work.

1474
01:05:52,800 --> 01:05:56,800
If they're spending more than 20% of their time fighting the platform, you failed.

1475
01:05:56,800 --> 01:05:59,800
The whole point of the platform is to take that burden away, not make it heavier.

1476
01:05:59,800 --> 01:06:01,800
Finally, you have to connect this to the business.

1477
01:06:01,800 --> 01:06:04,800
Does the platform speed up how fast you launch new features?

1478
01:06:04,800 --> 01:06:08,800
Does it reduce the cost of downtime? Does it make each deployment cheaper?

1479
01:06:08,800 --> 01:06:10,800
The research is very clear on this.

1480
01:06:10,800 --> 01:06:15,800
Organizations that measure their platform success improve two to three times faster than those that don't.

1481
01:06:15,800 --> 01:06:19,800
They aren't just better at measuring. They're better at improving the platform itself.

1482
01:06:19,800 --> 01:06:23,800
Because when you have data, you stop guessing and start making decisions based on evidence.

1483
01:06:23,800 --> 01:06:26,800
The biggest mistake is measuring activities instead of outcomes.

1484
01:06:26,800 --> 01:06:30,800
Don't count how many pipelines you build or how many teams signed up. Those are just numbers.

1485
01:06:30,800 --> 01:06:33,800
Instead, measure what those pipelines actually allow the business to do.

1486
01:06:33,800 --> 01:06:38,800
Focus on speed, safety and developer happiness, the activity doesn't matter. The outcome is everything.

1487
01:06:38,800 --> 01:06:44,800
These dashboards shouldn't be hidden away. They need to be visible to leadership, product teams and the people who control the budget.

1488
01:06:44,800 --> 01:06:49,800
Transparency creates accountability. The platform team has to earn its keep by showing measurable value every single day.

1489
01:06:49,800 --> 01:06:52,800
That's how you get the investment you need to keep growing.

1490
01:06:52,800 --> 01:07:00,800
Once you have these measurements in place, you finally have the data you need to build a roadmap that takes you from where you are to where you need to be.

1491
01:07:00,800 --> 01:07:10,800
Implementation roadmap. The model you've learned, federated teams, golden paths, ring deployments and automated gates looks perfect on paper, but in reality, it's a sequence.

1492
01:07:10,800 --> 01:07:17,800
You can't jump to fully automated gates if you don't have observability. You can't expect teams to use golden paths if those paths don't exist yet.

1493
01:07:17,800 --> 01:07:21,800
The roadmap is how you translate strategy into steps without breaking the organization.

1494
01:07:21,800 --> 01:07:29,800
Phase one, assessment, months one to three. You aren't building yet. You're measuring, inventory every CICD system you own.

1495
01:07:29,800 --> 01:07:36,800
GitHub actions, Azure DevOps, Jenkins, GitLab. Count the pipelines, are there hundreds, thousands, look for the patterns and look for the chaos.

1496
01:07:36,800 --> 01:07:42,800
Interview the people running them, what frustrates them, what takes the most time, what breaks every single week.

1497
01:07:42,800 --> 01:07:46,800
Measure your current state against door metrics like deployment frequency and lead time.

1498
01:07:46,800 --> 01:07:52,800
Measure developer satisfaction separately. How much of their week is spent fighting the pipeline?

1499
01:07:52,800 --> 01:07:58,800
This assessment is uncomfortable because it exposes inefficiency, but you need the data. It's your case for investment.

1500
01:07:58,800 --> 01:08:05,800
Phase two, core platform, months three to six. Now you build, but you build narrow, set up the shared infrastructure that every team needs.

1501
01:08:05,800 --> 01:08:12,800
Central logging, metrics collection, secret management, a shared runner pool, this infrastructure is expensive to duplicate.

1502
01:08:12,800 --> 01:08:20,800
But it's cheap to share. It isn't opinionated about templates yet. It's just the foundation. Create two or three golden path templates. Don't try to cover everything.

1503
01:08:20,800 --> 01:08:34,800
If 70% of your services are Java microservices, build a Java template. If power platform is critical, build a power platform, ALM template. These templates are opinionated. They encode your governance. And they make the safe choice, the default.

1504
01:08:34,800 --> 01:08:41,800
Phase three, pilot, months six to nine. Select five to ten teams to be your design partners. Don't just take volunteers who are desperate for help.

1505
01:08:41,800 --> 01:08:46,800
Choose teams that represent your diversity. A web team, a data team, a legacy monolith team.

1506
01:08:46,800 --> 01:08:53,800
Give them the templates and support them directly. Don't send them documentation. Sit with them. Build together. Understand what works and what doesn't.

1507
01:08:53,800 --> 01:08:59,800
Collect feedback ruthlessly. If the template doesn't fit, you need to know why. Is the use case an outlier? Or is your template too rigid?

1508
01:08:59,800 --> 01:09:09,800
Iteration is uncomfortable. But by the end of phase three, your templates have survived contact with reality. They are better because they've been tested in the real world.

1509
01:09:09,800 --> 01:09:18,800
Phase four, expansion, months nine to twelve. You move from five teams to 50. This is where adoption programs matter, training, champions networks, office hours.

1510
01:09:18,800 --> 01:09:28,800
Don't assume people will figure it out. The difference between adoption and rejection is whether teams feel supported. Invest in Slack channels, runbooks and guides written in plain language.

1511
01:09:28,800 --> 01:09:45,800
Phase five, full rollout, months twelve to eighteen. The platform becomes the standard path. Legacy systems migrate. Holdouts have to justify why they're building outside the platform. By now the value is proven. Faster deployments, reliable releases, developers prefer it because it removes the burden.

1512
01:09:45,800 --> 01:09:53,800
Phase six, continuous improvement, ongoing. You don't stop. Measure outcomes, identify friction, evolve the templates.

1513
01:09:53,800 --> 01:10:10,800
The platform that solves your problems in month eighteen will be incomplete by month twenty four. That isn't a failure. It's success. You're only enhancing it because adoption exceeded your expectations. The core philosophy is simple. Don't build perfect. Build useful. Get feedback. Improve. Repeat.

1514
01:10:10,800 --> 01:10:19,800
Most organizations fail by trying to design the perfect platform in isolation. By the time they launch requirements have shifted, teams have found workarounds. Adoption stalls.

1515
01:10:19,800 --> 01:10:26,800
Start small, build in public. Let real use it shape the platform. It's slower at the start, but it's much faster at the finish.

1516
01:10:26,800 --> 01:10:38,800
The organizational shift, implementing this model requires something you can't code. It's a shift in how teams think about ownership. For years the assumption was, we own our pipeline. We make the decisions. If it breaks, we fix it.

1517
01:10:38,800 --> 01:10:47,800
This autonomy feels like freedom. In reality, it's just distributed burden. The new model inverts that assumption. We use a shared platform. Our logic is ours, but the mechanics are shared.

1518
01:10:47,800 --> 01:10:53,800
The security is non-negotiable. The templates are standard because they've been proven to work. This feels restrictive at first.

1519
01:10:53,800 --> 01:11:06,800
But it's actually freedom, just freedom with guardrails. The resistance is real. The platform doesn't support my use case. Maybe it doesn't, but is that use case common enough to justify a parallel system? Or is it a five percent outlier that deserves a conversation?

1520
01:11:06,800 --> 01:11:16,800
The platform teams job is to listen. If multiple teams have the same need, it goes in the template. If one team is unique, that's where exceptions exist. Exceptions are expensive. So they attract. When they become common, the template evolves.

1521
01:11:16,800 --> 01:11:27,800
Champions are vital in this transition. Identify the power users who already believe automation saves time. Inlist them as evangelists. Give them early access. Let them experience the benefits first.

1522
01:11:27,800 --> 01:11:41,800
Peer-to-peer conversion is always more credible than a mandate from the top. Leadership alignment is the other half. If leadership values speed but not safety, teams will find ways around governance. They'll build shadow platforms. They'll ship things without testing because the business needs it now.

1523
01:11:41,800 --> 01:11:54,800
If leadership values both, the message is clear. Governance isn't the enemy of velocity. It's the mechanism that makes velocity sustainable. And the shift is also about identity. Platform teams aren't traditional ops teams. They are product teams that build infrastructure.

1524
01:11:54,800 --> 01:12:03,800
This distinction changes everything. They don't maintain systems that people are forced to use. They build products that teams choose to use. That mindset drives an obsession with usability and support.

1525
01:12:03,800 --> 01:12:15,800
The platform team measures developer satisfaction. Like a product team measures customers. They iterate when the feedback shows they're off track. They drive adoption by making the platform so useful that choosing it is the only rational move.

1526
01:12:15,800 --> 01:12:21,800
This is uncomfortable for old school central teams. Those teams were gatekeepers. They owned the infrastructure and they control the access.

1527
01:12:21,800 --> 01:12:30,800
In the new model, they become enablers. Their job is to reduce barriers, not create them. They make safe behavior the default. Teams that can't make this mental shift will struggle.

1528
01:12:30,800 --> 01:12:45,800
Resistance also comes from sunk cost. Teams have spent years on their current pipelines, starting over feels like waste. The reframe is this. The platform isn't asking you to throw away your learning. It's asking you to stop duplicating solutions. Your expertise doesn't disappear. It informs the platform.

1529
01:12:45,800 --> 01:12:58,800
Six to 12 months in, two groups emerge, the adopters and the holdouts. The adopters are shipping faster. Their deployments are reliable. Their on call shifts aren't chaotic. The holdouts see this. They see one team going home at 5pm while they're dealing with incidents at midnight.

1530
01:12:58,800 --> 01:13:05,800
That creates momentum, not enforcement, momentum. The organizational shift is the hardest part because it requires people to change who they are.

1531
01:13:05,800 --> 01:13:16,800
You move from owner of my pipeline to member of an ecosystem. That's a loss of control. But it's a gain in results. Teams get to ship faster. They get to focus on products instead of plumbing. They get to sleep at night. The shift is real.

1532
01:13:16,800 --> 01:13:27,800
Making it happen requires leadership, patience and a relentless focus on outcomes over compliance. The model behind the model. Scaling CI/CD isn't a tool problem. It's a structure problem.

1533
01:13:27,800 --> 01:13:35,800
The governance blueprint is deceptively simple. Standardize through templates, measure with observability, gate with policies, scale through automation.

1534
01:13:35,800 --> 01:13:40,800
Ring-based deployments reduce your blast radius. Canary releases validated scale. Automated gates remove the toil.

1535
01:13:40,800 --> 01:13:48,800
In this model, the platform team operates as a product team. Developers are the customers. Satisfaction and adoption matter more than features shipped.

1536
01:13:48,800 --> 01:13:55,800
This model works for traditional apps, cloud native services, power platform solutions, data pipelines and AI models every domain.

1537
01:13:55,800 --> 01:14:03,800
Not because one tool fits all contexts, but because the principles are universal start with one golden path. Measure the outcomes. Expand gradually.

1538
01:14:03,800 --> 01:14:10,800
The model scales because the structure repeats. Your next step is to assess your current state. Where is the biggest friction? Start there.

Mirko Peters Profile Photo

Founder of m365.fm, m365.show and m365con.net

Mirko Peters is a Microsoft 365 expert, content creator, and founder of m365.fm, a platform dedicated to sharing practical insights on modern workplace technologies. His work focuses on Microsoft 365 governance, security, collaboration, and real-world implementation strategies.

Through his podcast and written content, Mirko provides hands-on guidance for IT professionals, architects, and business leaders navigating the complexities of Microsoft 365. He is known for translating complex topics into clear, actionable advice, often highlighting common mistakes and overlooked risks in real-world environments.

With a strong emphasis on community contribution and knowledge sharing, Mirko is actively building a platform that connects experts, shares experiences, and helps organizations get the most out of their Microsoft 365 investments.