M365con.net Microsoft Community Conference 2027
M365 FM Podcast
M365 FM Podcast
The M365 FM Podcast is your daily destination for everything happening across the Microsoft cloud. We cover the full spectrum of Microsoft 365, including Teams, SharePoint, Exchange, OneDrive, and the tools driving the modern workplace. Each episode delivers practical insights, expert interviews, and hands-on strategies for IT admins, cloud architects, developers, power users, and decision-makers in the Microsoft ecosystem. We explore the latest M365 updates, dive into Power Platform topics like Power Apps, Power Automate, Power BI, Power Pages, and share real-world guidance on automation, digital transformation, and low-code development. You’ll also get deep insights into Azure, including cloud infrastructure, Azure AD / Entra ID, identity, hybrid cloud, and Azure security. The show features focused discussions on Microsoft 365 Security, Defender, compliance, DLP, Zero Trust, and the best practices needed to protect and optimize your environment. We also highlight how AI and Copilot for Microsoft 365 are transforming productivity, collaboration, and automation across the cloud. Whether you want to improve Teams collaboration, strengthen security, enhance cloud architecture, or stay ahead of the latest Microsoft 365, Azure, Power Platform, and AI announcements, The M365 Podcast is your essential guide. M365 FM Podcast is Part of the M365.Show Network.
Oct. 1, 2026

Building Azure That Survives the Real World — with Mike Martin [MVP]

Building Azure That Survives the Real World — with Mike Martin [MVP]
Building Azure That Survives the Real World — with Mike Martin [MVP]
M365 FM Podcast
Building Azure That Survives the Real World — with Mike Martin [MVP]

Key Takeaways

  • Azure architecture is not just about writing code or choosing the latest technologies; it requires balancing scalability, security, maintenance, and cost.
  • Mike Martin emphasizes that fundamental IT challenges like DNS and IP dependencies remain constant, while solutions have grown increasingly over-engineered.
  • Successful cloud design begins with clear business requirements—such as user load, compliance, and budget—rather than jumping straight to specific Azure services like Kubernetes.
  • Resilience goes beyond high availability by ensuring a system can absorb load, handle partial backend failures, and recover gracefully without a complete collapse.
  • Achieving 100 percent uptime is a myth; instead, architects should focus on fallback scenarios, quick recovery times, and defining realistic RTO and RPO metrics based on actual business impact.

Azure architecture looks easy when everything works.The real test starts when dependencies fail, regions become unavailable, traffic spikes, assumptions turn out to be wrong, requirements change, and someone eventually asks the uncomfortable question: why did we design it this way in the first place?In this episode of M365.FM, Mirko Peters talks with Mike Martin [MVP] about what Azure architecture looks like when it has to survive real production conditions rather than just look good on a diagram.Mike brings decades of experience across development, infrastructure, architecture, leadership, coaching, training, and Microsoft Azure. One of his strongest observations is that many of the problems architects face today are not actually new. DNS still breaks. IP dependencies still matter. Costs still become a problem. Integrations still fail. Dependencies still disappear at the worst possible moment.What has changed is the level of complexity we build around those problems.

FROM VISUAL BASIC TO MODERN AZURE ARCHITECTURE
Mike looks back at nearly three decades in IT, starting as a Visual Basic developer in the 1990s and moving through distributed systems, networking, enterprise software, infrastructure, and eventually Azure.His key observation is simple: the industry keeps solving many of the same fundamental problems, but the architectures around them have become much more complex.Modern systems have moved from client-server applications to distributed architectures, cloud platforms, containers, microservices, Kubernetes, hybrid environments, and now AI-assisted development.That creates enormous possibilities, but it also creates a new risk: overengineering.Mike argues that many teams today make simple problems unnecessarily complicated. Good architecture is often not about adding more technology. It is about knowing what not to add.

WHAT DOES AN AZURE ARCHITECT ACTUALLY DO?
For Mike, architecture is not about choosing the largest number of Azure services or producing an impressive diagram.It is about understanding which components belong together, which ones should be avoided, which ones are necessary, and how to build something that remains maintainable, scalable, secure, and resilient.Architecture includes much more than compute.

  • Identity
  • Networking
  • Data flows
  • Security
  • Monitoring
  • Integration
  • Dependencies
  • Scalability
  • Operations
  • Recovery
  • Deployment
  • Cost

Mike also challenges the idea that cloud-native automatically means Kubernetes or containers.Azure provides many managed and native services that can solve problems without introducing unnecessary operational overhead.The architect’s role is to understand the complete solution and choose the simplest architecture that still satisfies the real requirements.

START WITH BUSINESS REQUIREMENTS, NOT AZURE SERVICES
One of the most important lessons in this episode is simple: do not start with technology.Before deciding between Azure Kubernetes Service, App Service, Azure Functions, containers, Service Bus, or another platform, teams should first understand what the solution actually needs to do.Questions to ask:

  • Is it internal or customer-facing?
  • How many users will depend on it?
  • Does it need to scale globally?
  • How long can it be unavailable?
  • How much data can the business afford to lose?
  • Which compliance requirements apply?
  • What happens if the application disappears for several hours?
  • Who is affected?
  • What level of operational support is required?

These questions lead directly into concepts such as SLAs, SLOs, RTOs, and RPOs.They also determine whether the architecture should be single-region, multi-region, active-active, active-passive, or something much simpler.

RTO AND RPO WITHOUT THE BUZZWORDS
RTO and RPO are often discussed as technical acronyms, but their real meaning is business-oriented.

  • RTO — Recovery Time Objective: How quickly must the system return after a failure?
  • RPO — Recovery Point Objective: How much data loss is acceptable?

A system used for non-critical monitoring may tolerate several hours of downtime or lost data.A system supporting first responders, financial operations, commerce, or critical infrastructure may require recovery in minutes.The important point is that these numbers should not be invented by the architect.They should come from the actual business impact of failure.Once those requirements are understood, they can be translated into technical design decisions.

RESILIENCE IS NOT THE SAME AS HIGH AVAILABILITYMike uses a simple analogy to explain the difference between availability and resilience.A highly available system may have another component ready to take over when the primary one fails.A resilient system is designed to absorb problems, continue functioning, recover gracefully, and return to normal without collapsing completely.That distinction matters.An application can technically be available while still providing a poor experience.It may be:

  • Slow
  • Throttled
  • Partially unavailable
  • Dependent on a failing backend
  • Affected by an integration issue
  • Under regional pressure

Real resilience therefore requires more than uptime.It requires an architecture that can cope with load, application errors, partial failures, regional issues, broken dependencies, operational incidents, and recovery after the incident has passed.

AZURE DOES NOT MAKE YOUR APPLICATION RESILIENT AUTOMATICALLY
One dangerous assumption is that because Azure itself is highly available, any application running on Azure automatically inherits that resilience.It does not.Microsoft provides the services and capabilities that make resilient architectures possible.Customers still need to design for resilience.This is where architectural patterns matter:

  • Circuit breakers
  • Loose coupling
  • Asynchronous processing
  • Independent scaling
  • Multiple availability zones
  • Multiple regions
  • Failover mechanisms
  • Monitoring
  • Infrastructure as Code
  • Tested recovery procedures

Microsoft provides the building blocks.The architecture determines whether those building blocks actually create a resilient system.

WHY ASYNCHRONOUS DESIGN MATTERS
One of the strongest examples in the episode comes from systems that experience predictable traffic spikes, such as government tax portals.A common design problem occurs when every action depends on a synchronous backend operation.The user clicks a button.The frontend waits for the backend.The backend waits for another service.Another dependency slows down.Eventually the entire user experience is affected.Mike explains why loosely coupled systems are often more resilient.Instead of forcing the frontend to wait for every backend operation, applications can place work into queues or event-driven systems and allow components to process tasks independently.That creates several advantages:

  • Different components can scale independently
  • One service can fail without taking down the whole platform
  • Backends can process work asynchronously
  • Traffic spikes become easier to absorb
  • User-facing components become less dependent on backend timing
  • Failures become easier to isolate

THE MYTH OF 100 PERCENT UPTIME
Another major topic in the discussion is the obsession with 100 percent availability.Mike is very clear: 100 percent uptime is not a realistic architecture target.A solution is usually composed of multiple services.Each service has its own SLA.Once several dependencies are combined, the effective availability of the complete system changes.A better architecture focuses on:

  • Fallback scenarios
  • Recovery procedures
  • Redundancy
  • Failover
  • Monitoring
  • Mitigation strategies
  • Tested recovery plans

The better question is not: how do we guarantee zero downtime?The better question is: what happens when something fails, and how quickly can the business continue operating

WHEN ANOTHER NINE BECOMES TOO EXPENSIVE
More availability is not automatically better.Every additional level of resilience introduces cost, infrastructure, monitoring, operational complexity, and management overhead.Whether that investment makes sense depends entirely on the business impact of failure.An internal holiday request system can probably tolerate several hours of downtime.An order-processing platform handling millions in transactions cannot.Questions that matter:

  • How much money is lost during downtime?
  • How many users are affected?
  • What happens to customer trust?
  • What happens if an API becomes slow?
  • What happens if orders cannot be processed?
  • What does another level of redundancy actually cost?
  • Is the additional availability worth that cost?

Reliability is ultimately an economic decision as much as a technical one.

Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.

🚀 Want to be part of m365.fm?

Then stop just listening… and start showing up.

👉 Connect with me on LinkedIn and let’s make something happen:

  • 🎙️ Be a podcast guest and share your story
  • 🎧 Host your own episode (yes, seriously)
  • 💡 Pitch topics the community actually wants to hear
  • 🌍 Build your personal brand in the Microsoft 365 space

This isn’t just a podcast — it’s a platform for people who take action.

🔥 Most people wait. The best ones don’t.

👉 Connect with me on LinkedIn and send me a message:
"I want in"

Let’s build something awesome 👊

Frequently Asked Questions

Who is Mike Martin and what is his background in IT?

Mike Martin is an Azure MVP with nearly three decades of experience in IT, starting as a Visual Basic developer in the 1990s and evolving into distributed systems, infrastructure, and modern Microsoft Azure architecture.

What is the difference between RTO and RPO in Azure architecture?

RTO (Recovery Time Objective) measures how quickly a system must return to operation after a failure, while RPO (Recovery Point Objective) defines how much data loss is acceptable to the business.

Why is asynchronous design important for Azure applications?

Asynchronous design uses queues and event-driven patterns to decouple front-end and back-end components, allowing systems to absorb traffic spikes and continue functioning even if a backend service fails.

Do you always need Kubernetes when building on Azure?

No, cloud-native does not automatically mean Kubernetes or containers. Azure offers numerous native and managed services that can solve business problems without introducing unnecessary operational overhead.

1
00:00:00,000 --> 00:00:07,800
Hello everyone and welcome back to the M65 as M podcast.

2
00:00:07,800 --> 00:00:13,360
Azure Architation is easy when you.

3
00:00:13,360 --> 00:00:20,640
Yeah, when everything works, then interesting question starts when dependencies fail, regions

4
00:00:20,640 --> 00:00:29,160
become unavailable, assumptions turn out to be a wrong coast rise.

5
00:00:29,160 --> 00:00:38,600
And teams grow, we probably haven't changed and somebody everybody asked.

6
00:00:38,600 --> 00:00:42,360
Why did we design it this way?

7
00:00:42,360 --> 00:00:48,680
Mike Markin has spent almost three decades in IT across development, infrastructure, architecture,

8
00:00:48,680 --> 00:00:59,040
system design, leadership coaching, training and community work with deep experience across

9
00:00:59,040 --> 00:01:07,120
Microsoft Azure and particular focus on architecture and resistance.

10
00:01:07,120 --> 00:01:14,920
This is, yeah, we are, I make it short, we explore how the world, rather than walking

11
00:01:14,920 --> 00:01:18,880
through in middle Azure services.

12
00:01:18,880 --> 00:01:28,640
This interview is focused on architecture judgment, how we make better decision, how we

13
00:01:28,640 --> 00:01:37,360
design, tell you what architecture is really wrong and how the role of Azure architecture

14
00:01:37,360 --> 00:01:38,360
is changing.

15
00:01:38,360 --> 00:01:47,760
Yeah, Mike, you are being working in IT for 30 years.

16
00:01:47,760 --> 00:01:54,400
That's amazing when you look back, when you look back, where are the major turning points

17
00:01:54,400 --> 00:01:58,040
that shape the way you're thinking about technology?

18
00:01:58,040 --> 00:02:02,120
Yeah, first of all, it's like showing age, right?

19
00:02:02,120 --> 00:02:04,360
Well, the stuff you're really old.

20
00:02:04,360 --> 00:02:13,880
Yeah, 30 years, I started, professionally, I started in 96, 97.

21
00:02:13,880 --> 00:02:17,840
I started my career as a visual basic developer, by the way.

22
00:02:17,840 --> 00:02:19,440
Visual basic, yeah.

23
00:02:19,440 --> 00:02:20,440
Visual basic, yeah.

24
00:02:20,440 --> 00:02:21,440
Yeah, yeah.

25
00:02:21,440 --> 00:02:28,320
I have learned, in a, a new view was about it.

26
00:02:28,320 --> 00:02:35,400
So I was basically designing software, all really with distributed software and the

27
00:02:35,400 --> 00:02:40,080
price based, visual basic, so like, common, complex and deep come in all these things and

28
00:02:40,080 --> 00:02:43,840
transactions, sir, and all the cool things that we had at the time.

29
00:02:43,840 --> 00:02:50,480
At today, my main business is mostly now Azure and in between, I did a lot of things and

30
00:02:50,480 --> 00:02:54,960
you asked me what changed over these 30 years.

31
00:02:54,960 --> 00:02:59,320
Well, the one thing that is a constant is that all the problems keep reappearing.

32
00:02:59,320 --> 00:03:03,560
There's always things like DNS, there's always things like IP changes, there's always things

33
00:03:03,560 --> 00:03:08,760
like, like, dependencies that get broken, there's always these things.

34
00:03:08,760 --> 00:03:19,800
But the biggest change I think is the way that we move from a single application into

35
00:03:19,800 --> 00:03:23,120
a more complex architecture.

36
00:03:23,120 --> 00:03:29,640
So we used to have client server, we used to have distributed, compute where we had middle

37
00:03:29,640 --> 00:03:35,120
tiers and multiple tiers and then we had layers and then we came up with these new concepts

38
00:03:35,120 --> 00:03:38,480
of, let's move it into the cloud and do some hybrid.

39
00:03:38,480 --> 00:03:45,520
And then before we knew it, we had all these things like Kubernetes and microservices and

40
00:03:45,520 --> 00:03:46,520
all these things.

41
00:03:46,520 --> 00:03:52,840
I think the biggest change there is the way that we're looking at software, naming it for

42
00:03:52,840 --> 00:04:01,080
the better because we humans are kind of, you know, most happy spot if you can overcomplexify

43
00:04:01,080 --> 00:04:02,880
things.

44
00:04:02,880 --> 00:04:08,000
And I see that a lot nowadays because I see designs of applications, I go like, or solutions,

45
00:04:08,000 --> 00:04:11,640
I go like, oh my god, what have they built it now?

46
00:04:11,640 --> 00:04:15,320
There's a lot of a lot of varsity things that are thinking that they need to be like Dr.

47
00:04:15,320 --> 00:04:20,640
Frank and Stein in that they also need to create a monster instead of creating a solution.

48
00:04:20,640 --> 00:04:24,160
So that's basically that.

49
00:04:24,160 --> 00:04:31,320
But I think that is the thing that I see most in the change is the way we have all these

50
00:04:31,320 --> 00:04:38,800
technologies and we tend to complexify a lot of things and we cannot keep it simple anymore.

51
00:04:38,800 --> 00:04:41,880
Even when it's a simple solution, we cannot keep it simple anymore.

52
00:04:41,880 --> 00:04:45,840
And I think that is the biggest change and I think that will change for the better over the

53
00:04:45,840 --> 00:04:52,120
next course of say next 15 to 20 years, especially not with the coming of all these things that

54
00:04:52,120 --> 00:04:57,560
we now at hand like AI because, you know, developers don't need to know anymore what they're

55
00:04:57,560 --> 00:05:03,000
doing, they're just five coded and they go ahead and push it to production, right?

56
00:05:03,000 --> 00:05:04,760
So there's that.

57
00:05:04,760 --> 00:05:08,360
So yeah, that's, I think that's, I think that is the biggest change that I've seen over

58
00:05:08,360 --> 00:05:14,080
the course of 30 years that we still have the same problems.

59
00:05:14,080 --> 00:05:18,920
We just made them more complex to deeper into to do to detangle basically.

60
00:05:18,920 --> 00:05:21,800
That's that's what it comes down to.

61
00:05:21,800 --> 00:05:29,280
I see on your link the profile you are as you take.

62
00:05:29,280 --> 00:05:37,480
And that is a bit hard, but the king can almost, it can mean almost everything.

63
00:05:37,480 --> 00:05:50,400
What do the architecture mean to you and how, yeah, what the person's, the person's of architecture

64
00:05:50,400 --> 00:05:55,520
and technology or technical stuff that you do?

65
00:05:55,520 --> 00:06:02,440
What do, what's the real, what's the architecture?

66
00:06:02,440 --> 00:06:09,520
Okay, so I was very lucky to be on the very first basis of Azure.

67
00:06:09,520 --> 00:06:10,960
I started in the early days.

68
00:06:10,960 --> 00:06:19,520
So when Azure was still red dog, while it was still a little, a little puppy, I was there

69
00:06:19,520 --> 00:06:20,520
at the PDC.

70
00:06:20,520 --> 00:06:25,520
I was there when Ray O'Z announced that we're going to have a cloud.

71
00:06:25,520 --> 00:06:31,440
And I saw the first portal and actually have been actively playing around or building with

72
00:06:31,440 --> 00:06:36,120
it since 2009 since the first installments.

73
00:06:36,120 --> 00:06:40,680
So 2009 was really playing and having a sample size somewhere and things like that.

74
00:06:40,680 --> 00:06:43,440
And portals were like nonexistent.

75
00:06:43,440 --> 00:06:46,560
The first portal was one button remembered well.

76
00:06:46,560 --> 00:06:50,640
And then we had that silver lighting, then we had that silver lighting, which nobody remembers

77
00:06:50,640 --> 00:06:53,160
apparently.

78
00:06:53,160 --> 00:06:54,720
And that thing evolved.

79
00:06:54,720 --> 00:07:01,320
And with the evolving part of Azure complexity also came, I come from a developer back

80
00:07:01,320 --> 00:07:07,600
round, but I also have an infrastructure background.

81
00:07:07,600 --> 00:07:15,680
I came in from Windows 2,000 networking into, and also some Unix networking system monitoring.

82
00:07:15,680 --> 00:07:19,640
And Azure for me, it all came together.

83
00:07:19,640 --> 00:07:23,400
And this is where I found my interest.

84
00:07:23,400 --> 00:07:29,160
I didn't want to build a search as just writing code, but I wanted to make sure that the solutions

85
00:07:29,160 --> 00:07:38,720
could be existing, the way cloud was intended to be scalable and resilient at every level.

86
00:07:38,720 --> 00:07:45,520
And for me, architecture is about knowing which pieces go together, which pieces to avoid,

87
00:07:45,520 --> 00:07:50,080
which pieces are necessary and which pieces are not.

88
00:07:50,080 --> 00:07:57,120
And building something beautiful out of it, which is not complex, which makes it easy to

89
00:07:57,120 --> 00:08:07,240
maintain, and which makes it also as cloud native as possible without referring to cloud native.

90
00:08:07,240 --> 00:08:12,520
Because cloud native has just a connotation, it has to be Kubernetes, it has to be that.

91
00:08:12,520 --> 00:08:15,560
News flash, you don't always need Kubernetes, you don't always need containers.

92
00:08:15,560 --> 00:08:19,600
There's a lot of native services that you can just use and make them all work together.

93
00:08:19,600 --> 00:08:24,080
But for me, architecture is more about design a full solution.

94
00:08:24,080 --> 00:08:28,600
It's not about one single application, you need to be able to cope with integrations, you

95
00:08:28,600 --> 00:08:34,320
need to be able to cope with everything surrounding that solution, like monitoring security,

96
00:08:34,320 --> 00:08:35,720
it all comes together.

97
00:08:35,720 --> 00:08:38,760
And that's for me is the architectural part.

98
00:08:38,760 --> 00:08:45,080
And there's some annoyance when I look at certain other architics, and I'm not saying

99
00:08:45,080 --> 00:08:49,120
that they are bad, who are they doing.

100
00:08:49,120 --> 00:08:55,200
But typically most of the architics that I'm seeing today are being architated from a way

101
00:08:55,200 --> 00:08:59,520
that they were building architectures in a data center.

102
00:08:59,520 --> 00:09:05,720
VDIP-based, VLamb based, VDNS based, VR-based, VR-based, VR-based, VR-based, VR-based, VR-based,

103
00:09:05,720 --> 00:09:09,640
network oriented, which is always the best way.

104
00:09:09,640 --> 00:09:15,400
You should look at data flows, you should look at identity flows, you should look at security

105
00:09:15,400 --> 00:09:19,080
flows, you should look at how compute needs to be distributed.

106
00:09:19,080 --> 00:09:22,360
What are the capabilities that you need?

107
00:09:22,360 --> 00:09:23,360
These things.

108
00:09:23,360 --> 00:09:27,160
And it typically start designing wrong in my eyes.

109
00:09:27,160 --> 00:09:28,160
Okay?

110
00:09:28,160 --> 00:09:32,440
Does that sound like, does it sound like bragging or like me being uptight?

111
00:09:32,440 --> 00:09:38,240
No, it's just, I've seen things that do that work, and I try to make sure that people

112
00:09:38,240 --> 00:09:42,200
don't make the same mistakes that other people made.

113
00:09:42,200 --> 00:09:44,480
That's for me architics.

114
00:09:44,480 --> 00:09:59,120
But I have so many questions about the deep dive a little bit in our, we have the Kubernetes.

115
00:09:59,120 --> 00:10:01,360
It's such a cool service.

116
00:10:01,360 --> 00:10:11,400
We have to use function, the Kubernetes app service container service box.

117
00:10:11,400 --> 00:10:22,120
I think for normal business guys, we don't have a plan, what we use when, what should

118
00:10:22,120 --> 00:10:31,480
the, should we discuss first as not technical guys?

119
00:10:31,480 --> 00:10:35,680
What will your solution need to be doing, first of all?

120
00:10:35,680 --> 00:10:44,600
And with that, I mean, will it be an internal or an external application?

121
00:10:44,600 --> 00:10:46,360
How many users does it have?

122
00:10:46,360 --> 00:10:47,760
Does it need to scale?

123
00:10:47,760 --> 00:10:51,800
Does it need to be resilient?

124
00:10:51,800 --> 00:10:55,200
Is it something long lift, short lift?

125
00:10:55,200 --> 00:11:00,240
Is it something which is impacting a lot of users?

126
00:11:00,240 --> 00:11:03,240
And also is it impacting users worldwide?

127
00:11:03,240 --> 00:11:10,240
Even if you're using an internal application, it still could be that you have, let's say,

128
00:11:10,240 --> 00:11:16,120
a sub, a subsidiary in, in, in, in, in the United States, you have a subsidiary in, in, in,

129
00:11:16,120 --> 00:11:24,680
in, in, in, in, in, somewhere in Europe, and they all need to work together.

130
00:11:24,680 --> 00:11:30,120
Look at the way that those requirements need to be built up and typically that's a business

131
00:11:30,120 --> 00:11:31,640
approach.

132
00:11:31,640 --> 00:11:34,400
What compliance do we need to uphold?

133
00:11:34,400 --> 00:11:36,400
What's our rock time?

134
00:11:36,400 --> 00:11:40,560
As a lace, as a lose, our deals, our bureaus?

135
00:11:40,560 --> 00:11:43,280
And those are things that can be defined even without an architecture.

136
00:11:43,280 --> 00:11:47,920
There's a requirements that you need to set, and you need to take a look at it.

137
00:11:47,920 --> 00:11:57,400
The second thing is, how can we live with an application if you pull the plug on it?

138
00:11:57,400 --> 00:12:00,240
What will happen?

139
00:12:00,240 --> 00:12:04,560
Who will be responsible, but also who will be impacted and will they start yelling at

140
00:12:04,560 --> 00:12:07,560
you when you do pull the plug?

141
00:12:07,560 --> 00:12:12,520
And there's a, there's a, there's a brother from another mother, as I like to call him.

142
00:12:12,520 --> 00:12:19,720
He's called Derek Martin and he has a great series of, of sessions that he does when he

143
00:12:19,720 --> 00:12:20,960
goes to conferences.

144
00:12:20,960 --> 00:12:26,920
The one whose session is, is about resilience and his first rule of thumb, and I also live

145
00:12:26,920 --> 00:12:31,520
by that is never take a dependency on a single region.

146
00:12:31,520 --> 00:12:33,520
Always make a design for multiple regions.

147
00:12:33,520 --> 00:12:39,840
So that way you always have a plan B and even a plan C or a plan D because anything can happen.

148
00:12:39,840 --> 00:12:40,840
Yeah.

149
00:12:40,840 --> 00:12:45,000
A region can go down or you want to deploy quickly into another region.

150
00:12:45,000 --> 00:12:46,000
Capacity can go down.

151
00:12:46,000 --> 00:12:48,000
It's, it's a given.

152
00:12:48,000 --> 00:12:50,120
I mean, we have to fight it.

153
00:12:50,120 --> 00:12:55,120
Capacity issues are existing in any cloud, not on Microsoft cloud in any cloud.

154
00:12:55,120 --> 00:13:00,240
Although with Microsoft, it's more visible apparently than down, for instance, with AWS or

155
00:13:00,240 --> 00:13:05,240
Google, although they have the same capacity issues that just don't want to admit it.

156
00:13:05,240 --> 00:13:06,240
Yeah.

157
00:13:06,240 --> 00:13:10,520
The reason why is because of the, of the growth of everything, which is AI, which is taking

158
00:13:10,520 --> 00:13:15,880
more power and the traditional data centers cannot follow anymore with all the power that

159
00:13:15,880 --> 00:13:17,200
they need.

160
00:13:17,200 --> 00:13:21,560
So they cannot expand their data centers anymore because of the powers capacity to give you

161
00:13:21,560 --> 00:13:27,880
an idea. One of the reasons why, why Microsoft in Western Europe cannot deploy anymore is

162
00:13:27,880 --> 00:13:33,160
typically because there's a lot of power there because the Netherlands have power issue.

163
00:13:33,160 --> 00:13:34,280
And that's the same everywhere.

164
00:13:34,280 --> 00:13:41,840
So and everybody's working, working their, their butts off to actually make it work and

165
00:13:41,840 --> 00:13:44,080
seeing that there's enough capacity.

166
00:13:44,080 --> 00:13:49,760
But it's, it's actually the, the, the, the, the, the victims of their own success.

167
00:13:49,760 --> 00:13:51,520
Yeah.

168
00:13:51,520 --> 00:13:56,200
And when we never planned on, on that kind of a success, although we have been growing all

169
00:13:56,200 --> 00:14:01,960
these, this cloud business over the last 10 years, like him, muslims, I mean, we see year

170
00:14:01,960 --> 00:14:09,200
over year growth in most of the cloud providers for about 30%, which is huge, right?

171
00:14:09,200 --> 00:14:17,200
So, and that brings us to do, if you need to start designing, what are the questions that

172
00:14:17,200 --> 00:14:19,360
you need to be asking your business?

173
00:14:19,360 --> 00:14:24,880
Without the questions that you need to be asking also, not only your business, maybe also

174
00:14:24,880 --> 00:14:29,560
your users or even your developers in your, in your system managers, how do, how do they

175
00:14:29,560 --> 00:14:30,560
want to cope with it?

176
00:14:30,560 --> 00:14:32,560
Do they want to do it 24/7 support?

177
00:14:32,560 --> 00:14:39,000
Do they want to do a full high NCICD and they want to be bothered anymore with deployments

178
00:14:39,000 --> 00:14:40,000
on Friday evening?

179
00:14:40,000 --> 00:14:41,040
These kind of things.

180
00:14:41,040 --> 00:14:43,440
Those are the questions that you need to be asking.

181
00:14:43,440 --> 00:14:45,880
And that's also our signature.

182
00:14:45,880 --> 00:14:49,160
If you're moving to the cloud, what kind of art do you need to follow?

183
00:14:49,160 --> 00:14:50,760
You know, there's eight alt, right?

184
00:14:50,760 --> 00:14:57,080
And one of them is we platform you, so have re-hosting, but also do we need to move it

185
00:14:57,080 --> 00:14:58,680
to the cloud?

186
00:14:58,680 --> 00:15:05,800
Can we know just like, put it into the, into the grave for us and just retire it?

187
00:15:05,800 --> 00:15:08,800
It's all these things that you need to be asking yourself.

188
00:15:08,800 --> 00:15:11,960
Yeah.

189
00:15:11,960 --> 00:15:21,400
I think some people are not really family with these, I think, bospers at your IPO.

190
00:15:21,400 --> 00:15:27,360
I think at your, it's really famous for my daughter.

191
00:15:27,360 --> 00:15:29,120
She's here for years.

192
00:15:29,120 --> 00:15:33,200
How long can I stay online?

193
00:15:33,200 --> 00:15:35,880
At your, I don't know.

194
00:15:35,880 --> 00:15:41,880
It's more for the, how would you explain this in?

195
00:15:41,880 --> 00:15:45,400
For a deep dive, how will you explain this?

196
00:15:45,400 --> 00:15:47,920
So RTN, right?

197
00:15:47,920 --> 00:15:49,440
Was that the question?

198
00:15:49,440 --> 00:15:50,440
Yeah.

199
00:15:50,440 --> 00:15:56,440
So yeah, I'm going to wait for a second.

200
00:15:56,440 --> 00:15:59,320
So because I heard the fire trick or what is it?

201
00:15:59,320 --> 00:16:00,320
Nambless?

202
00:16:00,320 --> 00:16:01,320
I don't know.

203
00:16:01,320 --> 00:16:07,520
No, but if we talk about RTN, so we have a recovery time objective and we have a recovery

204
00:16:07,520 --> 00:16:08,520
point of jitter.

205
00:16:08,520 --> 00:16:13,320
So how much, how fast do you want to get online?

206
00:16:13,320 --> 00:16:15,760
RTO.

207
00:16:15,760 --> 00:16:17,400
How do you want to get back online?

208
00:16:17,400 --> 00:16:19,440
And I know my data loss is acceptable.

209
00:16:19,440 --> 00:16:20,440
That's RPO.

210
00:16:20,440 --> 00:16:23,240
So these are the points that you need to define.

211
00:16:23,240 --> 00:16:29,320
So if you have an accessible, an acceptable loss of four hours for your business because it's,

212
00:16:29,320 --> 00:16:33,480
it's not about or the management, but it's more about monitoring of something and it's

213
00:16:33,480 --> 00:16:36,240
not really needed that, that is real time data.

214
00:16:36,240 --> 00:16:39,320
It can actually move along without having that.

215
00:16:39,320 --> 00:16:41,400
Then it's fine.

216
00:16:41,400 --> 00:16:49,320
But if it's something critical, let's say, for first responders, the fire trick came, came

217
00:16:49,320 --> 00:16:51,920
in as an example.

218
00:16:51,920 --> 00:16:57,320
Then an RTO, which is really fast, is necessary, like five minutes, ten minutes, these kind

219
00:16:57,320 --> 00:16:58,320
of things.

220
00:16:58,320 --> 00:17:09,320
But these are design decisions that will also help you ground and build your solution with,

221
00:17:09,320 --> 00:17:15,040
because that's an necessity for resilient application too.

222
00:17:15,040 --> 00:17:20,240
We only think about the architecture decisions.

223
00:17:20,240 --> 00:17:26,480
What's a really famous topic in what most you can, you can write it or say a little bit

224
00:17:26,480 --> 00:17:37,800
about this identity, network, data architecture, region surgery, platform engineer decisions,

225
00:17:37,800 --> 00:17:43,760
ten design, what, what, what, what did you think?

226
00:17:43,760 --> 00:17:49,320
To me, architecture is, is all about those because if you do, if you're building a solution,

227
00:17:49,320 --> 00:17:50,320
you need all of them.

228
00:17:50,320 --> 00:17:56,040
But it depends on what kind of an architect you are and there's multiple levels.

229
00:17:56,040 --> 00:18:02,640
I see myself as a solution slash enterprise architect because I also need them taking,

230
00:18:02,640 --> 00:18:06,160
taking to account integrations that are there.

231
00:18:06,160 --> 00:18:13,280
Dependencies that you have on extended systems or auxiliary systems like, let's say, for instance,

232
00:18:13,280 --> 00:18:17,200
you need to integrate within a, within a AP system.

233
00:18:17,200 --> 00:18:20,800
That doesn't mean that you need to be running it, but you need to know that if your system

234
00:18:20,800 --> 00:18:26,360
goes down, it might be a dependency on SAP, which will bring down the SAP flow, which will

235
00:18:26,360 --> 00:18:31,880
probably break down your invoicing or your order management or whatever kind of part that

236
00:18:31,880 --> 00:18:34,200
you that you're relying on that.

237
00:18:34,200 --> 00:18:38,240
So that's an important thing.

238
00:18:38,240 --> 00:18:43,240
But it also comes down to, I'm going to give you an example.

239
00:18:43,240 --> 00:18:47,360
If you, let's say, for instance, I'm going to give a real live example of something that

240
00:18:47,360 --> 00:18:50,080
I see every single year.

241
00:18:50,080 --> 00:18:51,520
It's taxes.

242
00:18:51,520 --> 00:18:57,440
So we have to, you have to push our taxes into, yeah, I don't know, you're in Germany, right?

243
00:18:57,440 --> 00:19:00,120
Or are you in Switzerland?

244
00:19:00,120 --> 00:19:01,440
And they are in Germany.

245
00:19:01,440 --> 00:19:02,440
I'm going to come back.

246
00:19:02,440 --> 00:19:03,440
Let's put that down.

247
00:19:03,440 --> 00:19:04,440
Okay.

248
00:19:04,440 --> 00:19:09,760
So I don't know how you do it in Germany, but each year we need to fill out our taxes

249
00:19:09,760 --> 00:19:13,640
on a certain site and we, and they have all that data.

250
00:19:13,640 --> 00:19:18,960
But everybody typically waits to the last few days and then the system breaks because

251
00:19:18,960 --> 00:19:24,440
there's not enough scale, there's not enough space, there's, there's timeouts because of

252
00:19:24,440 --> 00:19:26,240
things being troubled.

253
00:19:26,240 --> 00:19:28,240
I know it.

254
00:19:28,240 --> 00:19:34,560
So that's a typical use case of, of, of how a cloud application should work.

255
00:19:34,560 --> 00:19:42,200
And one of the things that always fails is our Belgian government has a, it's own identity

256
00:19:42,200 --> 00:19:43,200
system.

257
00:19:43,200 --> 00:19:50,760
It's an EID and you need to like, jacket your EID cart into a car, and, and, and, and that

258
00:19:50,760 --> 00:19:51,760
always gets worked.

259
00:19:51,760 --> 00:19:57,400
Especially on the, and it doesn't, and it doesn't always work because it always gives you

260
00:19:57,400 --> 00:20:02,920
timeouts and then, and then you need to update it and does work and then, and it fails again.

261
00:20:02,920 --> 00:20:07,160
And then once you get in, the token doesn't get recognized on the site and there's always

262
00:20:07,160 --> 00:20:08,160
something.

263
00:20:08,160 --> 00:20:12,800
Now, there's a small company, and there's a small company, which is now expanding out

264
00:20:12,800 --> 00:20:15,600
throughout Europe, a company which I'm really proud of.

265
00:20:15,600 --> 00:20:23,200
It's called It's Me and they build a, a digital identity based upon a phone verification and

266
00:20:23,200 --> 00:20:29,480
I mean numbers and they use the identity of your EID or your banking account because a banking

267
00:20:29,480 --> 00:20:35,520
account is a validated identity, a determinator.

268
00:20:35,520 --> 00:20:37,280
And they can really scale.

269
00:20:37,280 --> 00:20:40,800
So they don't, you don't need to validate against the state anymore.

270
00:20:40,800 --> 00:20:45,240
They do the validation for you, but they really scale and their application is existing into

271
00:20:45,240 --> 00:20:52,200
a, into a cloud building solution, in a, to a cloud build solution, which can on each of

272
00:20:52,200 --> 00:20:55,600
the levels of the application can scale and fail independently.

273
00:20:55,600 --> 00:21:00,440
So their backend will not be troubled by the front end and so on.

274
00:21:00,440 --> 00:21:05,680
Bringing that back then to the texting, once you authenticate, you feel again that the

275
00:21:05,680 --> 00:21:11,680
state's application is, is terrible because then you press the submit, there's a process

276
00:21:11,680 --> 00:21:13,840
being filed, but it's in real time.

277
00:21:13,840 --> 00:21:17,040
They don't do the a sync part and that's the thing.

278
00:21:17,040 --> 00:21:21,760
You need to be looking into how many of my features can I build a sync?

279
00:21:21,760 --> 00:21:23,920
How many of my features do I need in real time?

280
00:21:23,920 --> 00:21:29,560
It's all these design decisions and that's an application part, but also your, with that

281
00:21:29,560 --> 00:21:34,120
application part, you also need to design of the infrastructure behind it because a front

282
00:21:34,120 --> 00:21:39,080
end can easily, if you press a button multiple times and it does it immediately respond,

283
00:21:39,080 --> 00:21:42,360
will it break your front end?

284
00:21:42,360 --> 00:21:43,360
Right?

285
00:21:43,360 --> 00:21:49,440
But also, if you press it once and it keeps rolling, keeps doing pulse to the backend, that's

286
00:21:49,440 --> 00:21:52,080
a bad design decision because it's waiting for the backend.

287
00:21:52,080 --> 00:21:55,960
It should be loosely coupled and should be just pushing something into a queue and event

288
00:21:55,960 --> 00:22:00,000
source or whatever and making sure that it comes back like that.

289
00:22:00,000 --> 00:22:05,000
That way each of your layers of your application will be able to fail and scale independently

290
00:22:05,000 --> 00:22:08,480
and making sure that your application is resilient.

291
00:22:08,480 --> 00:22:13,080
After that, multiple regions and you have a full resilient application, but most of the

292
00:22:13,080 --> 00:22:16,240
applications, especially with the state, are not built that way.

293
00:22:16,240 --> 00:22:17,240
Why?

294
00:22:17,240 --> 00:22:18,760
Because we're running it in our own data center.

295
00:22:18,760 --> 00:22:21,760
It needs to be, it needs to be in our country.

296
00:22:21,760 --> 00:22:24,280
Germans are really good at that, by the way.

297
00:22:24,280 --> 00:22:25,280
Yeah.

298
00:22:25,280 --> 00:22:26,280
And I get that.

299
00:22:26,280 --> 00:22:29,680
You want to stay sovereign and you want to make sure that you have your data under

300
00:22:29,680 --> 00:22:31,520
control.

301
00:22:31,520 --> 00:22:35,320
But then you don't need to be starting to pretend that you're doing cloud applications.

302
00:22:35,320 --> 00:22:38,880
Then you're just doing hosting of these applications yourself and that's not the way you should

303
00:22:38,880 --> 00:22:40,520
be looking at it.

304
00:22:40,520 --> 00:22:44,880
If you're talking about applications which are being run in Europe, everything in Europe

305
00:22:44,880 --> 00:22:49,840
is valid as a landing zone for European applications.

306
00:22:49,840 --> 00:22:53,760
And sometimes you have these exceptions like in France, medical data and it's a stay in

307
00:22:53,760 --> 00:22:55,760
France in Germany.

308
00:22:55,760 --> 00:22:59,640
There's also certain data that needs to be in Germany.

309
00:22:59,640 --> 00:23:00,880
Belgium is just the same thing.

310
00:23:00,880 --> 00:23:06,920
We only have one type of application that needs to be on Belgium soil, especially in the

311
00:23:06,920 --> 00:23:12,000
cloud, but also in a data center that is gambling data because police need to be able to jack

312
00:23:12,000 --> 00:23:17,480
out the discs to see where there's white color crime involved and where there's fraud involved.

313
00:23:17,480 --> 00:23:22,400
So it's all these things that come to play.

314
00:23:22,400 --> 00:23:26,960
And when we were designing, a lot of developers don't think that way.

315
00:23:26,960 --> 00:23:31,880
They think about, "I want to make it beautiful and want to make it like technical and all

316
00:23:31,880 --> 00:23:35,360
display with that technology and do it by that."

317
00:23:35,360 --> 00:23:40,620
But if you then push it to production at scale, you see things break because they never

318
00:23:40,620 --> 00:23:43,000
thought of these things at scale.

319
00:23:43,000 --> 00:23:49,640
They never came into contact with things like distributed computing or event sourcing

320
00:23:49,640 --> 00:23:51,120
for that matter.

321
00:23:51,120 --> 00:23:54,360
So today we have a lot of tools that can help us with that.

322
00:23:54,360 --> 00:24:00,680
Think about tooling like Dapper, think about tooling like Radius, think about tooling like Aspire,

323
00:24:00,680 --> 00:24:09,320
which allows us to actually instantiate certain things in more instances, more pots, more

324
00:24:09,320 --> 00:24:14,560
services that can roll with those resilience factors.

325
00:24:14,560 --> 00:24:17,320
So that's basically what it comes down to.

326
00:24:17,320 --> 00:24:27,760
Yeah, I lose by on your LinkedIn and other profiles around the net and something you

327
00:24:27,760 --> 00:24:32,720
are really work on or your famous for is recent architecture.

328
00:24:32,720 --> 00:24:39,200
And I have some questions about the recent architecture.

329
00:24:39,200 --> 00:24:43,920
What is the difference between a variable?

330
00:24:43,920 --> 00:24:48,720
What is the variable and what is the difference?

331
00:24:48,720 --> 00:24:51,880
Yeah, it's a real.

332
00:24:51,880 --> 00:24:53,400
What did you think?

333
00:24:53,400 --> 00:24:59,920
So how available and resilient they go hand in hand and still they exist separately.

334
00:24:59,920 --> 00:25:07,640
So I always compare resilient with look at a tropical island.

335
00:25:07,640 --> 00:25:12,380
And what's typical for a tropical island is always palm trees at the beach and a palm

336
00:25:12,380 --> 00:25:14,620
tree is a remarkable thing.

337
00:25:14,620 --> 00:25:20,820
A palm tree is like you see them hanging over the water and they like high and they always

338
00:25:20,820 --> 00:25:23,860
go in bending in a bending shape.

339
00:25:23,860 --> 00:25:29,060
Now let's imagine that you are on that island in a tropical storm.

340
00:25:29,060 --> 00:25:31,260
El Nino for instance, passes there.

341
00:25:31,260 --> 00:25:37,860
What happens is all those palm trees when wind arises, all those palm trees they typically

342
00:25:37,860 --> 00:25:39,900
hit the floor.

343
00:25:39,900 --> 00:25:44,700
And they stay at the floor while the wind is there while the storm is still happening.

344
00:25:44,700 --> 00:25:50,200
At a certain point in time the storm stops and what happens is you see palm trees arise

345
00:25:50,200 --> 00:25:51,700
again like very naturally.

346
00:25:51,700 --> 00:25:55,540
Like you come back up.

347
00:25:55,540 --> 00:25:59,820
You can imagine it like in cartoons.

348
00:25:59,820 --> 00:26:01,580
That's resilience.

349
00:26:01,580 --> 00:26:02,820
It doesn't break.

350
00:26:02,820 --> 00:26:06,220
The palm tree stays a palm tree.

351
00:26:06,220 --> 00:26:11,700
And afterwards you can still see that a palm tree didn't or almost didn't suffer any damage.

352
00:26:11,700 --> 00:26:12,700
Right?

353
00:26:12,700 --> 00:26:20,720
High available means that it will break but you have a secondary palm tree that will take

354
00:26:20,720 --> 00:26:24,660
over its place but it's on another spot.

355
00:26:24,660 --> 00:26:30,740
So you still have, let's say that we just replant that palm tree afterwards.

356
00:26:30,740 --> 00:26:32,700
That's high available.

357
00:26:32,700 --> 00:26:37,260
Does it mean that it will be resilience?

358
00:26:37,260 --> 00:26:39,060
No, because there is a difference there.

359
00:26:39,060 --> 00:26:44,180
But resilience will also make sure that it will not go down but also there is no issue

360
00:26:44,180 --> 00:26:48,780
with regions that it will also be able to cope with the load that it will be able to cope

361
00:26:48,780 --> 00:26:53,340
with the scale that it will be able to cope with any mechanisms that happen underneath.

362
00:26:53,340 --> 00:26:58,220
That it can actually function while there is application errors and so on and so on.

363
00:26:58,220 --> 00:27:03,740
But available just means that we can ping it and that we can still have a available somehow.

364
00:27:03,740 --> 00:27:06,940
And if there is an issue we can just fail over.

365
00:27:06,940 --> 00:27:08,940
That's the biggest difference.

366
00:27:08,940 --> 00:27:11,340
I like this.

367
00:27:11,340 --> 00:27:15,940
I often hear one phrase.

368
00:27:15,940 --> 00:27:20,620
It's, yeah, it's, it's, it's highly available.

369
00:27:20,620 --> 00:27:24,620
I heard so often.

370
00:27:24,620 --> 00:27:34,740
I used to statement a little bit dangerous and where does Microsoft responsible end and

371
00:27:34,740 --> 00:27:40,860
where does the customer respond the begin from your perspective.

372
00:27:40,860 --> 00:27:47,020
So there is indeed the false premise that cloud is high available and high resilient.

373
00:27:47,020 --> 00:27:49,500
Now high resilience, that's a given.

374
00:27:49,500 --> 00:27:53,580
I mean, you have all these opportunities and all these solutions to make it out of

375
00:27:53,580 --> 00:27:55,580
the box resilient as possible.

376
00:27:55,580 --> 00:28:01,020
And think about how you can deploy Vm in multiple zones, for instance, or any compute for

377
00:28:01,020 --> 00:28:02,740
that matter.

378
00:28:02,740 --> 00:28:04,740
Think about storage accounts.

379
00:28:04,740 --> 00:28:10,660
A storage account has automatically underneath the hood three copies within a data center.

380
00:28:10,660 --> 00:28:17,220
And it will pass on on different, on different parts of the stamp to make sure that data can

381
00:28:17,220 --> 00:28:21,660
be consistently written to that stamp, to that data stamp.

382
00:28:21,660 --> 00:28:26,980
So Microsoft does some work, but that doesn't mean that your application will be fully resilient

383
00:28:26,980 --> 00:28:31,220
or fully high available because there's still some design patterns that you need to adhere

384
00:28:31,220 --> 00:28:32,220
to.

385
00:28:32,220 --> 00:28:38,940
And if you don't deploy your Vm, if you choose not to deploy your Vm in multiple zones, it

386
00:28:38,940 --> 00:28:39,940
will not happen.

387
00:28:39,940 --> 00:28:45,060
It will just be deployed in one zone or in multiple regions for that matter or in multi

388
00:28:45,060 --> 00:28:47,660
zone or regional regions and so on and so on.

389
00:28:47,660 --> 00:28:54,260
So these are design decisions. They come with some challenges, typically it's costs and

390
00:28:54,260 --> 00:29:00,740
also operations because you need to monitor them in a different way.

391
00:29:00,740 --> 00:29:05,500
But they also can help you speed up.

392
00:29:05,500 --> 00:29:12,900
And as I said earlier, where does the responsibility of Microsoft and well, once you start deploying,

393
00:29:12,900 --> 00:29:15,620
you choose to deploy only one in one region.

394
00:29:15,620 --> 00:29:22,960
Yeah, the break Microsoft will not be held responsible for your application to move, unless you

395
00:29:22,960 --> 00:29:27,580
say, of course, I want an auto failover, let's say with my Cosmos DB to another region

396
00:29:27,580 --> 00:29:35,660
because I designed it that way or where I have a read only copy there for that matter.

397
00:29:35,660 --> 00:29:42,220
These are things that you can have in your own control because you designed them that

398
00:29:42,220 --> 00:29:43,620
way.

399
00:29:43,620 --> 00:29:52,100
But out of the box, a lot of these services are not being enabled by a checkbox out of the

400
00:29:52,100 --> 00:29:53,300
box.

401
00:29:53,300 --> 00:29:54,660
You need to choose that yourself.

402
00:29:54,660 --> 00:29:57,700
And some some services do have that.

403
00:29:57,700 --> 00:30:04,140
They do an underneath under the layer scaling or they do an under the layer high available

404
00:30:04,140 --> 00:30:06,540
thing.

405
00:30:06,540 --> 00:30:11,060
Mostly if they're built upon upon containers, you will actually see that that they will

406
00:30:11,060 --> 00:30:14,540
have the Kubernetes lifecycle and needed.

407
00:30:14,540 --> 00:30:19,580
So these things do exist, but there's still a lot of decisions that you need to take your

408
00:30:19,580 --> 00:30:20,580
own.

409
00:30:20,580 --> 00:30:26,060
And you need to have some certain design patterns in fact in in mind like the circuit break

410
00:30:26,060 --> 00:30:30,420
if for instance, if you have a circuit break it does.

411
00:30:30,420 --> 00:30:34,020
Then if one service goes down, that doesn't mean that your entire solution need to be going

412
00:30:34,020 --> 00:30:35,020
down.

413
00:30:35,020 --> 00:30:39,540
Those were the examples that I already gave earlier in the discussion.

414
00:30:39,540 --> 00:30:41,900
So it's these things that you need to keep in mind.

415
00:30:41,900 --> 00:30:45,620
And luckily Microsoft has good guidance there.

416
00:30:45,620 --> 00:30:49,820
It's called it's called the bellar should take the framework where you can have design

417
00:30:49,820 --> 00:30:54,100
patterns and you can have all these design patterns like singleton circuit break or you

418
00:30:54,100 --> 00:30:58,900
cannot imagine how many they are, but they're all like well explained.

419
00:30:58,900 --> 00:31:04,860
And if you then combine with things like cloud adoption framework, you can actually build

420
00:31:04,860 --> 00:31:07,660
up on a full resilient application.

421
00:31:07,660 --> 00:31:08,660
Is that sufficient?

422
00:31:08,660 --> 00:31:13,620
No, you also need to be having some experience in how to do these things because documentation

423
00:31:13,620 --> 00:31:17,820
is one thing, but you also need to be monitoring that in a correct way.

424
00:31:17,820 --> 00:31:21,980
You also need to make sure that you understand what you're doing.

425
00:31:21,980 --> 00:31:23,380
Following documentation is not enough.

426
00:31:23,380 --> 00:31:28,660
You also need to think about these things and what happens if.

427
00:31:28,660 --> 00:31:34,220
And then there's an additional part because we can be prepared for that and we can build

428
00:31:34,220 --> 00:31:35,220
like that.

429
00:31:35,220 --> 00:31:39,540
Microsoft can offer us that, but are we sure now?

430
00:31:39,540 --> 00:31:44,020
And then there's one thing to do that's testing these things, right?

431
00:31:44,020 --> 00:31:46,580
And then you start doing like a what will happen?

432
00:31:46,580 --> 00:31:49,740
Will my resilience break if I put a lot of stress on it?

433
00:31:49,740 --> 00:31:54,820
If I put a lot of load on it and you can do things like playwright service, testing or

434
00:31:54,820 --> 00:31:56,660
Azure load testing.

435
00:31:56,660 --> 00:31:58,540
But what if something else happens?

436
00:31:58,540 --> 00:32:01,060
What if a region goes down?

437
00:32:01,060 --> 00:32:03,140
Well, you cannot just yak away the service.

438
00:32:03,140 --> 00:32:09,540
Then you go looking at things like Azure Chaos Studio and do fault injections, home production,

439
00:32:09,540 --> 00:32:12,780
on staging, whatever, whatever works for you.

440
00:32:12,780 --> 00:32:15,220
And use these as fire drills.

441
00:32:15,220 --> 00:32:17,500
Use these as exercises.

442
00:32:17,500 --> 00:32:20,420
Make sure that these things can work for you.

443
00:32:20,420 --> 00:32:23,580
And this is how resilience is being built up.

444
00:32:23,580 --> 00:32:26,500
It's not just one thing that you need to keep in mind of.

445
00:32:26,500 --> 00:32:30,860
It's the some of all fears that you need to combine.

446
00:32:30,860 --> 00:32:35,460
And bring in that to a single solution in one single architecture.

447
00:32:35,460 --> 00:32:40,020
So, so my, so home like let me more of 11 do it here.

448
00:32:40,020 --> 00:32:45,860
And I pay for the Azure SRA for the service level agreement.

449
00:32:45,860 --> 00:32:54,780
So I think I pay a lot of money for it.

450
00:32:54,780 --> 00:32:57,780
What?

451
00:32:57,780 --> 00:33:02,260
What's my misunderstanding about the SAs?

452
00:33:02,260 --> 00:33:06,860
So the service, the service level, the service agreement that you have with Microsoft is more

453
00:33:06,860 --> 00:33:10,220
about the support that you have.

454
00:33:10,220 --> 00:33:14,780
Preemptive or reactive, right?

455
00:33:14,780 --> 00:33:19,660
So they can do proactive support like, hey, they're going to go over with you over research,

456
00:33:19,660 --> 00:33:20,660
for instance.

457
00:33:20,660 --> 00:33:23,620
So you're going to see what the incident reports look like.

458
00:33:23,620 --> 00:33:24,780
How you can report that.

459
00:33:24,780 --> 00:33:31,660
But they will also go with you over your solution and take a look at, wow, what you built

460
00:33:31,660 --> 00:33:36,380
here, you should, and they're going to give the same advice like I do.

461
00:33:36,380 --> 00:33:38,100
But then it's grounded from their art.

462
00:33:38,100 --> 00:33:40,900
So it's actually hard evidence that they're going to give you and you're going to give

463
00:33:40,900 --> 00:33:43,500
you some reports and some assessments.

464
00:33:43,500 --> 00:33:48,620
And typically this comes with things like unify or support contracts where you can actually

465
00:33:48,620 --> 00:33:51,780
sign them all.

466
00:33:51,780 --> 00:33:55,780
So are they useful?

467
00:33:55,780 --> 00:33:57,100
Yes, the certain level they are useful.

468
00:33:57,100 --> 00:34:01,380
You always need them if you get good architics and you have a lot of credits and you have a

469
00:34:01,380 --> 00:34:03,540
lot of money also.

470
00:34:03,540 --> 00:34:06,500
You can just design these things as well.

471
00:34:06,500 --> 00:34:11,700
Use things like advise to have you well informed about the things that there are, but also take

472
00:34:11,700 --> 00:34:18,260
a look at service retirements, service updates and so on and so on because they also play.

473
00:34:18,260 --> 00:34:24,340
Microsoft actually, with their service agreement, they actually do this exactly, they do preemptive

474
00:34:24,340 --> 00:34:29,980
calls, they do proactive calls saying, hey, look, these things will happen.

475
00:34:29,980 --> 00:34:32,740
Maybe it's time to move into V next.

476
00:34:32,740 --> 00:34:34,460
Maybe it's time to start designing this.

477
00:34:34,460 --> 00:34:38,420
And by the way, you have a year and a half now to do that because in the year and a half,

478
00:34:38,420 --> 00:34:39,940
this will change.

479
00:34:39,940 --> 00:34:48,340
So this is where that kind of support comes from and where that kind of a service level is

480
00:34:48,340 --> 00:34:50,540
coming from at Microsoft.

481
00:34:50,540 --> 00:34:54,300
So it's not like that they are going to do a full switching.

482
00:34:54,300 --> 00:34:58,180
There will be a support when there's when there's when shit hits the fan basically and

483
00:34:58,180 --> 00:35:04,220
because they are being there for severity a severity, be severity, see, we'll have dedicated

484
00:35:04,220 --> 00:35:06,060
support people on that.

485
00:35:06,060 --> 00:35:08,260
Yes, that's true.

486
00:35:08,260 --> 00:35:16,500
But they will not guarantee you that they will monitor your service and put it up 24/7.

487
00:35:16,500 --> 00:35:18,580
That's your responsibility.

488
00:35:18,580 --> 00:35:27,060
But they will help you try to maintain that, like supporting you in that in that operation.

489
00:35:27,060 --> 00:35:34,820
I think in a little bit of we have the tailing of the guys, they have the energy or the RPO

490
00:35:34,820 --> 00:35:43,780
and we have as a side, I think there are more people, I call it the business reality.

491
00:35:43,780 --> 00:35:55,060
So how often do companies claim they need a zero downtime with, yeah, understanding what

492
00:35:55,060 --> 00:35:59,300
the requirements really cost?

493
00:35:59,300 --> 00:36:00,300
We're talking about the 9-0.

494
00:36:00,300 --> 00:36:01,820
That's always funny to see.

495
00:36:01,820 --> 00:36:07,180
I actually had it once, a partner coming up, I was working at Microsoft and a partner coming

496
00:36:07,180 --> 00:36:10,060
up and I, we need an Isla of 100%.

497
00:36:10,060 --> 00:36:12,820
And then we're like, good luck with that.

498
00:36:12,820 --> 00:36:15,740
You will never have an Isla of 100%.

499
00:36:15,740 --> 00:36:17,260
Not even at Microsoft.

500
00:36:17,260 --> 00:36:23,180
Not even if you're doing it yourself, you will never have 100%.

501
00:36:23,180 --> 00:36:29,540
Because the Isla is the product of all the Isla is of the different services that you have.

502
00:36:29,540 --> 00:36:36,020
Some services have 99.9, other services have 99.99, a couple of services are reaching up

503
00:36:36,020 --> 00:36:38,060
to five nines.

504
00:36:38,060 --> 00:36:42,980
But if you then have five nines, time, three nines, you only will end up with four nines.

505
00:36:42,980 --> 00:36:48,700
So that narrows down your Isla time and that narrows down your up time, your 100%.

506
00:36:48,700 --> 00:36:49,700
Is that a bad thing?

507
00:36:49,700 --> 00:36:51,700
No.

508
00:36:51,700 --> 00:36:56,780
But people need to be aware of their selling their service at 100% is a lay.

509
00:36:56,780 --> 00:36:59,700
They will shoot themselves in the photo already before they even sell it.

510
00:36:59,700 --> 00:37:09,700
[BLANK_AUDIO]

511
00:37:09,700 --> 00:37:19,700
[BLANK_AUDIO]

512
00:37:19,700 --> 00:37:29,700
[BLANK_AUDIO]

513
00:37:29,700 --> 00:37:39,700
[BLANK_AUDIO]

514
00:37:39,700 --> 00:37:49,700
[BLANK_AUDIO]

515
00:37:49,700 --> 00:37:59,700
[BLANK_AUDIO]

516
00:37:59,700 --> 00:38:09,700
[BLANK_AUDIO]

517
00:38:09,700 --> 00:38:11,700
[BLANK_AUDIO]

518
00:38:11,700 --> 00:38:22,060
Yeah, we have these are dudes, they say 100% are online and we have the outside, they have

519
00:38:22,060 --> 00:38:25,700
these SLA's, the service, that would be mean.

520
00:38:25,700 --> 00:38:36,220
What, Mike did you think about and can we reach the 100% really?

521
00:38:36,220 --> 00:38:38,780
Never, that's the biggest joke ever.

522
00:38:38,780 --> 00:38:44,540
People going up with that that you can have 100% isla, good luck with that because you always

523
00:38:44,540 --> 00:38:47,460
have component which is less than that isla.

524
00:38:47,460 --> 00:38:53,220
Don't forget that your isla is always calculated to the products of all these isla is combined.

525
00:38:53,220 --> 00:38:58,580
It's funny when I was murking at Microsoft at one partner, we were actually requested at

526
00:38:58,580 --> 00:38:59,820
a certain point in time.

527
00:38:59,820 --> 00:39:05,260
Yeah, we want a service which is, we want to build our service 100% isla up time towards

528
00:39:05,260 --> 00:39:06,260
our customers.

529
00:39:06,260 --> 00:39:10,180
You know, and like, good luck with that because if you're running on Azure, the highest capable

530
00:39:10,180 --> 00:39:18,500
service that you can have at, at, as a lay is 99.99, so five nines, that's the highest

531
00:39:18,500 --> 00:39:19,500
that you can reach.

532
00:39:19,500 --> 00:39:22,180
And you're like, oh, really?

533
00:39:22,180 --> 00:39:27,940
So combined that with an additional service, let's say 99.99, you're all of a sudden

534
00:39:27,940 --> 00:39:34,340
are back to three nines and so on and so on and they were, they were like really surprised.

535
00:39:34,340 --> 00:39:38,780
So it's a false premise that you can have 100% up time that's never given.

536
00:39:38,780 --> 00:39:45,660
But you want to make sure that you can be as close as possible and that you have mitigations

537
00:39:45,660 --> 00:39:55,380
and fallback scenarios to avoid that, let's say, 0.0, X percent.

538
00:39:55,380 --> 00:40:00,460
So to make sure that you don't fail, that you don't have a downtime, that you don't have

539
00:40:00,460 --> 00:40:06,740
to do a failover and it does failover, just make sure that you stay calm.

540
00:40:06,740 --> 00:40:12,060
And you have that plan B checked that it works, that you tested it before and then we're

541
00:40:12,060 --> 00:40:15,540
hoping that everything comes back up in no time, right?

542
00:40:15,540 --> 00:40:21,660
Because it's not only like moving into another region or like putting that, that backup plan

543
00:40:21,660 --> 00:40:25,860
or that as a lay plan and the interaction, but it's also going back to the normal situation

544
00:40:25,860 --> 00:40:30,380
afterwards because that's also something that a lot of people don't realize that if something

545
00:40:30,380 --> 00:40:35,380
breaks and if something fails, at a certain point in time, you need to head back over to the

546
00:40:35,380 --> 00:40:40,740
normal situation because your backup situation can never be your normal situation.

547
00:40:40,740 --> 00:40:47,300
Unless you can really say it's 100% on par on each side, that's also something that you

548
00:40:47,300 --> 00:40:51,060
can do, but then you're going to start paying a lot of money, for instance, right?

549
00:40:51,060 --> 00:40:52,900
And a lot of service will help you with that.

550
00:40:52,900 --> 00:40:59,300
Think about things like service buzz that will now copy cues into one region into another

551
00:40:59,300 --> 00:41:02,780
or high resilient and are available.

552
00:41:02,780 --> 00:41:06,580
Think about new service like Azure resilience manager.

553
00:41:06,580 --> 00:41:10,700
Think about new services like not all new services.

554
00:41:10,700 --> 00:41:13,140
Think about Azure site recovery.

555
00:41:13,140 --> 00:41:16,780
Think about Cosmos DB.

556
00:41:16,780 --> 00:41:21,660
All these services have these capabilities to quickly fail over into another region or

557
00:41:21,660 --> 00:41:28,140
into another zone for that matter and making sure that your service stays up and running.

558
00:41:28,140 --> 00:41:30,900
But 100% forget about that.

559
00:41:30,900 --> 00:41:33,900
That will never fly that will never happen.

560
00:41:33,900 --> 00:41:45,460
Yeah, for me, it's a, I never have, one time, one time, but it was Microsoft Teams Failure,

561
00:41:45,460 --> 00:41:53,180
something deployed, Microsoft something and I cannot use Teams, and it's really my most

562
00:41:53,180 --> 00:41:54,340
use tool.

563
00:41:54,340 --> 00:42:05,580
But when I think on the Azure services, for me is, it's, it's near zero-dalt time.

564
00:42:05,580 --> 00:42:07,500
That's what Microsoft aims to do too.

565
00:42:07,500 --> 00:42:09,060
Yeah, exactly.

566
00:42:09,060 --> 00:42:17,220
I mean, and we, Microsoft is being very helpful with globally deployed services.

567
00:42:17,220 --> 00:42:25,460
Think about, Intra, think about helping you with traffic management, thinking about front

568
00:42:25,460 --> 00:42:30,420
or service that they also use because they came from their own stables and they productized

569
00:42:30,420 --> 00:42:31,740
it afterwards.

570
00:42:31,740 --> 00:42:35,260
Front or for instance, it's been existing for over so many years.

571
00:42:35,260 --> 00:42:40,780
It was first used in service like OneNote, Sky for Business, but also Halo for instance.

572
00:42:40,780 --> 00:42:43,340
As a front, as an entry point.

573
00:42:43,340 --> 00:42:47,020
So that way it worked at scale, but it also worked for resilient because you know, you

574
00:42:47,020 --> 00:42:54,460
don't want to like piss off 150,000 gamers when they're doing a dead match in Halo for instance,

575
00:42:54,460 --> 00:42:55,460
right?

576
00:42:55,460 --> 00:42:56,700
You don't want to, you don't want to see that happen.

577
00:42:56,700 --> 00:42:58,940
Do you have claims going on?

578
00:42:58,940 --> 00:43:03,180
So that's something you try to avoid at all costs.

579
00:43:03,180 --> 00:43:08,860
And with the knowledge that they have of managing their own services, they now deliver services

580
00:43:08,860 --> 00:43:18,220
that can be up to 100% or at least five, nine's availability.

581
00:43:18,220 --> 00:43:24,060
But again, as I stated, you can get as close as possible, but it's always a sum of all

582
00:43:24,060 --> 00:43:25,060
fears.

583
00:43:25,060 --> 00:43:35,060
If you have multiple services at 99.9, well, it will never reach even two nines behind the

584
00:43:35,060 --> 00:43:38,780
coma.

585
00:43:38,780 --> 00:43:48,380
I think over three years, we have downtime for over 30 seconds.

586
00:43:48,380 --> 00:43:58,020
I think all my clients come in the 30 seconds, but in our sound, no joke.

587
00:43:58,020 --> 00:44:06,300
And we have talk about the IPO and the IPO and the business.

588
00:44:06,300 --> 00:44:11,700
And I think a little bit about it.

589
00:44:11,700 --> 00:44:23,380
Sorry for those because I'm only one person company and from you, nearly, yeah, I don't

590
00:44:23,380 --> 00:44:26,620
care.

591
00:44:26,620 --> 00:44:35,500
But what point do is another nine of the labels to stop making financial sentence when you

592
00:44:35,500 --> 00:44:42,580
talk to your clients as architecture and how did you see it?

593
00:44:42,580 --> 00:44:47,700
Is it the nine, nine, nine, nine, nine, nine, nine, nine, nine, nine, nine, nine, nine, nine,

594
00:44:47,700 --> 00:44:52,540
or I think, I say, let's have a present.

595
00:44:52,540 --> 00:44:53,540
It's good.

596
00:44:53,540 --> 00:44:55,660
I don't know.

597
00:44:55,660 --> 00:44:56,580
How did you see it?

598
00:44:56,580 --> 00:45:01,180
It depends on the business use case and it also depends on the way that they're looking

599
00:45:01,180 --> 00:45:09,380
at a customer management internally or externally because it's not always of external facing

600
00:45:09,380 --> 00:45:10,380
applications.

601
00:45:10,380 --> 00:45:13,140
If your sauce company, it becomes more important.

602
00:45:13,140 --> 00:45:19,120
If it's internal tooling or internal sites or internal applications, you can cure less

603
00:45:19,120 --> 00:45:23,580
because there's always something that breaks internally.

604
00:45:23,580 --> 00:45:31,580
To give you an idea, if a holiday application form internally for your company doesn't

605
00:45:31,580 --> 00:45:34,700
work for a couple of hours, that's an issue.

606
00:45:34,700 --> 00:45:35,940
It can be bypassed.

607
00:45:35,940 --> 00:45:41,500
I mean, if you manage it, it says go by mail, you fill it out tomorrow and so good.

608
00:45:41,500 --> 00:45:45,780
But if you have an order fulfillment system, for instance, where you are the next, let's

609
00:45:45,780 --> 00:45:53,180
say, Amazon or you are the next AliExpress or sheen, that's a different story because every

610
00:45:53,180 --> 00:45:58,260
minute that you have downtime brings you, breaks your bank, right?

611
00:45:58,260 --> 00:45:59,260
The costs are high.

612
00:45:59,260 --> 00:46:06,940
So you need to see for yourself what a financial impact will be of a non-resilient or a non-high

613
00:46:06,940 --> 00:46:10,180
available application.

614
00:46:10,180 --> 00:46:17,340
The resilience in the available applications here then become really something important

615
00:46:17,340 --> 00:46:23,340
because high available is one thing.

616
00:46:23,340 --> 00:46:25,340
That's bad luck in the cost of your bank.

617
00:46:25,340 --> 00:46:31,340
But if your application is not resilient, you always have some issues like tropling or whatever.

618
00:46:31,340 --> 00:46:32,340
Customers will stay away.

619
00:46:32,340 --> 00:46:37,740
That's even worse because if you have, let's say, for instance, that you are building the

620
00:46:37,740 --> 00:46:43,980
next Amazon and every click that you do to buy a product.

621
00:46:43,980 --> 00:46:49,980
And every click that you do to buy a product takes about five minutes to store it and then

622
00:46:49,980 --> 00:46:53,460
you see that your basket is empty, your shopping basket is empty.

623
00:46:53,460 --> 00:46:58,180
You know, people will stay away from this site and they will no longer be shopping on

624
00:46:58,180 --> 00:46:59,180
your site.

625
00:46:59,180 --> 00:47:03,420
They will go look and look somewhere else, even if you are cheaper than that because they

626
00:47:03,420 --> 00:47:09,540
cannot be guaranteed that their shopping basket will not be emptied every single time.

627
00:47:09,540 --> 00:47:18,220
And they try to buy it is my money paid, yes or no, will I get my stuff that I ordered?

628
00:47:18,220 --> 00:47:20,100
It's all these things that come to play.

629
00:47:20,100 --> 00:47:26,460
And this also comes with credibility of a brand, a key credibility of a platform.

630
00:47:26,460 --> 00:47:27,860
And certainly a front end.

631
00:47:27,860 --> 00:47:31,340
It's also the API that you're using.

632
00:47:31,340 --> 00:47:37,500
It's also about a trockling there.

633
00:47:37,500 --> 00:47:43,700
On the other hand, you should also foresee the bypass on that because it doesn't always

634
00:47:43,700 --> 00:47:49,420
work in the direction of, ah, your customers need to be resilient for your customers.

635
00:47:49,420 --> 00:47:52,540
It also needs to be resilient for you.

636
00:47:52,540 --> 00:47:58,940
What happens if somebody tries to attack you and if you'll bring you down, right?

637
00:47:58,940 --> 00:48:02,580
That's also something that the denial of all that service is a reality, right?

638
00:48:02,580 --> 00:48:10,500
You see a lot of SMS or SMS attacks because SMS OTP is a very nice thing to do and if

639
00:48:10,500 --> 00:48:12,420
you'll break your bank, right?

640
00:48:12,420 --> 00:48:16,340
Because every single SMS that you send out will cost you money and it will not cost the

641
00:48:16,340 --> 00:48:19,660
attacker's money.

642
00:48:19,660 --> 00:48:20,660
Is that a good design pattern?

643
00:48:20,660 --> 00:48:22,780
No, you try to do it in a different way.

644
00:48:22,780 --> 00:48:25,100
So it's all these little things that come to play.

645
00:48:25,100 --> 00:48:31,180
If you have an API which is constantly being troubled and it will also trouble then your

646
00:48:31,180 --> 00:48:37,420
application that is actually like a, a details attack that you have on your, on your API,

647
00:48:37,420 --> 00:48:39,180
even it was not intended like that.

648
00:48:39,180 --> 00:48:41,420
So it's also a protection mechanism.

649
00:48:41,420 --> 00:48:43,060
So it works in both directions.

650
00:48:43,060 --> 00:48:46,180
You want to make sure that you're protected, but also your customers are protected.

651
00:48:46,180 --> 00:48:49,180
It's not only about customer UX and experience.

652
00:48:49,180 --> 00:48:52,620
It's also about protecting the platform and making sure that everything is there.

653
00:48:52,620 --> 00:48:58,420
So as you can see, our shit actually is more than just of the parts that are moving.

654
00:48:58,420 --> 00:49:04,460
It's also protecting the parts encapsulating them, making sure they're safe, warm, well-nurtured,

655
00:49:04,460 --> 00:49:05,460
well monitored.

656
00:49:05,460 --> 00:49:08,860
Actually, you can see in our shit texture as a baby, right?

657
00:49:08,860 --> 00:49:10,420
It needs to grow.

658
00:49:10,420 --> 00:49:14,540
It needs to make sure that it gets warm and godly.

659
00:49:14,540 --> 00:49:16,500
It needs to be fed right on time.

660
00:49:16,500 --> 00:49:21,500
It needs to be provided and all in all primary necessities.

661
00:49:21,500 --> 00:49:22,500
That's a thing.

662
00:49:22,500 --> 00:49:25,500
And application of solutions are just the same thing.

663
00:49:25,500 --> 00:49:26,500
You need to nurture them.

664
00:49:26,500 --> 00:49:27,500
You need to monitor them.

665
00:49:27,500 --> 00:49:28,500
You need to feed them.

666
00:49:28,500 --> 00:49:30,580
You need to make sure that they're safe.

667
00:49:30,580 --> 00:49:33,100
So that's basically what comes out.

668
00:49:33,100 --> 00:49:34,100
Yeah.

669
00:49:34,100 --> 00:49:46,340
Or this was really cool, but I think what's about we like to get nearly the 100% and there's,

670
00:49:46,340 --> 00:49:56,140
I think it's not really new, but Microsoft allowed the multi-region architecture.

671
00:49:56,140 --> 00:50:02,860
For me, movie region, Asica sounds like.

672
00:50:02,860 --> 00:50:05,020
Yeah.

673
00:50:05,020 --> 00:50:11,380
It's the best result I can get.

674
00:50:11,380 --> 00:50:21,260
When multi-asica actually got a little bit wrong is the cost, is the data consensus, the

675
00:50:21,260 --> 00:50:28,420
agency is, so I don't know, application design, I don't know, I ask you, Mike.

676
00:50:28,420 --> 00:50:30,580
So we were talking about multi-regency.

677
00:50:30,580 --> 00:50:32,340
There's a couple of approaches that you can have.

678
00:50:32,340 --> 00:50:37,700
I do you do a failover or you do a full active active, just like you would do in the boss

679
00:50:37,700 --> 00:50:41,220
in the data center where you have two servers and you have an active server.

680
00:50:41,220 --> 00:50:44,180
Or you switch on the other server only when you need it.

681
00:50:44,180 --> 00:50:46,140
It depends on the choices that you have.

682
00:50:46,140 --> 00:50:54,060
But again, comes down to the RTO, RPO, that discussion, SLA, resiliency, that part.

683
00:50:54,060 --> 00:50:59,580
Designing these are indeed more complex because you need to make sure what will I do with

684
00:50:59,580 --> 00:51:01,740
my data?

685
00:51:01,740 --> 00:51:03,620
How will I design that?

686
00:51:03,620 --> 00:51:08,940
Because, all right, I have now, for instance, I have an application which needs to stay

687
00:51:08,940 --> 00:51:10,420
on European soil.

688
00:51:10,420 --> 00:51:12,420
Okay, fine.

689
00:51:12,420 --> 00:51:16,580
I will take multiple European regions.

690
00:51:16,580 --> 00:51:19,300
But I also have customers in the United States.

691
00:51:19,300 --> 00:51:27,100
Let's say, for instance, for some reason, an entire region goes down on the internet because

692
00:51:27,100 --> 00:51:32,860
somebody broke a main DNS or somebody broke an on the C cable which brought down the entire

693
00:51:32,860 --> 00:51:33,860
region.

694
00:51:33,860 --> 00:51:38,180
So there's no more USAID or no more South America, no India or whatever.

695
00:51:38,180 --> 00:51:43,020
Yeah, now we need to start looking into another multi-region design.

696
00:51:43,020 --> 00:51:44,500
Can we go cross-boundary?

697
00:51:44,500 --> 00:51:45,980
Can we move data?

698
00:51:45,980 --> 00:51:48,020
Can we are we allowed to?

699
00:51:48,020 --> 00:51:49,020
Right?

700
00:51:49,020 --> 00:51:51,100
So these are questions that you should ask yourself.

701
00:51:51,100 --> 00:51:59,740
There is a good set of design patterns which Microsoft calls mission critical.

702
00:51:59,740 --> 00:52:01,940
And there's a full design set on that.

703
00:52:01,940 --> 00:52:09,820
There's actually, I think there's even an A-K-Aid on the miss on that, I don't know, by

704
00:52:09,820 --> 00:52:13,140
hard, I think it's mission critical.

705
00:52:13,140 --> 00:52:16,660
I can look it up.

706
00:52:16,660 --> 00:52:25,340
But there's a full design spec on that which actually allows you to add it through to

707
00:52:25,340 --> 00:52:26,860
allow should take it framework.

708
00:52:26,860 --> 00:52:33,740
And specifically for those scenarios you get additional things like multi-region but also

709
00:52:33,740 --> 00:52:37,900
multi-data, center multi-region, what with the data, how should I approach these, what are

710
00:52:37,900 --> 00:52:41,540
the patterns behind it, what are design areas that are after it.

711
00:52:41,540 --> 00:52:45,940
We have indeed the data platform, we have the networking connectivity because, in typically

712
00:52:45,940 --> 00:52:55,380
you also start speaking about multi-outbreaks with the A-Kamon express routes.

713
00:52:55,380 --> 00:53:04,260
You have multiple hub spots, hub spoke designs on that level but you need to design multiple

714
00:53:04,260 --> 00:53:06,700
regions that way.

715
00:53:06,700 --> 00:53:11,580
You need to have a circuit breaker on networking level, you need to have a circuit breaker

716
00:53:11,580 --> 00:53:16,620
on the in-s level, you need to have a circuit breaker on public endpoints.

717
00:53:16,620 --> 00:53:20,780
The security will be different because you need to monitor on different levels.

718
00:53:20,780 --> 00:53:25,500
You also need to make sure that your security is being able to move on to another side so

719
00:53:25,500 --> 00:53:30,020
that means that you need to have a multiple log analytics workspace available for that

720
00:53:30,020 --> 00:53:31,020
failovers on these.

721
00:53:31,020 --> 00:53:34,620
So it's all these things that come to play.

722
00:53:34,620 --> 00:53:37,260
Rishing critical gets you there.

723
00:53:37,260 --> 00:53:43,660
There's a full design methodology behind that and specifically on reliability that you're

724
00:53:43,660 --> 00:53:50,380
going to design it with resiliency into that.

725
00:53:50,380 --> 00:53:55,780
There's a design for zero downtime deployments making sure that you can just deploy in a snap

726
00:53:55,780 --> 00:53:58,820
without interfering with all the other stuff.

727
00:53:58,820 --> 00:54:05,380
Then we talk about infrastructure, code, CI/CD pipelines that can also be warmed up very

728
00:54:05,380 --> 00:54:11,140
quickly so that way you're there, and immediately, and that your design is actually basically stateless

729
00:54:11,140 --> 00:54:14,140
and that you can just push it from one side to another.

730
00:54:14,140 --> 00:54:18,340
It's all these things that come to play.

731
00:54:18,340 --> 00:54:26,180
It's actually pretty cool from an architecture point of view because there are so many gears

732
00:54:26,180 --> 00:54:32,020
involved that it's actually very nice to see but it also comes with a surgical complexity

733
00:54:32,020 --> 00:54:40,540
that people need to be able to grasp it in a really correct way.

734
00:54:40,540 --> 00:54:44,100
That's the biggest issue, typically with mission critical.

735
00:54:44,100 --> 00:54:47,340
But there's design patterns to that as such.

736
00:54:47,340 --> 00:54:59,340
It's a really interesting question, Microsoft and all these other vendors deliver design

737
00:54:59,340 --> 00:55:04,420
patterns for how secure stuff.

738
00:55:04,420 --> 00:55:15,020
When I'm a hacker, I can look at these guys or these patterns, not secure.

739
00:55:15,020 --> 00:55:18,260
So this is my tech part.

740
00:55:18,260 --> 00:55:25,980
And I think a little bit about ROM somewhere or what I see is ROM somewhere.

741
00:55:25,980 --> 00:55:27,460
It's cool.

742
00:55:27,460 --> 00:55:30,140
I can talk a long time about it.

743
00:55:30,140 --> 00:55:43,540
But what really, what Microsoft really have an issue is the identity, the complexity, I

744
00:55:43,540 --> 00:55:45,580
say, hey, I'm Mercopeaters.

745
00:55:45,580 --> 00:55:48,860
I started with this a while ago.

746
00:55:48,860 --> 00:55:50,900
Hey, I am from the University.

747
00:55:50,900 --> 00:55:58,900
Can you give me the name for the professor that I call him and I get online every?

748
00:55:58,900 --> 00:56:10,860
What was Microsoft or the Azure do in this compromised?

749
00:56:10,860 --> 00:56:12,860
You mean like social engineering?

750
00:56:12,860 --> 00:56:16,620
Why is there more like?

751
00:56:16,620 --> 00:56:21,260
It's actually a good one.

752
00:56:21,260 --> 00:56:29,540
Microsoft has the security response center and they're actively monitoring a lot of things.

753
00:56:29,540 --> 00:56:38,340
If identity gets compromised at scale, they will probably know it before you know it.

754
00:56:38,340 --> 00:56:43,820
And again, there's a lot of things that you need to put up play to the responsibility also

755
00:56:43,820 --> 00:56:45,940
comes in your hand, right?

756
00:56:45,940 --> 00:56:51,340
Microsoft does a lot of work when it comes down to threat management and they have a threat

757
00:56:51,340 --> 00:56:57,740
intelligence center, which is looking actively at hacker groups and in all these ransomware

758
00:56:57,740 --> 00:57:01,340
groups.

759
00:57:01,340 --> 00:57:04,820
But it also depends on the design that you have.

760
00:57:04,820 --> 00:57:09,580
I am, for instance, I hate VMs just because of that reason.

761
00:57:09,580 --> 00:57:14,780
A VM is a liability because you still have to patch it yourself.

762
00:57:14,780 --> 00:57:16,500
Somebody can log into it.

763
00:57:16,500 --> 00:57:21,140
You typically cannot see it that somebody logged into it unless you are monitoring in a

764
00:57:21,140 --> 00:57:27,100
correct way, but even that the noise is getting lost in the bigger noise of the entire solution.

765
00:57:27,100 --> 00:57:29,020
So you don't always see that.

766
00:57:29,020 --> 00:57:33,820
So your first defense becomes your identity, but a lot of people just come in with a local

767
00:57:33,820 --> 00:57:37,220
user on a server where there's still a portal.

768
00:57:37,220 --> 00:57:38,620
And so people can just enter it.

769
00:57:38,620 --> 00:57:45,420
And before you know, they just confirm the entire VM, the entire network and so on and so on.

770
00:57:45,420 --> 00:57:51,340
Microsoft will be able to help you or assist you in his kind of attacks, but a lot of responsibility

771
00:57:51,340 --> 00:57:52,940
also lies on your hands.

772
00:57:52,940 --> 00:57:54,820
Are you monitoring correctly?

773
00:57:54,820 --> 00:58:01,140
The things installed like conditional access, do you have multifactor authentication?

774
00:58:01,140 --> 00:58:09,140
It's not for a no reason that Microsoft enforces now MFA on intra identity that it forces you

775
00:58:09,140 --> 00:58:15,620
to use pass keys on all your identity systems nowadays, just like any other security vendor

776
00:58:15,620 --> 00:58:19,020
or any other identity vendor for that matter.

777
00:58:19,020 --> 00:58:22,220
So it's all these things that come to play.

778
00:58:22,220 --> 00:58:30,020
Microsoft has the security response center, also the threat intelligence center.

779
00:58:30,020 --> 00:58:36,620
And they try to work with large customers on that level, but a lot of responsibility still

780
00:58:36,620 --> 00:58:43,380
ends up on your lap because you still need to have a good insight on what has been deployed.

781
00:58:43,380 --> 00:58:47,380
Microsoft can only just do as much stay, stay at the late look.

782
00:58:47,380 --> 00:58:49,660
We have found the CV in this software.

783
00:58:49,660 --> 00:58:52,740
We have packages that are gone rogue.

784
00:58:52,740 --> 00:58:58,660
We have a we have fixed an issue on the Azure port over tokens can get stolen.

785
00:58:58,660 --> 00:59:03,220
For instance, not saying that it is a fact, but these things happen, right?

786
00:59:03,220 --> 00:59:07,020
Everybody's guilty about because you building software, there's always liabilities and there's

787
00:59:07,020 --> 00:59:08,220
always dependencies.

788
00:59:08,220 --> 00:59:14,340
If you're building something at scale, anybody is vulnerable to it, where it's a cloud vendor,

789
00:59:14,340 --> 00:59:21,580
whether it's an AI vendor, whether it's a gaming industry part, there's always something

790
00:59:21,580 --> 00:59:25,420
which is vulnerable that nobody discovered yet.

791
00:59:25,420 --> 00:59:27,940
And there's always an entry point.

792
00:59:27,940 --> 00:59:32,260
And these companies can just do as much as that.

793
00:59:32,260 --> 00:59:40,100
They have no 100% guarantee that they will give you 100% safety.

794
00:59:40,100 --> 00:59:45,060
It's just like within the delay, they give you as much safety as possible.

795
00:59:45,060 --> 00:59:47,980
But it comes with a grain of salt.

796
00:59:47,980 --> 00:59:52,820
Ask within anybody as within anything else.

797
00:59:52,820 --> 00:59:53,820
Awesome.

798
00:59:53,820 --> 01:00:01,180
I think a little about, yeah, normally I have a rapid fire round.

799
01:00:01,180 --> 01:00:10,260
I ask your short questions or ask a question.

800
01:00:10,260 --> 01:00:12,260
And you give a short answer.

801
01:00:12,260 --> 01:00:13,780
Okay, I try to.

802
01:00:13,780 --> 01:00:17,740
It's just okay when it's longer.

803
01:00:17,740 --> 01:00:18,740
We have.

804
01:00:18,740 --> 01:00:21,540
I have a lot of time.

805
01:00:21,540 --> 01:00:22,540
Perfect.

806
01:00:22,540 --> 01:00:31,300
So, human needs is for application with three microservices, which one?

807
01:00:31,300 --> 01:00:33,860
So come again, sorry?

808
01:00:33,860 --> 01:00:40,580
My Kubernetes for applications with three microservices.

809
01:00:40,580 --> 01:00:42,580
With three microservices.

810
01:00:42,580 --> 01:00:48,860
A user, no, or should you or should you not?

811
01:00:48,860 --> 01:00:54,900
You get through everything you are on single island and you have only.

812
01:00:54,900 --> 01:00:57,180
It's just free microservices.

813
01:00:57,180 --> 01:01:00,940
Then Kubernetes is a little bit too bloated for that.

814
01:01:00,940 --> 01:01:07,220
I would say no, don't go for that and just run it as a function or run it as an app service.

815
01:01:07,220 --> 01:01:10,100
It's going to be cheaper too.

816
01:01:10,100 --> 01:01:13,860
Yeah, you still a little bit out of the transmission.

817
01:01:13,860 --> 01:01:15,980
Okay, I understand.

818
01:01:15,980 --> 01:01:19,620
I come from you, Cryatman, 99.99.

819
01:01:19,620 --> 01:01:20,580
99.99.

820
01:01:20,580 --> 01:01:25,620
Is there ever a way of it?

821
01:01:25,620 --> 01:01:28,620
But we're choosing what the region cost.

822
01:01:28,620 --> 01:01:33,620
What's your answer?

823
01:01:33,620 --> 01:01:37,740
Good question.

824
01:01:37,740 --> 01:01:38,940
Repeat it?

825
01:01:38,940 --> 01:01:41,260
So a company has 99.

826
01:01:41,260 --> 01:01:47,980
99.99, you can make it a long time.

827
01:01:47,980 --> 01:01:53,660
Available, but they refuse to pay for the multi-region.

828
01:01:53,660 --> 01:01:55,540
What's your answer?

829
01:01:55,540 --> 01:01:58,260
Yeah, that part doesn't get.

830
01:01:58,260 --> 01:02:01,420
I would say good luck.

831
01:02:01,420 --> 01:02:05,020
I love it.

832
01:02:05,020 --> 01:02:10,660
At that base, back up has never been restored.

833
01:02:10,660 --> 01:02:12,020
MNCS previously.

834
01:02:12,020 --> 01:02:13,020
Good luck.

835
01:02:13,020 --> 01:02:20,900
You should always test mistakes.

836
01:02:20,900 --> 01:02:27,140
If you called, no, no, no, no, no, my favorite question.

837
01:02:27,140 --> 01:02:32,060
Safial Dalia called you and said, Mike, A, all cool.

838
01:02:32,060 --> 01:02:38,500
I give you all the responsibilities and the money to develop something you like, which

839
01:02:38,500 --> 01:02:40,860
feature will it be?

840
01:02:40,860 --> 01:02:44,260
A feature or a service?

841
01:02:44,260 --> 01:02:46,060
You are free.

842
01:02:46,060 --> 01:02:49,740
I love it.

843
01:02:49,740 --> 01:02:57,980
Oh, I would say one of my web dreams is still to have a full graphical designer to design

844
01:02:57,980 --> 01:03:05,300
architecture, like a real designing in a graphical way your Azure architecture and that you

845
01:03:05,300 --> 01:03:09,980
also can immediately deploy the applications that you need for that.

846
01:03:09,980 --> 01:03:11,300
That would be cool.

847
01:03:11,300 --> 01:03:19,580
So think of a visual like, but then specifically for Azure, plugging in things and making it simple

848
01:03:19,580 --> 01:03:23,340
and stupid, but also based upon well architecture.

849
01:03:23,340 --> 01:03:30,580
And actually come to think of it, shouldn't be that hard to build nowadays with AI, right?

850
01:03:30,580 --> 01:03:32,580
Yeah.

851
01:03:32,580 --> 01:03:37,740
Yeah, yeah, this AI stuff.

852
01:03:37,740 --> 01:03:47,900
What did you think about Kubernetes for applications with our or should we use Microsoft, Microsoft?

853
01:03:47,900 --> 01:03:50,460
What's your own stuff, Mike?

854
01:03:50,460 --> 01:03:57,380
It really depends on the scale that you need, but also the complexity because Kubernetes is

855
01:03:57,380 --> 01:03:59,820
like steep learning curve.

856
01:03:59,820 --> 01:04:02,180
And if you have never done it, stay away from it.

857
01:04:02,180 --> 01:04:06,460
If you just want to focus on the applications and please go with container services or just

858
01:04:06,460 --> 01:04:11,980
app services and run it as a container there, I mean, depending on the use case, I would

859
01:04:11,980 --> 01:04:12,980
say.

860
01:04:12,980 --> 01:04:20,900
Mike, this was also awesome experience.

861
01:04:20,900 --> 01:04:31,020
I think, yeah, we have learned so much about Azure and applications and we are not believe

862
01:04:31,020 --> 01:04:32,020
it.

863
01:04:32,020 --> 01:04:39,340
I think I believe, but we believe of the percent of available time.

864
01:04:39,340 --> 01:04:41,020
Thank you so much.

865
01:04:41,020 --> 01:04:44,180
It's more a great interview with you.

866
01:04:44,180 --> 01:04:46,020
Yeah, it was great to make her.

867
01:04:46,020 --> 01:04:47,020
Thank you for having me.

868
01:04:47,020 --> 01:04:48,700
It was, it was my pleasure.

869
01:04:48,700 --> 01:04:53,420
It was also funny that we talked about resilience and that my internet went down during the

870
01:04:53,420 --> 01:04:54,420
discussions.

871
01:04:54,420 --> 01:04:58,020
That made it even made it more more powerful.

872
01:04:58,020 --> 01:04:59,980
Yeah, that's a thing.

873
01:04:59,980 --> 01:05:03,380
So that made it even more sarcastic and ironic, right?

874
01:05:03,380 --> 01:05:05,300
So that's a lot of fun.

875
01:05:05,300 --> 01:05:10,740
No, thank you for having me and yeah, I'm looking forward when this thing comes online.

876
01:05:10,740 --> 01:05:10,940
Thank you.

877
01:05:10,940 --> 01:05:29,780
[ Music ]

Mirko Peters Profile Photo

Founder of m365.fm, m365.show and m365con.net

Mirko Peters is a Microsoft 365 expert, content creator, and founder of m365.fm, a platform dedicated to sharing practical insights on modern workplace technologies. His work focuses on Microsoft 365 governance, security, collaboration, and real-world implementation strategies.

Through his podcast and written content, Mirko provides hands-on guidance for IT professionals, architects, and business leaders navigating the complexities of Microsoft 365. He is known for translating complex topics into clear, actionable advice, often highlighting common mistakes and overlooked risks in real-world environments.

With a strong emphasis on community contribution and knowledge sharing, Mirko is actively building a platform that connects experts, shares experiences, and helps organizations get the most out of their Microsoft 365 investments.

Mike Martin Profile Photo

Technical Evangelist - architect - strategist

As a Technical Evangelist, Mike is an Azure goto for anyone in the community. He’s been active in the IT industry for almost 30 years and has performed almost all types of job profiles, going from coaching and leading a team to architecting and systems design and training. Today he’s primarily into the Microsoft Cloud Platform and Application Lifecycle Management. He’s not a stranger to both dev and IT Pro topics, they even call him the perfect hybrid solution.

In January 2012 he became a crew member of AZUG, the Belgian Microsoft Azure User Group. As an active member he’s both involved in giving presentations and organizing events (like Techorama and Cloudbrew). Mike was also a Microsoft Azure MVP befire blue badging (awarded 5 times since 2013) and received his renewal in December 2025!.
Helping out in the community and introducing new & young people into the world of Microsoft and technology is also one of his passions.

Related to this Episode

Why Kubernetes Is Overkill for Most Azure Apps (And What to Use Instead)

When designing cloud architecture, many teams immediately default to Kubernetes and complex microservices, assuming modern cloud-native systems require container orchestration. However, forcing every application into Kubernetes introduces unnecessar…