Phi-4 is the Runtime. MAI-1 is the Reason.
Microsoft’s AI strategy becomes much easier to understand once we stop asking which model is better. The important question is no longer whether MAI-1 can beat Phi-4, whether Phi-4 is more efficient, or which model should become the enterprise standard. Those questions assume that both models are competing for the same job. They are not. Microsoft is building toward an architecture in which different forms of intelligence perform different roles. Phi-4 represents the fast, efficient Runtime layer. MAI-1 represents the deeper Reasoning layer. One executes close to the workload; the other handles problems that justify substantially more reasoning capability. This distinction matters because enterprise AI is moving beyond the era of connecting every application to one enormous general-purpose model. Organizations increasingly need to think about AI as an architecture consisting of models, routing, orchestration, governance, infrastructure, and specialized workloads. The competitive advantage may therefore come less from having access to the most powerful model and more from knowing when that power is actually necessary.
THE WRONG QUESTION: WHICH MODEL IS BETTER?
AI model launches are usually treated like sporting events. Benchmark scores are compared, parameter counts are examined, and eventually somebody declares a winner. That approach makes sense for consumers choosing between individual AI assistants, but enterprise architecture has never worked according to that principle. Companies don't use the same compute configuration for every application, the same storage tier for every file, or the same database architecture for every workload. There is little reason to assume intelligence should be different. The source describes this assumption as “The Model” thinking: the belief that one model ultimately needs to become the standard intelligence layer. Microsoft's emerging architecture points in another direction. Instead of selecting a winner, organizations need to understand the division of labor between models. Reasoning systems determine what should happen. Runtime systems execute work efficiently. Once intelligence is viewed this way, directly comparing Phi-4 and MAI-1 becomes much less useful. The meaningful comparison is between the requirements of a workload and the characteristics of the model handling it.
MICROSOFT IS BUILDING A DIVISION OF LABOR
The broader signal from Microsoft's model strategy is specialization. Instead of concentrating exclusively on one universal flagship model, Microsoft is developing multiple model families and capabilities spanning reasoning, coding, voice, images, transcription, multimodal processing, and efficient local execution. That suggests an architecture in which intelligence becomes distributed. Some models can live close to users and devices. Others can remain centralized because their workloads require substantially greater compute and context. Specialized models can handle specific modalities or business processes while deeper reasoning models become escalation points for problems requiring judgment and planning. The result begins to resemble a modern computing architecture more than a traditional chatbot. Different layers perform different jobs, and an orchestration mechanism connects those layers into what appears to the user to be one intelligent system.
DENSE AND SPARSE REPRESENT DIFFERENT DESIGN PHILOSOPHIES
Phi-4 and MAI-1 also demonstrate two different approaches to building intelligence. Phi-4 emphasizes density and efficiency. The objective is to produce substantial capability from a comparatively compact architecture. MAI-1 represents the opposite side of the equation, where substantially greater total capacity can be combined with selective activation through a Mixture-of-Experts architecture. A useful analogy is organizational structure. A small company may employ fewer specialists but expect almost everyone to participate whenever work arrives. A much larger organization can maintain hundreds of specialists while involving only the employees relevant to a particular problem. Both structures can work extremely well, but they optimize for different environments. That difference becomes crucial when AI reaches enterprise scale. Sending a simple classification request to a massive reasoning system can be unnecessary. Sending an extremely complicated planning problem to a small model can leave the system without enough reasoning capacity. Neither outcome means that the underlying model is bad. It means the workload was assigned to the wrong layer.
AI ARCHITECTURE IS ALSO AI ECONOMICS
Model architecture quickly becomes a financial issue once organizations move from experiments into production. During a proof of concept, the difference between a cheap inference request and an expensive one may appear insignificant. Multiply that difference across millions of requests and the architecture begins determining whether the use case is financially sustainable. The important metric therefore isn't simply cost per token. Organizations need to understand cost relative to workload complexity. A request that requires deep reasoning may justify substantially greater inference cost because the business problem itself is valuable. A routine classification request performed millions of times should be optimized very differently. This is why the Runtime-and-Reason distinction matters financially. The organization gains the ability to reserve expensive intelligence for workloads that actually benefit from it while moving repetitive execution into a significantly more efficient layer.
PHI-4 AS THE RUNTIME LAYER
A Runtime is responsible for execution. It sits close to the point where work happens and responds quickly enough that intelligence becomes part of the application experience rather than a remote service users are constantly waiting for. The source positions Phi-4's compact models naturally within this role. Phi-4-mini and Phi-4-multimodal are designed around relatively small footprints, while capabilities such as function calling allow the model to participate in agentic workflows rather than merely generate text. MIT licensing also creates considerably more flexibility for developers considering embedded and specialized deployments. Consider a local coding assistant examining a file, an endpoint agent classifying a request, an application deciding which internal function should execute, or an assistant performing routine processing against local information. These tasks require intelligence, but they do not necessarily require a frontier model with enormous context and deep planning capabilities. They need low latency, predictable execution, and an economic model that works at high volume. That is the Runtime role.
LOCAL AI CHANGES THE SOVEREIGNTY DISCUSSION
Local inference also changes the relationship between AI and data sovereignty. Traditionally, organizations have concentrated on securing the journey between corporate information and a remote AI service. If a workload can instead be processed directly on an endpoint or within a controlled local environment, some information may never need to make that journey. That doesn't eliminate governance. It changes where governance can be enforced. Data residency and sovereignty can potentially become characteristics of the architecture itself rather than controls applied after information has already been transmitted somewhere else. For highly regulated organizations, this distinction can be significant. Some workloads may be perfectly suitable for cloud reasoning while others should remain local by design. The question becomes workload-specific rather than forcing an organization into a binary choice between cloud AI and local AI.
MAI-1 AS THE REASONING LAYER
Reasoning serves a fundamentally different purpose. It is needed when a system must evaluate alternatives, maintain relationships across a large amount of information, understand dependencies, plan multiple steps ahead, or make sense of a problem where the correct next action is not immediately obvious. The source positions MAI-Thinking-1 within this deeper reasoning category and discusses its larger active reasoning capacity, substantial context capabilities, multi-step problem solving, training lineage, and software-engineering performance. Think about an architecture review involving multiple systems, a complicated root-cause investigation, a large codebase where changes have cascading consequences, or a strategic planning problem involving dozens of constraints. Those aren't primarily execution problems. They require the model to maintain the structure of the problem while evaluating what should happen next. That is where the Reason layer belongs.
PHI-4 EXECUTES WHILE MAI-1 DECIDES
The central idea of the architecture can therefore be reduced to a very simple distinction: Phi-4 executes. MAI-1 decides. The important part, however, is what connects those two layers. A Runtime without a reasoning layer eventually encounters problems beyond its capabilities. A reasoning layer without an efficient Runtime wastes expensive intelligence on routine execution. The architecture only becomes powerful when requests can move intelligently between them. That handoff may become considerably more important than the individual models themselves. If organizations can reliably determine when a problem requires escalation, they can create systems that feel fast for routine work while still providing sophisticated intelligence when complexity demands it.
Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-modern-work-security-and-productivity-with-microsoft-365--6704921/support.
🚀 Want to be part of m365.fm?
Then stop just listening… and start showing up.
👉 Connect with me on LinkedIn and let’s make something happen:
- 🎙️ Be a podcast guest and share your story
- 🎧 Host your own episode (yes, seriously)
- 💡 Pitch topics the community actually wants to hear
- 🌍 Build your personal brand in the Microsoft 365 space
This isn’t just a podcast — it’s a platform for people who take action.
🔥 Most people wait. The best ones don’t.
👉 Connect with me on LinkedIn and send me a message:
"I want in"
Let’s build something awesome 👊
00:00:00,000 --> 00:00:01,720
Everyone keeps asking the wrong question,
2
00:00:01,720 --> 00:00:03,920
is MAI one better than 5/4?
3
00:00:03,920 --> 00:00:05,640
Which one should we standardize on?
4
00:00:05,640 --> 00:00:07,140
That question was already outdated,
5
00:00:07,140 --> 00:00:09,360
the moment Microsoft shipped both of them on purpose
6
00:00:09,360 --> 00:00:10,160
in the same month.
7
00:00:10,160 --> 00:00:11,800
Here is the tension nobody is naming.
8
00:00:11,800 --> 00:00:13,960
Microsoft just released two completely different types
9
00:00:13,960 --> 00:00:15,920
of intelligence, not a bigger version
10
00:00:15,920 --> 00:00:17,640
and a smaller version of the same thing,
11
00:00:17,640 --> 00:00:19,720
two different organs built for two different jobs.
12
00:00:19,720 --> 00:00:21,120
Most of the coverage treats them
13
00:00:21,120 --> 00:00:23,680
like they are fighting for the same spot on the podium.
14
00:00:23,680 --> 00:00:26,240
But they are not MI1 and 5/4 are not rivals,
15
00:00:26,240 --> 00:00:29,480
they are not competing for the same job inside your stack.
16
00:00:29,480 --> 00:00:31,560
By the end of this, you are going to understand
17
00:00:31,560 --> 00:00:34,040
the actual architecture Microsoft is building toward,
18
00:00:34,040 --> 00:00:37,280
not a feature list, not a benchmark chart, the real system.
19
00:00:37,280 --> 00:00:39,400
If you are the kind of person who has to design this stuff
20
00:00:39,400 --> 00:00:41,160
for a living, subscribe now.
21
00:00:41,160 --> 00:00:42,600
Because this split is going to define
22
00:00:42,600 --> 00:00:45,360
how M365 admins and architects build systems
23
00:00:45,360 --> 00:00:46,960
for the next five years.
24
00:00:46,960 --> 00:00:48,800
The question everyone is getting wrong.
25
00:00:48,800 --> 00:00:51,120
So let's start with how most people are covering this.
26
00:00:51,120 --> 00:00:52,840
It explains why so many teams are about
27
00:00:52,840 --> 00:00:54,440
to make expensive mistakes.
28
00:00:54,440 --> 00:00:56,320
Watch any breakdown of MI1 and 5/4
29
00:00:56,320 --> 00:00:58,160
and it turns into a leaderboard battle,
30
00:00:58,160 --> 00:01:00,040
which one scores higher on this benchmark,
31
00:01:00,040 --> 00:01:02,920
which one is faster, which one is smarter per parameter,
32
00:01:02,920 --> 00:01:03,920
bigger versus smaller.
33
00:01:03,920 --> 00:01:06,360
Like it is a heavyweight fight and somebody has to lose.
34
00:01:06,360 --> 00:01:08,480
That framing comes from consumer AI thinking.
35
00:01:08,480 --> 00:01:10,560
It is the same instinct that made everyone argue
36
00:01:10,560 --> 00:01:12,800
about which chatbot writes better emails.
37
00:01:12,800 --> 00:01:14,520
And look, that instinct isn't stupid,
38
00:01:14,520 --> 00:01:16,520
it is just built for the wrong problem.
39
00:01:16,520 --> 00:01:19,360
When you are a consumer picking one assistant for one phone,
40
00:01:19,360 --> 00:01:21,960
which model is better is actually the right question.
41
00:01:21,960 --> 00:01:23,160
You only get to pick one,
42
00:01:23,160 --> 00:01:25,280
but enterprise system design does not work that way.
43
00:01:25,280 --> 00:01:26,120
It never did.
44
00:01:26,120 --> 00:01:27,840
Nobody runs one server for every workload.
45
00:01:27,840 --> 00:01:30,160
Nobody uses one storage tier for every file.
46
00:01:30,160 --> 00:01:32,080
So why would anybody assume there is supposed
47
00:01:32,080 --> 00:01:33,880
to be one model for every request?
48
00:01:33,880 --> 00:01:35,120
This is worth naming directly.
49
00:01:35,120 --> 00:01:37,320
It is the assumption sitting underneath almost everything
50
00:01:37,320 --> 00:01:38,600
written about this launch.
51
00:01:38,600 --> 00:01:40,120
Call it the model thinking.
52
00:01:40,120 --> 00:01:42,560
The old belief that one system has to win.
53
00:01:42,560 --> 00:01:44,520
The idea that intelligence is a single throne
54
00:01:44,520 --> 00:01:45,920
and somebody is sitting on it.
55
00:01:45,920 --> 00:01:48,840
That belief made sense when there was basically one option,
56
00:01:48,840 --> 00:01:50,440
rent GPT or rent something else,
57
00:01:50,440 --> 00:01:52,280
but that is not the world Microsoft just built.
58
00:01:52,280 --> 00:01:53,280
Here is the reframe.
59
00:01:53,280 --> 00:01:56,560
Microsoft is not choosing a winner between MI1 and 5/4.
60
00:01:56,560 --> 00:01:58,040
It is building a division of labor.
61
00:01:58,040 --> 00:01:59,080
Once you see it that way,
62
00:01:59,080 --> 00:02:00,880
the leaderboard question falls apart
63
00:02:00,880 --> 00:02:02,440
because you are not supposed to pick one.
64
00:02:02,440 --> 00:02:04,240
You are supposed to know when to use which.
65
00:02:04,240 --> 00:02:05,480
So let's set up the two words
66
00:02:05,480 --> 00:02:07,840
that are going to anchor everything from here forward.
67
00:02:07,840 --> 00:02:09,360
Reason and runtime.
68
00:02:09,360 --> 00:02:11,640
Reason is the deep layer, the part that plans,
69
00:02:11,640 --> 00:02:12,840
the part that weighs options,
70
00:02:12,840 --> 00:02:14,600
it figures out what should actually happen next
71
00:02:14,600 --> 00:02:16,640
in a complicated multi-step problem.
72
00:02:16,640 --> 00:02:18,880
Runtime is the fast layer, the part that executes,
73
00:02:18,880 --> 00:02:20,640
the part that responds instantly.
74
00:02:20,640 --> 00:02:23,000
It lives close to wherever the work is actually happening.
75
00:02:23,000 --> 00:02:24,520
MI1 sits in the reason slot,
76
00:02:24,520 --> 00:02:26,200
5/4 sits in the runtime slot.
77
00:02:26,200 --> 00:02:27,480
Once you frame it that way,
78
00:02:27,480 --> 00:02:30,080
comparing their benchmark scores directly is a mistake.
79
00:02:30,080 --> 00:02:32,440
It is like comparing a company's strategy team
80
00:02:32,440 --> 00:02:33,840
to its shipping department
81
00:02:33,840 --> 00:02:35,560
and asking which one is better.
82
00:02:35,560 --> 00:02:36,600
Wrong question.
83
00:02:36,600 --> 00:02:38,080
They are not measured against each other.
84
00:02:38,080 --> 00:02:40,320
They are measured against the job they were built for.
85
00:02:40,320 --> 00:02:42,640
That is why benchmarks alone cannot tell you what to do
86
00:02:42,640 --> 00:02:44,280
with either of these models.
87
00:02:44,280 --> 00:02:45,440
A number on a leaderboard
88
00:02:45,440 --> 00:02:48,320
does not tell you where a model belongs in your architecture.
89
00:02:48,320 --> 00:02:49,640
What tells you that is understanding
90
00:02:49,640 --> 00:02:52,080
what actually happened at Bill 2026?
91
00:02:52,080 --> 00:02:54,600
That is when Microsoft laid this whole strategy out in public
92
00:02:54,600 --> 00:02:57,440
and what actually happened at Bill 2026.
93
00:02:57,440 --> 00:02:59,200
So here is what actually got announced.
94
00:02:59,200 --> 00:03:01,120
The framing matters more than the coverage.
95
00:03:01,120 --> 00:03:02,960
Mustafa Suleiman walked on stage
96
00:03:02,960 --> 00:03:05,840
and unveiled seven new MI models in one sitting,
97
00:03:05,840 --> 00:03:07,960
not one flagship, not a single headline model.
98
00:03:07,960 --> 00:03:09,400
Seven spanning reasoning, coding,
99
00:03:09,400 --> 00:03:11,240
image generation, voice and transcription
100
00:03:11,240 --> 00:03:13,800
all released together, all trained by Microsoft's own team
101
00:03:13,800 --> 00:03:14,760
from scratch.
102
00:03:14,760 --> 00:03:16,960
And alongside that, 5/4 kept expanding too.
103
00:03:16,960 --> 00:03:18,920
More variance, more sizes,
104
00:03:18,920 --> 00:03:20,520
still shipping on its own track,
105
00:03:20,520 --> 00:03:23,160
two families growing at the same time in the same event.
106
00:03:23,160 --> 00:03:25,280
Microsoft didn't call this a product launch.
107
00:03:25,280 --> 00:03:27,800
They called it humanist superintelligence.
108
00:03:27,800 --> 00:03:30,440
State of the art capability across modalities,
109
00:03:30,440 --> 00:03:33,280
but explicitly designed to serve people and organizations
110
00:03:33,280 --> 00:03:34,120
not replace them.
111
00:03:34,120 --> 00:03:36,600
You can take that branding with a grain of salt if you want.
112
00:03:36,600 --> 00:03:37,560
That's fair.
113
00:03:37,560 --> 00:03:39,840
But underneath the phrase is a real claim.
114
00:03:39,840 --> 00:03:43,120
These models are meant to work across text, code, image, voice,
115
00:03:43,120 --> 00:03:45,320
and reasoning as one coordinated system,
116
00:03:45,320 --> 00:03:46,960
not a scattered side projects.
117
00:03:46,960 --> 00:03:48,920
Now here is the part that actually explains why
118
00:03:48,920 --> 00:03:51,240
this happened right now in this exact window
119
00:03:51,240 --> 00:03:52,320
and not two years ago.
120
00:03:52,320 --> 00:03:54,840
Until late 2025, Microsoft literally
121
00:03:54,840 --> 00:03:56,320
could not build frontier models.
122
00:03:56,320 --> 00:03:57,720
That wasn't a strategy choice.
123
00:03:57,720 --> 00:03:59,320
It was contractual.
124
00:03:59,320 --> 00:04:02,520
The original OpenAI partnership signed back in 2019,
125
00:04:02,520 --> 00:04:05,200
restricted Microsoft from independently pursuing
126
00:04:05,200 --> 00:04:08,320
frontier level AI or superintelligence on its own.
127
00:04:08,320 --> 00:04:10,880
So for years, Microsoft's entire AI story
128
00:04:10,880 --> 00:04:14,840
was really OpenAI's AI story, wearing a Microsoft badge.
129
00:04:14,840 --> 00:04:17,360
Every leap in capability ran through somebody else's lab
130
00:04:17,360 --> 00:04:17,760
first.
131
00:04:17,760 --> 00:04:20,520
That changed when the contract got renegotiated.
132
00:04:20,520 --> 00:04:22,040
And once that restriction lifted,
133
00:04:22,040 --> 00:04:24,360
Microsoft didn't ease into building its own models.
134
00:04:24,360 --> 00:04:25,200
It sprinted.
135
00:04:25,200 --> 00:04:27,280
Think about what that actually means strategically.
136
00:04:27,280 --> 00:04:29,440
In under a year, Microsoft went from being
137
00:04:29,440 --> 00:04:31,760
a distributor of intelligence to being a builder of it.
138
00:04:31,760 --> 00:04:33,560
Those are completely different businesses.
139
00:04:33,560 --> 00:04:36,120
A distributor's job is picking the best thing
140
00:04:36,120 --> 00:04:38,480
off somebody else's shelf and packaging it well.
141
00:04:38,480 --> 00:04:41,120
A builder's job is deciding what gets made in the first place,
142
00:04:41,120 --> 00:04:43,720
what trade-offs get baked in, what the model is actually for.
143
00:04:43,720 --> 00:04:46,640
Microsoft crossed that line in months, not years.
144
00:04:46,640 --> 00:04:48,760
And normally, when you hear seven new models
145
00:04:48,760 --> 00:04:50,240
shipped in one announcement, you
146
00:04:50,240 --> 00:04:53,200
assume some giant sprawling org chart behind it.
147
00:04:53,200 --> 00:04:55,720
Thousands of researchers, endless review cycles.
148
00:04:55,720 --> 00:04:57,240
The usual big company machinery.
149
00:04:57,240 --> 00:04:58,360
That's not what happened here.
150
00:04:58,360 --> 00:05:00,440
Suleiman described teams of around 10 people
151
00:05:00,440 --> 00:05:01,760
building these models.
152
00:05:01,760 --> 00:05:02,680
Flat structure.
153
00:05:02,680 --> 00:05:04,600
No bureaucracy stacked on top of them.
154
00:05:04,600 --> 00:05:06,440
More like a trading floor than a research lab.
155
00:05:06,440 --> 00:05:08,240
People working shoulder to shoulder.
156
00:05:08,240 --> 00:05:10,080
Moving fast because there was nobody above them
157
00:05:10,080 --> 00:05:11,240
slowing the decision down.
158
00:05:11,240 --> 00:05:13,720
That detail matters more than it sounds like it should.
159
00:05:13,720 --> 00:05:16,240
Small teams building frontier class models in parallel
160
00:05:16,240 --> 00:05:18,960
only works if each team has a narrow, clear job.
161
00:05:18,960 --> 00:05:21,040
You can't have 10 people own AI broadly.
162
00:05:21,040 --> 00:05:22,800
You can have 10 people own reasoning.
163
00:05:22,800 --> 00:05:24,160
10 people own transcription.
164
00:05:24,160 --> 00:05:25,640
10 people own image generation.
165
00:05:25,640 --> 00:05:28,280
The org structure itself is already a division of labor
166
00:05:28,280 --> 00:05:30,200
before a single line of code gets written.
167
00:05:30,200 --> 00:05:32,560
And that's the real signal buried in this announcement.
168
00:05:32,560 --> 00:05:33,640
Speed like this.
169
00:05:33,640 --> 00:05:36,480
Across seven models plus a whole separate 5/4 track
170
00:05:36,480 --> 00:05:37,920
doesn't happen when everyone's building
171
00:05:37,920 --> 00:05:40,800
slightly different versions of the same general purpose thing.
172
00:05:40,800 --> 00:05:42,600
It happens when the jobs are already split.
173
00:05:42,600 --> 00:05:44,880
Small teams move fast because they're not fighting
174
00:05:44,880 --> 00:05:45,920
over what the model should be.
175
00:05:45,920 --> 00:05:46,880
They already know.
176
00:05:46,880 --> 00:05:49,040
Which means the reason and runtime split we just laid out
177
00:05:49,040 --> 00:05:50,960
wasn't an accident that emerged later.
178
00:05:50,960 --> 00:05:53,080
It was baked into how Microsoft built these models
179
00:05:53,080 --> 00:05:55,920
from day one to different design philosophies.
180
00:05:55,920 --> 00:05:58,400
So now that you know why these teams moved fast,
181
00:05:58,400 --> 00:06:00,600
let's get into why they built two different things
182
00:06:00,600 --> 00:06:02,040
instead of one better thing.
183
00:06:02,040 --> 00:06:04,520
MAI-1 is built for scale and reasoning depth.
184
00:06:04,520 --> 00:06:07,120
5/4 is built for density and deployment speed.
185
00:06:07,120 --> 00:06:08,800
That's not a difference in quality.
186
00:06:08,800 --> 00:06:10,880
It's a difference in what problem each model is solving
187
00:06:10,880 --> 00:06:12,800
before it even gets asked a question.
188
00:06:12,800 --> 00:06:14,760
Here's where you need two words that sound technical
189
00:06:14,760 --> 00:06:15,840
but really aren't.
190
00:06:15,840 --> 00:06:17,160
Dense and Sparse.
191
00:06:17,160 --> 00:06:19,240
Think about a company where every single employee
192
00:06:19,240 --> 00:06:22,040
has to sit in on every single meeting no matter what it's about.
193
00:06:22,040 --> 00:06:22,920
That's a dense model.
194
00:06:22,920 --> 00:06:25,280
Every parameter fires for every request.
195
00:06:25,280 --> 00:06:27,360
Whether the question is what's two plus two
196
00:06:27,360 --> 00:06:29,880
or restructure our entire supply chain.
197
00:06:29,880 --> 00:06:30,880
Nothing gets to sit out.
198
00:06:30,880 --> 00:06:32,680
That's 5/4 approach and on purpose.
199
00:06:32,680 --> 00:06:35,640
It's a dense transformer with a small parameter count.
200
00:06:35,640 --> 00:06:37,880
Meaning there's less total headcount in the building
201
00:06:37,880 --> 00:06:39,600
but everyone who's there is working every time.
202
00:06:39,600 --> 00:06:41,160
Now picture a different company.
203
00:06:41,160 --> 00:06:42,800
Hundreds of specialists on staff.
204
00:06:42,800 --> 00:06:45,360
But for any given meeting only the three or four people
205
00:06:45,360 --> 00:06:46,960
who actually know that topic show up.
206
00:06:46,960 --> 00:06:48,360
Everyone else stays at their desk.
207
00:06:48,360 --> 00:06:50,560
That's Sparse. That's MI1.
208
00:06:50,560 --> 00:06:52,320
It's a mixture of experts model.
209
00:06:52,320 --> 00:06:55,480
Meaning there's a massive total roster of expertise sitting inside it
210
00:06:55,480 --> 00:06:58,080
but only a slice of that roster activates per token.
211
00:06:58,080 --> 00:07:01,800
Huge capacity on paper, partial activation in practice.
212
00:07:01,800 --> 00:07:04,200
That distinction is why you can't just say bigger is smarter
213
00:07:04,200 --> 00:07:05,200
and move on.
214
00:07:05,200 --> 00:07:07,400
MI1's total size is enormous
215
00:07:07,400 --> 00:07:09,320
but it's not paying the full cost of that size
216
00:07:09,320 --> 00:07:10,600
on every single request
217
00:07:10,600 --> 00:07:13,120
because most of the model is sitting quiet at any given moment.
218
00:07:13,120 --> 00:07:15,560
If 5/4 doesn't have that luxury and doesn't need it,
219
00:07:15,560 --> 00:07:18,800
it's small enough that using all of itself every time is still cheap.
220
00:07:18,800 --> 00:07:21,600
High quality per parameter is the whole design goal.
221
00:07:21,600 --> 00:07:23,280
Nothing wasted, nothing held in reserve
222
00:07:23,280 --> 00:07:25,400
because there's nothing to spare in the first place.
223
00:07:25,400 --> 00:07:28,720
Here is why this isn't just an engineering detail you can skip past.
224
00:07:28,720 --> 00:07:31,840
Picking the wrong one for a workload doesn't just make things slightly worse.
225
00:07:31,840 --> 00:07:34,320
It actively burns money or staff's capability
226
00:07:34,320 --> 00:07:36,000
depending on which way you get it wrong.
227
00:07:36,000 --> 00:07:38,520
Send a two sentence customer question through a Sparse,
228
00:07:38,520 --> 00:07:40,320
Frontier Scale Reasoning System
229
00:07:40,320 --> 00:07:42,000
and you're paying for a roster of specialists
230
00:07:42,000 --> 00:07:44,520
to show up for a meeting that needed one intern.
231
00:07:44,520 --> 00:07:47,360
Send a genuinely complex multi-step planning problem
232
00:07:47,360 --> 00:07:48,840
through a small dense model
233
00:07:48,840 --> 00:07:51,200
and you're asking a lean, fast team
234
00:07:51,200 --> 00:07:54,480
to solve something that actually needed deep, specialized reasoning
235
00:07:54,480 --> 00:07:56,400
they don't have on staff.
236
00:07:56,400 --> 00:07:58,880
Neither failure shows up as the model was bad.
237
00:07:58,880 --> 00:08:01,200
It shows up as a budget line that's too high
238
00:08:01,200 --> 00:08:04,160
or an output that's quietly, consistently shallow.
239
00:08:04,160 --> 00:08:06,520
That's the stakes, not which model wins a benchmark
240
00:08:06,520 --> 00:08:08,240
but which one you handed the wrong job.
241
00:08:08,240 --> 00:08:11,160
And this is where the abstract idea of dense versus Sparse
242
00:08:11,160 --> 00:08:13,400
turns into something you can actually put a number on
243
00:08:13,400 --> 00:08:16,120
because once you look at what these architectures cost to run
244
00:08:16,120 --> 00:08:18,240
and what they can actually hold in memory at once,
245
00:08:18,240 --> 00:08:21,480
the gap between reason and runtime stops being conceptual
246
00:08:21,480 --> 00:08:23,880
and starts being a line item.
247
00:08:23,880 --> 00:08:25,960
The compute economics nobody talks about,
248
00:08:25,960 --> 00:08:27,440
let's put actual numbers on this.
249
00:08:27,440 --> 00:08:29,800
The story of dense versus Sparse models
250
00:08:29,800 --> 00:08:32,560
only becomes real once you see what it costs to run them.
251
00:08:32,560 --> 00:08:33,880
Start with the context window.
252
00:08:33,880 --> 00:08:37,280
MI1 reportedly holds two million tokens in a single window
253
00:08:37,280 --> 00:08:39,840
which is roughly double what GPT-5 carries.
254
00:08:39,840 --> 00:08:41,800
But this isn't just about having more room.
255
00:08:41,800 --> 00:08:44,360
It is the difference between feeding a model a single chapter
256
00:08:44,360 --> 00:08:46,560
and feeding it the entire book, the footnotes
257
00:08:46,560 --> 00:08:49,040
and the author's previous three drafts all at once.
258
00:08:49,040 --> 00:08:50,920
That window matters for the exact jobs
259
00:08:50,920 --> 00:08:52,720
MI1 is supposed to handle.
260
00:08:52,720 --> 00:08:55,760
We are talking about sprawling multi-step reasoning problems.
261
00:08:55,760 --> 00:08:57,640
We're losing context halfway through means
262
00:08:57,640 --> 00:08:59,080
the entire answer falls apart.
263
00:08:59,080 --> 00:09:00,200
Now look at the price.
264
00:09:00,200 --> 00:09:04,680
MI1 runs around $0.001 per 1,000 tokens.
265
00:09:04,680 --> 00:09:07,480
5.4 has a marginal cost that is close to zero
266
00:09:07,480 --> 00:09:09,600
once it is sitting on your local hardware.
267
00:09:09,600 --> 00:09:12,840
It is not literally zero because electricity costs money
268
00:09:12,840 --> 00:09:14,280
but it is close enough that it stops
269
00:09:14,280 --> 00:09:15,840
feeling like a meter of expense.
270
00:09:15,840 --> 00:09:17,120
That is a massive gap.
271
00:09:17,120 --> 00:09:19,920
It is the difference between paying a monthly utility bill
272
00:09:19,920 --> 00:09:22,600
and using a subscription you already bought and paid for.
273
00:09:22,600 --> 00:09:25,040
And here is why that gap exists in the first place.
274
00:09:25,040 --> 00:09:26,560
MI1 is a cloud-hosted model
275
00:09:26,560 --> 00:09:29,200
so every single request has to travel to a Dytacenter
276
00:09:29,200 --> 00:09:31,640
and hit Microsoft's infrastructure to get billed.
277
00:09:31,640 --> 00:09:34,280
Even with Sparse activation keeping the costs down
278
00:09:34,280 --> 00:09:35,840
you are still paying for the trip.
279
00:09:35,840 --> 00:09:37,320
5.4 does not make that trip at all
280
00:09:37,320 --> 00:09:40,200
once it is loaded onto a laptop or a Windows endpoint.
281
00:09:40,200 --> 00:09:41,320
You already own the hardware
282
00:09:41,320 --> 00:09:43,240
and the model is just sitting there ready to work.
283
00:09:43,240 --> 00:09:44,960
The throughput number is back this up too.
284
00:09:44,960 --> 00:09:48,840
MI1 reportedly pushes around 145 tokens per second
285
00:09:48,840 --> 00:09:51,680
on Microsoft's own Maya 200 chips.
286
00:09:51,680 --> 00:09:53,040
That is a real performance number
287
00:09:53,040 --> 00:09:54,600
rather than a marketing figure.
288
00:09:54,600 --> 00:09:56,480
And it determines if a frontier class model
289
00:09:56,480 --> 00:09:58,160
is actually usable at scale.
290
00:09:58,160 --> 00:10:00,640
It is fast enough to serve real traffic in production
291
00:10:00,640 --> 00:10:03,040
instead of just looking good during a keynote demo.
292
00:10:03,040 --> 00:10:04,600
The reason that throughput exists
293
00:10:04,600 --> 00:10:06,800
traces back to how the model was trained.
294
00:10:06,800 --> 00:10:09,680
Microsoft claims they cut training costs by 40%
295
00:10:09,680 --> 00:10:11,800
by running on Maya co-designed infrastructure
296
00:10:11,800 --> 00:10:13,640
instead of standard Nvidia clusters
297
00:10:13,640 --> 00:10:14,960
that is not just a footnote.
298
00:10:14,960 --> 00:10:16,360
Training a frontier scale model
299
00:10:16,360 --> 00:10:18,680
is one of the largest expenses in this industry
300
00:10:18,680 --> 00:10:20,520
and saving 40% changes
301
00:10:20,520 --> 00:10:22,640
what Microsoft can afford to build next.
302
00:10:22,640 --> 00:10:25,320
It also changes how aggressively they can price the models
303
00:10:25,320 --> 00:10:26,760
they have already built.
304
00:10:26,760 --> 00:10:28,320
None of this is really a technical story.
305
00:10:28,320 --> 00:10:30,800
It is a business story wearing technical clothes.
306
00:10:30,800 --> 00:10:33,000
Cost per token is not an abstract metric
307
00:10:33,000 --> 00:10:34,760
that only engineers care about.
308
00:10:34,760 --> 00:10:37,600
It is the number that decides if a project is even viable.
309
00:10:37,600 --> 00:10:39,200
If you drop the cost far enough,
310
00:10:39,200 --> 00:10:41,760
a use case that used to be too expensive to automate
311
00:10:41,760 --> 00:10:44,120
suddenly gets greenlit without a second thought.
312
00:10:44,120 --> 00:10:47,160
If the cost stays high, even a useful feature gets shelved
313
00:10:47,160 --> 00:10:48,600
because the math does not work.
314
00:10:48,600 --> 00:10:51,360
Microsoft is not just cutting costs for the sake of it.
315
00:10:51,360 --> 00:10:53,040
They are expanding the list of problems
316
00:10:53,040 --> 00:10:55,040
that are actually worth solving with AI.
317
00:10:55,040 --> 00:10:57,200
But here is the catch with these efficiency numbers.
318
00:10:57,200 --> 00:10:59,440
A low cost per token and fast throughput
319
00:10:59,440 --> 00:11:01,240
do not mean anything on their own.
320
00:11:01,240 --> 00:11:03,600
They only matter once you know where each model is supposed
321
00:11:03,600 --> 00:11:05,880
to live and what kind of request is supposed to reach it.
322
00:11:05,880 --> 00:11:07,280
Cheap and fast is great,
323
00:11:07,280 --> 00:11:08,720
but cheap and fast in the wrong place
324
00:11:08,720 --> 00:11:11,120
is just a different way to waste your money.
325
00:11:11,120 --> 00:11:13,800
Five four as the runtime, what that actually means.
326
00:11:13,800 --> 00:11:15,440
Let's define the word runtime plainly
327
00:11:15,440 --> 00:11:17,640
because it is doing a lot of work in this episode.
328
00:11:17,640 --> 00:11:19,840
A runtime is the thing that executes.
329
00:11:19,840 --> 00:11:22,080
It is not the strategist deciding what should happen,
330
00:11:22,080 --> 00:11:24,080
but the part that actually does the work right
331
00:11:24,080 --> 00:11:25,000
where the action is.
332
00:11:25,000 --> 00:11:26,360
It happens instantly.
333
00:11:26,360 --> 00:11:28,400
There is no queue, no roundtrip,
334
00:11:28,400 --> 00:11:30,880
and no waiting on a data center somewhere to weigh in.
335
00:11:30,880 --> 00:11:33,280
That is the specific job fee four was built for.
336
00:11:33,280 --> 00:11:36,320
Look at the actual sizes and the picture gets concrete fast.
337
00:11:36,320 --> 00:11:39,080
Five four mini sits at 3.8 billion parameters,
338
00:11:39,080 --> 00:11:41,760
while five four multi-model sits at 5.6 billion.
339
00:11:41,760 --> 00:11:43,800
Both of them ship under the MIT license.
340
00:11:43,800 --> 00:11:45,600
That licensing detail sounds boring
341
00:11:45,600 --> 00:11:47,920
until you are the person who has to get a model approved
342
00:11:47,920 --> 00:11:48,800
for production.
343
00:11:48,800 --> 00:11:51,240
Most frontier models come wrapped in usage restrictions
344
00:11:51,240 --> 00:11:53,480
or revenue thresholds that need a lawyer sign off
345
00:11:53,480 --> 00:11:54,640
before anyone touches them.
346
00:11:54,640 --> 00:11:56,640
MIT licensing skips that entire process
347
00:11:56,640 --> 00:11:57,720
that there is no legal friction
348
00:11:57,720 --> 00:11:59,440
and no need to check with procurement
349
00:11:59,440 --> 00:12:01,040
before you embed the model.
350
00:12:01,040 --> 00:12:04,000
You can drop five four into a product or an internal tool
351
00:12:04,000 --> 00:12:07,040
without waiting on a contract review for a runtime layer.
352
00:12:07,040 --> 00:12:09,760
That matters because these tools are supposed to be everywhere,
353
00:12:09,760 --> 00:12:12,120
rather than gated behind a vendor agreement.
354
00:12:12,120 --> 00:12:14,840
Function calling is built directly into five four mini,
355
00:12:14,840 --> 00:12:16,760
and that is what makes it function as a runtime
356
00:12:16,760 --> 00:12:18,360
instead of just a small chatbot.
357
00:12:18,360 --> 00:12:20,200
The model is not just generating text
358
00:12:20,200 --> 00:12:23,120
because it can recognize when a task requires a decision point.
359
00:12:23,120 --> 00:12:26,760
It then hands off the task to a specific tool or function
360
00:12:26,760 --> 00:12:27,680
to act on it.
361
00:12:27,680 --> 00:12:29,160
That is the difference between a model
362
00:12:29,160 --> 00:12:30,680
that describes what should happen
363
00:12:30,680 --> 00:12:32,800
and one that can actually trigger the action.
364
00:12:32,800 --> 00:12:35,280
Small agents that need to make a call and act quickly
365
00:12:35,280 --> 00:12:37,640
run perfectly on this kind of setup,
366
00:12:37,640 --> 00:12:40,040
picture where this actually shows up in the real world.
367
00:12:40,040 --> 00:12:42,440
You could have a laptop running a local coding assistant
368
00:12:42,440 --> 00:12:44,400
that checks a file for errors and fixes them
369
00:12:44,400 --> 00:12:46,440
without ever sending that data anywhere.
370
00:12:46,440 --> 00:12:48,320
A Windows endpoint could run a small agent
371
00:12:48,320 --> 00:12:49,840
that triages an incoming request
372
00:12:49,840 --> 00:12:52,480
and decides which internal tool should handle it next.
373
00:12:52,480 --> 00:12:54,680
There is no cloud call, no latency spike,
374
00:12:54,680 --> 00:12:57,080
and no data leaving the machine during the process.
375
00:12:57,080 --> 00:12:59,640
That is execution living close to the work.
376
00:12:59,640 --> 00:13:02,160
The response feels instant because there is no travel time
377
00:13:02,160 --> 00:13:05,240
involved and the device has everything it needs locally.
378
00:13:05,240 --> 00:13:06,920
This is also where data sovereignty stops
379
00:13:06,920 --> 00:13:09,320
being an abstract conversation about compliance.
380
00:13:09,320 --> 00:13:11,160
If the model runs on the endpoint,
381
00:13:11,160 --> 00:13:13,240
the request never leaves the endpoint.
382
00:13:13,240 --> 00:13:15,640
That is not a policy you have to enforce after the fact.
383
00:13:15,640 --> 00:13:18,160
It is just how the architecture of the system works.
384
00:13:18,160 --> 00:13:20,080
But we have to be honest about the limits here.
385
00:13:20,080 --> 00:13:22,240
Execution alone does not explain intelligence.
386
00:13:22,240 --> 00:13:24,440
A runtime can act fast and act locally,
387
00:13:24,440 --> 00:13:26,480
but it cannot plan five moves ahead
388
00:13:26,480 --> 00:13:28,720
through a genuinely complicated problem.
389
00:13:28,720 --> 00:13:30,760
Five fork and catch an error or trigger a function
390
00:13:30,760 --> 00:13:32,200
to resolve a routine case.
391
00:13:32,200 --> 00:13:34,440
It cannot sit there and reason through something
392
00:13:34,440 --> 00:13:36,720
that requires holding a massive set of trade-offs
393
00:13:36,720 --> 00:13:37,840
in its head at once.
394
00:13:37,840 --> 00:13:39,000
That is not a flaw in the system.
395
00:13:39,000 --> 00:13:39,840
It is the design.
396
00:13:39,840 --> 00:13:42,520
The deep planning work was never supposed to live at the edge.
397
00:13:42,520 --> 00:13:44,040
That is the job for MI1.
398
00:13:44,040 --> 00:13:46,000
And we need to see what actual reason looks like
399
00:13:46,000 --> 00:13:48,160
once you get past the marketing label.
400
00:13:48,160 --> 00:13:51,040
MI1 as the reason, what that actually means.
401
00:13:51,040 --> 00:13:53,480
Reasoning means something very specific in this context.
402
00:13:53,480 --> 00:13:54,640
It isn't the part of the brain
403
00:13:54,640 --> 00:13:56,400
that remembers what happened yesterday.
404
00:13:56,400 --> 00:13:58,200
It's the part that decides what should happen next.
405
00:13:58,200 --> 00:14:00,320
It weighs options, it holds trade-offs,
406
00:14:00,320 --> 00:14:03,560
and it plans out five steps before a single line of code is written.
407
00:14:03,560 --> 00:14:05,680
That is a different job than what five orders.
408
00:14:05,680 --> 00:14:07,040
And because the job is different,
409
00:14:07,040 --> 00:14:08,880
it needs a different model underneath it.
410
00:14:08,880 --> 00:14:10,840
Microsoft put that job into my thinking one.
411
00:14:10,840 --> 00:14:13,000
It runs on 35 billion active parameters
412
00:14:13,000 --> 00:14:16,520
with a context window of 256,000 tokens.
413
00:14:16,520 --> 00:14:18,480
On the AIME 2025 math benchmark,
414
00:14:18,480 --> 00:14:20,040
it scores 97%.
415
00:14:20,040 --> 00:14:22,120
That isn't just a high score on a test.
416
00:14:22,120 --> 00:14:24,640
That specific benchmark is designed to trip models up
417
00:14:24,640 --> 00:14:26,840
with layered multi-step problems.
418
00:14:26,840 --> 00:14:30,320
Hitting 97% means the model can hold a complex problem together
419
00:14:30,320 --> 00:14:33,080
across a long chain of logic without losing the thread.
420
00:14:33,080 --> 00:14:36,120
But there is a detail here that matters more than the math scores.
421
00:14:36,120 --> 00:14:39,040
MI thinking one was built with zero distillation.
422
00:14:39,040 --> 00:14:41,160
In this industry, most models are built
423
00:14:41,160 --> 00:14:43,920
by training a small system to mimic a big one.
424
00:14:43,920 --> 00:14:46,600
You essentially teach the student to copy the teacher's homework.
425
00:14:46,600 --> 00:14:49,160
That's distillation, it's fast, but it has a hidden cost.
426
00:14:49,160 --> 00:14:51,120
You inherit every bias, every gap,
427
00:14:51,120 --> 00:14:52,760
and every licensing red flag
428
00:14:52,760 --> 00:14:54,640
that was baked into the original model.
429
00:14:54,640 --> 00:14:56,600
Zero distillation means none of that.
430
00:14:56,600 --> 00:14:58,880
MI thinking one was built from the ground up,
431
00:14:58,880 --> 00:15:01,960
using data with a clean, commercially licensed lineage.
432
00:15:01,960 --> 00:15:04,080
Nothing was borrowed from a black box source.
433
00:15:04,080 --> 00:15:07,200
If you work in a regulated industry like healthcare or finance,
434
00:15:07,200 --> 00:15:08,680
this isn't just a nice feature.
435
00:15:08,680 --> 00:15:11,840
It is the difference between a model your compliance team signs off on
436
00:15:11,840 --> 00:15:13,680
and one that gets stuck in review forever
437
00:15:13,680 --> 00:15:15,840
because nobody knows where its knowledge came from.
438
00:15:15,840 --> 00:15:17,440
This isn't just about math either.
439
00:15:17,440 --> 00:15:21,560
On SWEBench Pro, which tests real world software engineering,
440
00:15:21,560 --> 00:15:23,720
MI thinking one hits 53%.
441
00:15:23,720 --> 00:15:25,800
That puts it right next to Opus class models
442
00:15:25,800 --> 00:15:29,000
for a long time, Opus has been the gold standard for complex coding.
443
00:15:29,000 --> 00:15:32,400
Landing in that range proves MI one can handle the messy, ambiguous,
444
00:15:32,400 --> 00:15:35,000
multi-file reasoning that actual engineering requires.
445
00:15:35,000 --> 00:15:36,840
But we need to break a common instinct here.
446
00:15:36,840 --> 00:15:38,520
We've talked about how 5.4 is efficient,
447
00:15:38,520 --> 00:15:40,520
but bigger doesn't always mean better.
448
00:15:40,520 --> 00:15:42,440
Sparse activation proves that smaller footprints
449
00:15:42,440 --> 00:15:44,160
can outperform brute force.
450
00:15:44,160 --> 00:15:47,360
However, deeper reasoning still requires more active compute.
451
00:15:47,360 --> 00:15:48,960
That part is not negotiable.
452
00:15:48,960 --> 00:15:51,160
You cannot compress genuine, multi-step planning
453
00:15:51,160 --> 00:15:53,960
into a few billion parameters and expect the same result.
454
00:15:53,960 --> 00:15:56,960
Those 35 billion active parameters aren't there for show.
455
00:15:56,960 --> 00:15:59,800
They are the literal cost of holding a hard problem together
456
00:15:59,800 --> 00:16:01,080
long enough to solve it.
457
00:16:01,080 --> 00:16:03,360
This is where the two layers stop being separate stories.
458
00:16:03,360 --> 00:16:06,080
5.4 executes its fast, cheap, and local.
459
00:16:06,080 --> 00:16:08,720
MI one reasons its deep, deliberate, and expensive.
460
00:16:08,720 --> 00:16:10,200
Neither one replaces the other
461
00:16:10,200 --> 00:16:11,960
and neither one is complete on its own.
462
00:16:11,960 --> 00:16:14,440
The real architecture isn't about picking one model.
463
00:16:14,440 --> 00:16:15,560
It's about the handoff.
464
00:16:15,560 --> 00:16:18,240
It's about how a request starts at the runtime layer
465
00:16:18,240 --> 00:16:21,000
and moves to the reason layer the moment it needs real judgment.
466
00:16:21,000 --> 00:16:23,040
That handoff is the part nobody is talking about.
467
00:16:23,040 --> 00:16:25,880
And it's the most important piece of engineering to understand.
468
00:16:25,880 --> 00:16:28,360
The handoff, how reasoning becomes execution.
469
00:16:28,360 --> 00:16:30,200
Here's what that handoff looks like when you treat it
470
00:16:30,200 --> 00:16:32,000
as a system instead of an abstraction.
471
00:16:32,000 --> 00:16:33,120
A request comes in.
472
00:16:33,120 --> 00:16:36,040
Something has to look at it first before either model touches it
473
00:16:36,040 --> 00:16:37,760
to decide what it actually is.
474
00:16:37,760 --> 00:16:39,280
It isn't looking at the surface level.
475
00:16:39,280 --> 00:16:41,800
It's looking at what kind of problem it represents.
476
00:16:41,800 --> 00:16:44,400
Is this a two-line question with an obvious answer?
477
00:16:44,400 --> 00:16:46,040
Or is this something that needs five steps
478
00:16:46,040 --> 00:16:48,200
of planning before a response makes sense?
479
00:16:48,200 --> 00:16:49,760
That classification happens first.
480
00:16:49,760 --> 00:16:52,320
Then and only then the request gets routed.
481
00:16:52,320 --> 00:16:54,560
Simple, high volume tasks go to 5.4.
482
00:16:54,560 --> 00:16:55,840
It might be running on your device
483
00:16:55,840 --> 00:16:58,040
or hosted as a small model in Foundry.
484
00:16:58,040 --> 00:17:01,200
Complex multi-step reasoning gets escalated to MAI1.
485
00:17:01,200 --> 00:17:02,280
The pattern is simple.
486
00:17:02,280 --> 00:17:03,480
Classify, then route.
487
00:17:03,480 --> 00:17:04,560
But here's the problem.
488
00:17:04,560 --> 00:17:07,080
Everyone wants to talk about which model is smarter.
489
00:17:07,080 --> 00:17:09,400
Almost nobody wants to talk about the routing layer.
490
00:17:09,400 --> 00:17:11,520
This is the infrastructure that decides which model
491
00:17:11,520 --> 00:17:13,320
even sees the request in the first place.
492
00:17:13,320 --> 00:17:14,600
It isn't a glamorous build.
493
00:17:14,600 --> 00:17:16,160
It won't get a keynote moment.
494
00:17:16,160 --> 00:17:18,560
But it is the difference between a system that works
495
00:17:18,560 --> 00:17:20,720
and two expensive models sitting next to each other
496
00:17:20,720 --> 00:17:22,480
with no logic connecting them.
497
00:17:22,480 --> 00:17:24,200
Think about a support ticket system.
498
00:17:24,200 --> 00:17:25,920
80% of the tickets are routine.
499
00:17:25,920 --> 00:17:28,440
People want to reset passwords, check order status,
500
00:17:28,440 --> 00:17:29,720
or find an invoice.
501
00:17:29,720 --> 00:17:31,160
These questions have one clear answer
502
00:17:31,160 --> 00:17:32,840
and don't require weighing trade-offs.
503
00:17:32,840 --> 00:17:35,360
Those tickets resolve entirely at the 5.4 layer.
504
00:17:35,360 --> 00:17:37,560
They are fast, they are cheap, and they are done.
505
00:17:37,560 --> 00:17:39,240
The other 20% are different.
506
00:17:39,240 --> 00:17:42,120
Maybe it's a billing dispute that touches three different systems
507
00:17:42,120 --> 00:17:44,840
or a technical issue that needs a root cause analysis
508
00:17:44,840 --> 00:17:46,520
across a chain of dependencies.
509
00:17:46,520 --> 00:17:48,560
Those get escalated and MAI1 picks them up.
510
00:17:48,560 --> 00:17:50,160
That 80/20 split isn't a guess.
511
00:17:50,160 --> 00:17:52,680
It matches the data we see across the industry.
512
00:17:52,680 --> 00:17:55,080
Most requests are routine and only a small minority
513
00:17:55,080 --> 00:17:56,240
need deep reasoning.
514
00:17:56,240 --> 00:17:59,000
The routing layer's entire job is telling the difference
515
00:17:59,000 --> 00:18:02,480
correctly every single time without a human checking the work.
516
00:18:02,480 --> 00:18:04,400
This matters for more than just the budget.
517
00:18:04,400 --> 00:18:07,120
Yes, it's cheaper to solve routine tickets with a small model
518
00:18:07,120 --> 00:18:09,440
instead of paying frontier rates for a simple question.
519
00:18:09,440 --> 00:18:10,960
But the real win is latency.
520
00:18:10,960 --> 00:18:13,400
Think about the user experience for that 80%.
521
00:18:13,400 --> 00:18:15,920
They ask a question and the answer comes back almost instantly
522
00:18:15,920 --> 00:18:18,160
because the request never left the local layer.
523
00:18:18,160 --> 00:18:19,080
There is no queue.
524
00:18:19,080 --> 00:18:20,800
There is no wait for a massive system
525
00:18:20,800 --> 00:18:23,360
to spin up its experts for a question that didn't need them.
526
00:18:23,360 --> 00:18:25,280
The handoff to MI1 happens invisibly.
527
00:18:25,280 --> 00:18:28,480
It only happens for the cases that actually justify the wait.
528
00:18:28,480 --> 00:18:31,200
Most users will never notice a handoff even occurred.
529
00:18:31,200 --> 00:18:33,800
They just experience a system that feels fast all the time.
530
00:18:33,800 --> 00:18:35,960
It only takes longer on the genuinely hard stuff
531
00:18:35,960 --> 00:18:37,480
and they never have to know why.
532
00:18:37,480 --> 00:18:39,160
That invisibility is the whole point.
533
00:18:39,160 --> 00:18:41,640
A well-built routing layer doesn't announce itself.
534
00:18:41,640 --> 00:18:44,200
It makes the system feel like one single intelligence
535
00:18:44,200 --> 00:18:46,320
even though two different models are doing the work.
536
00:18:46,320 --> 00:18:47,600
And here is the thing to consider.
537
00:18:47,600 --> 00:18:50,120
This pattern of classifying, routing, and escalating
538
00:18:50,120 --> 00:18:52,920
only when necessary isn't something Microsoft invented
539
00:18:52,920 --> 00:18:53,880
in a vacuum.
540
00:18:53,880 --> 00:18:55,960
It isn't unique to MI1 and 5.4.
541
00:18:55,960 --> 00:18:57,240
This pattern is showing up everywhere.
542
00:18:57,240 --> 00:18:59,320
Company after company is moving toward this,
543
00:18:59,320 --> 00:19:01,280
regardless of which models they use.
544
00:19:01,280 --> 00:19:02,560
And that raises a real question.
545
00:19:02,560 --> 00:19:04,800
If everyone is converging on the same architecture,
546
00:19:04,800 --> 00:19:07,320
what does that tell you about where this is headed?
547
00:19:07,320 --> 00:19:09,560
Why the industry already agrees on this split?
548
00:19:09,560 --> 00:19:12,360
This shift is bigger than one company's product strategy.
549
00:19:12,360 --> 00:19:15,080
Gardner is projecting that by 2027,
550
00:19:15,080 --> 00:19:18,840
organizations will use small, task-specific models three times
551
00:19:18,840 --> 00:19:21,960
more often than general-purpose LLMs, not slightly more.
552
00:19:21,960 --> 00:19:22,880
Three times.
553
00:19:22,880 --> 00:19:25,200
That isn't a niche trend tucked into a research footnote.
554
00:19:25,200 --> 00:19:27,720
It's analysts describing what's about to become the default
555
00:19:27,720 --> 00:19:29,520
shape of enterprise AI.
556
00:19:29,520 --> 00:19:32,200
The do everything models, stop being the norm.
557
00:19:32,200 --> 00:19:34,040
The task-specific small models take over
558
00:19:34,040 --> 00:19:35,680
the majority of the workload.
559
00:19:35,680 --> 00:19:37,120
And this isn't some future prediction
560
00:19:37,120 --> 00:19:38,680
with nothing behind it yet.
561
00:19:38,680 --> 00:19:41,400
In AI-major markets, 68% of enterprises
562
00:19:41,400 --> 00:19:43,720
are already running at least one small model
563
00:19:43,720 --> 00:19:44,640
in production today.
564
00:19:44,640 --> 00:19:46,560
Not piloting, not testing in a sandbox,
565
00:19:46,560 --> 00:19:49,240
running it in production right now.
566
00:19:49,240 --> 00:19:51,000
That number alone should tell you the industry
567
00:19:51,000 --> 00:19:53,600
didn't wait around for a keynote to figure this out.
568
00:19:53,600 --> 00:19:56,280
Companies were already building toward exactly this split
569
00:19:56,280 --> 00:20:00,000
before Microsoft ever said the words MI1 or 5-4 out loud.
570
00:20:00,000 --> 00:20:03,120
Here's the economic driver underneath all of it, stated plainly.
571
00:20:03,120 --> 00:20:05,040
Serving a 7-billion parameter model
572
00:20:05,040 --> 00:20:07,240
can be 10 to 30 times cheaper than serving a model
573
00:20:07,240 --> 00:20:10,400
in the 70-billion plus range, 10 to 30 times.
574
00:20:10,400 --> 00:20:12,880
That isn't a rounding error you absorb into overhead.
575
00:20:12,880 --> 00:20:15,560
It's the difference between a workload that scales sustainably
576
00:20:15,560 --> 00:20:18,960
and one that gets quietly killed in a budget review six months in.
577
00:20:18,960 --> 00:20:20,840
Because the ongoing cost never made sense
578
00:20:20,840 --> 00:20:22,120
once volume actually showed up.
579
00:20:22,120 --> 00:20:24,360
Once you see that math, the industry wide shift
580
00:20:24,360 --> 00:20:25,760
stops looking like a trend.
581
00:20:25,760 --> 00:20:28,880
It starts looking like the only rational move available.
582
00:20:28,880 --> 00:20:31,560
If a small model handles the routine 80% of requests
583
00:20:31,560 --> 00:20:33,000
that are fraction of the cost,
584
00:20:33,000 --> 00:20:34,680
and the frontier model only gets pulled in
585
00:20:34,680 --> 00:20:36,480
for the genuinely hard 20%,
586
00:20:36,480 --> 00:20:38,520
you aren't choosing between quality and savings,
587
00:20:38,520 --> 00:20:39,520
you're getting both.
588
00:20:39,520 --> 00:20:42,200
Because you stopped paying frontier prices for questions
589
00:20:42,200 --> 00:20:44,640
that never needed frontier reasoning in the first place.
590
00:20:44,640 --> 00:20:46,400
So here's the point worth being honest about.
591
00:20:46,400 --> 00:20:48,080
Microsoft didn't invent this pattern.
592
00:20:48,080 --> 00:20:49,560
This isn't some proprietary insight
593
00:20:49,560 --> 00:20:51,080
that only came out of a Redmond Lab.
594
00:20:51,080 --> 00:20:54,200
What Microsoft did was productize it at their own scale,
595
00:20:54,200 --> 00:20:55,560
running on their own silicon.
596
00:20:55,560 --> 00:20:59,040
Maya chips underneath MIA1, MIT licensed 5-4 models
597
00:20:59,040 --> 00:21:00,320
built to sit on endpoints,
598
00:21:00,320 --> 00:21:02,040
and the whole reason and runtime split
599
00:21:02,040 --> 00:21:04,080
dressed up in shipped as a coordinated platform.
600
00:21:04,080 --> 00:21:06,440
But the underlying logic, small models for volume,
601
00:21:06,440 --> 00:21:07,760
large models for depth,
602
00:21:07,760 --> 00:21:10,280
was already the direction the entire industry was moving.
603
00:21:10,280 --> 00:21:11,920
Microsoft just built the most visible,
604
00:21:11,920 --> 00:21:14,040
most vertically integrated version of it,
605
00:21:14,040 --> 00:21:15,480
which raises an uncomfortable question
606
00:21:15,480 --> 00:21:17,240
for anyone who hasn't caught up yet.
607
00:21:17,240 --> 00:21:20,000
If this split is already the industry consensus,
608
00:21:20,000 --> 00:21:22,560
and the data shows 68% adoption
609
00:21:22,560 --> 00:21:25,400
with a three-fold shift coming within a couple of years,
610
00:21:25,400 --> 00:21:27,080
what does it actually cost a company
611
00:21:27,080 --> 00:21:29,360
that hasn't built this architecture on purpose?
612
00:21:29,360 --> 00:21:31,880
What happens when every request, simple or complex,
613
00:21:31,880 --> 00:21:34,640
still gets rooted through the exact same frontier model?
614
00:21:34,640 --> 00:21:36,440
Because nobody ever built the routing logic
615
00:21:36,440 --> 00:21:37,600
to do anything else?
616
00:21:37,600 --> 00:21:38,840
That isn't a hypothetical.
617
00:21:38,840 --> 00:21:40,840
That's the default state most organizations
618
00:21:40,840 --> 00:21:42,080
are still sitting in right now.
619
00:21:42,080 --> 00:21:46,120
And it's worth looking at exactly what that default costs.
620
00:21:46,120 --> 00:21:48,280
The old model, one model to rule everything,
621
00:21:48,280 --> 00:21:49,960
here's what that default actually looks like
622
00:21:49,960 --> 00:21:52,040
inside a company that never built the split,
623
00:21:52,040 --> 00:21:53,520
every request comes in,
624
00:21:53,520 --> 00:21:55,360
and every request goes to the same place,
625
00:21:55,360 --> 00:21:58,520
a frontier model, general purpose, expensive.
626
00:21:58,520 --> 00:22:00,800
Handling a two-sentence password reset
627
00:22:00,800 --> 00:22:03,360
the exact same way it handles a genuinely hard
628
00:22:03,360 --> 00:22:04,720
multi-step planning problem.
629
00:22:04,720 --> 00:22:07,400
There's no classification step, no routing logic,
630
00:22:07,400 --> 00:22:09,440
no decision about what a request actually needs.
631
00:22:09,440 --> 00:22:11,880
There's just one door, and everything walks through it.
632
00:22:11,880 --> 00:22:13,440
This is the floor default,
633
00:22:13,440 --> 00:22:16,160
and it's still the reality for most organizations right now,
634
00:22:16,160 --> 00:22:18,080
not because anyone chose it deliberately,
635
00:22:18,080 --> 00:22:19,920
but because nobody built anything else.
636
00:22:19,920 --> 00:22:21,800
When you only have one model available,
637
00:22:21,800 --> 00:22:24,280
route everything through it isn't a strategy.
638
00:22:24,280 --> 00:22:26,040
It's just what happens by omission.
639
00:22:26,040 --> 00:22:27,760
Here's what that omission costs.
640
00:22:27,760 --> 00:22:29,320
Companies stuck on this pattern report
641
00:22:29,320 --> 00:22:33,720
monthly cloud AI bills running $50,000 to $100,000 or more.
642
00:22:33,720 --> 00:22:36,160
For workloads that never needed frontier level reasoning
643
00:22:36,160 --> 00:22:37,200
in the first place.
644
00:22:37,200 --> 00:22:39,720
That isn't the cost of doing hard, valuable work.
645
00:22:39,720 --> 00:22:41,680
It's the cost of asking an expensive specialist
646
00:22:41,680 --> 00:22:43,960
to sit through a meeting that only needed an intern.
647
00:22:43,960 --> 00:22:45,760
Over and over, thousands of times a month
648
00:22:45,760 --> 00:22:47,640
because there was no cheaper option on the roster,
649
00:22:47,640 --> 00:22:49,560
and the cost isn't only financial.
650
00:22:49,560 --> 00:22:51,640
There's a latency tax buried in here too,
651
00:22:51,640 --> 00:22:53,440
one that's easy to miss because it doesn't show up
652
00:22:53,440 --> 00:22:54,400
on an invoice.
653
00:22:54,400 --> 00:22:56,600
Picture a model built to hold a two million token
654
00:22:56,600 --> 00:22:58,720
context window, architected to reason
655
00:22:58,720 --> 00:23:01,520
across entire code bases and sprawling document sets,
656
00:23:01,520 --> 00:23:03,840
and then someone asks it a two-sentence question.
657
00:23:03,840 --> 00:23:05,240
The system still has to spin up,
658
00:23:05,240 --> 00:23:07,400
it still has to prepare for the possibility of holding
659
00:23:07,400 --> 00:23:09,920
that much context, even though this particular request
660
00:23:09,920 --> 00:23:11,200
needed almost none of it.
661
00:23:11,200 --> 00:23:12,400
The user waits on infrastructure
662
00:23:12,400 --> 00:23:15,200
that was built for a different kind of problem entirely.
663
00:23:15,200 --> 00:23:16,840
Here's the belief worth breaking,
664
00:23:16,840 --> 00:23:19,040
and it's an uncomfortable one for a lot of teams.
665
00:23:19,040 --> 00:23:21,040
When the setup feels slow or expensive,
666
00:23:21,040 --> 00:23:22,600
the instinct is to blame the model.
667
00:23:22,600 --> 00:23:24,760
Upgraded, tune it, throw more compute at it,
668
00:23:24,760 --> 00:23:26,480
but that's misdiagnosing the problem.
669
00:23:26,480 --> 00:23:29,480
This isn't a technology failure, it's an architecture failure.
670
00:23:29,480 --> 00:23:30,960
The model was never the bottleneck.
671
00:23:30,960 --> 00:23:32,680
The bottleneck was the absence of a system
672
00:23:32,680 --> 00:23:34,120
that knew which request deserved
673
00:23:34,120 --> 00:23:36,800
that model's full capability and which ones didn't.
674
00:23:36,800 --> 00:23:38,960
A frontier model performing exactly as designed
675
00:23:38,960 --> 00:23:41,560
on a task that never needed frontier level reasoning
676
00:23:41,560 --> 00:23:42,520
isn't a broken model.
677
00:23:42,520 --> 00:23:45,240
It's a broken decision about where that model belongs.
678
00:23:45,240 --> 00:23:47,640
No amount of upgrading fixes a routing problem
679
00:23:47,640 --> 00:23:49,800
because the model was never what needed fixing.
680
00:23:49,800 --> 00:23:52,960
This is exactly the Gap MI1 and 5.4 were built to close.
681
00:23:52,960 --> 00:23:55,040
Not by making one model do everything better,
682
00:23:55,040 --> 00:23:57,120
but by making sure everything doesn't have to go through
683
00:23:57,120 --> 00:23:58,760
one model in the first place.
684
00:23:58,760 --> 00:24:01,720
To understand how that Gap actually gets closed in practice,
685
00:24:01,720 --> 00:24:04,400
it helps to look inside what each of these two models
686
00:24:04,400 --> 00:24:07,080
is actually built from, starting with the one designed
687
00:24:07,080 --> 00:24:09,160
to live closest to the work.
688
00:24:09,160 --> 00:24:11,400
Inside 5.4's actual architecture.
689
00:24:11,400 --> 00:24:14,560
Let's open this up and look at what is sitting inside 5.4
690
00:24:14,560 --> 00:24:16,680
because the architecture explains why it can live
691
00:24:16,680 --> 00:24:18,520
at the edge in the first place.
692
00:24:18,520 --> 00:24:20,720
It starts with grouped query attention.
693
00:24:20,720 --> 00:24:23,120
In a normal transformer, every single query head
694
00:24:23,120 --> 00:24:25,120
does its own separate lookup work.
695
00:24:25,120 --> 00:24:28,160
It is thorough, but it is also expensive for your memory.
696
00:24:28,160 --> 00:24:29,800
Grouped query attention changes that
697
00:24:29,800 --> 00:24:32,800
by having clusters of query heads share the same lookup work
698
00:24:32,800 --> 00:24:34,520
instead of each one duplicating it.
699
00:24:34,520 --> 00:24:37,760
The same basic job gets done, but you spend less memory doing it.
700
00:24:37,760 --> 00:24:39,920
This is not a shortcut that hurts quality.
701
00:24:39,920 --> 00:24:41,760
It is a shortcut that removes waste
702
00:24:41,760 --> 00:24:43,360
that was never buying you anything.
703
00:24:43,360 --> 00:24:45,680
When you pair that with a 200,000 word vocabulary
704
00:24:45,680 --> 00:24:48,240
and shared input output embeddings, the model changes.
705
00:24:48,240 --> 00:24:50,400
A bigger vocabulary means the model can represent
706
00:24:50,400 --> 00:24:52,320
language more precisely without breaking words
707
00:24:52,320 --> 00:24:53,720
into awkward fragments.
708
00:24:53,720 --> 00:24:55,760
By sharing the embeddings between input and output,
709
00:24:55,760 --> 00:24:57,840
the model does not have to store two separate copies
710
00:24:57,840 --> 00:25:00,760
of the same vocabulary knowledge for reading and writing.
711
00:25:00,760 --> 00:25:02,800
It reuses the same table for both directions.
712
00:25:02,800 --> 00:25:05,440
This means fewer parameters are spent on redundancy
713
00:25:05,440 --> 00:25:06,880
and more of the model's budget goes
714
00:25:06,880 --> 00:25:09,000
toward actually understanding what you are asking.
715
00:25:09,000 --> 00:25:11,520
The multi-modal version follows a similar logic.
716
00:25:11,520 --> 00:25:13,760
5.4 multi-modal is not just a text model
717
00:25:13,760 --> 00:25:16,120
with a camera bolted on as an afterthought.
718
00:25:16,120 --> 00:25:19,000
It runs a vision encoder and an audio encoder side by side
719
00:25:19,000 --> 00:25:21,440
and both feed into that same compact backbone
720
00:25:21,440 --> 00:25:23,800
using a mixture of low-res approach.
721
00:25:23,800 --> 00:25:25,720
Instead of training three entirely separate models
722
00:25:25,720 --> 00:25:28,800
for text, image and audio, Microsoft trains small,
723
00:25:28,800 --> 00:25:30,920
lightweight adapters that plug into the core model
724
00:25:30,920 --> 00:25:32,920
depending on what kind of input shows up.
725
00:25:32,920 --> 00:25:35,240
You have one backbone and three sets of ears
726
00:25:35,240 --> 00:25:37,920
and each one only activates when it is actually needed.
727
00:25:37,920 --> 00:25:40,320
There is one detail here that is genuinely surprising.
728
00:25:40,320 --> 00:25:42,480
5.4 was trained on five trillion tokens.
729
00:25:42,480 --> 00:25:44,440
That number sounds huge until you compare it
730
00:25:44,440 --> 00:25:46,840
to frontier scale models that often train
731
00:25:46,840 --> 00:25:48,520
on two or three times that amount.
732
00:25:48,520 --> 00:25:51,520
5.4 is working with a noticeably smaller training diet
733
00:25:51,520 --> 00:25:53,160
than the giants it competes with.
734
00:25:53,160 --> 00:25:54,680
So why does it still perform well?
735
00:25:54,680 --> 00:25:56,560
It works because those five trillion tokens
736
00:25:56,560 --> 00:25:58,680
are heavily synthetic and heavily curated.
737
00:25:58,680 --> 00:26:01,120
Microsoft is not just scraping the open internet
738
00:26:01,120 --> 00:26:03,160
and hoping the quality averages out.
739
00:26:03,160 --> 00:26:05,080
They are generating targeted training data
740
00:26:05,080 --> 00:26:07,080
specifically designed to teach the model
741
00:26:07,080 --> 00:26:09,040
math coding and reasoning patterns.
742
00:26:09,040 --> 00:26:11,640
They filter hard for quality before any of it gets used.
743
00:26:11,640 --> 00:26:13,200
This is quality of a volume.
744
00:26:13,200 --> 00:26:15,680
A frontier model might need 10 times the raw data
745
00:26:15,680 --> 00:26:17,320
because a huge chunk of what it eats
746
00:26:17,320 --> 00:26:19,040
is noisy or repetitive.
747
00:26:19,040 --> 00:26:21,240
5.4's approach front loads the curation work
748
00:26:21,240 --> 00:26:23,680
so every token it sees is pulling more weight.
749
00:26:23,680 --> 00:26:25,560
You have less data but that data was built
750
00:26:25,560 --> 00:26:28,160
to teach exactly what the model needed to learn.
751
00:26:28,160 --> 00:26:30,160
Then there is SambaY.
752
00:26:30,160 --> 00:26:32,320
This variant is a genuine architectural departure
753
00:26:32,320 --> 00:26:34,080
rather than just a size adjustment.
754
00:26:34,080 --> 00:26:35,960
SambaY is a hybrid design that blends
755
00:26:35,960 --> 00:26:39,160
Mamba style state space layers in with traditional attention.
756
00:26:39,160 --> 00:26:41,640
State space layers process long sequences differently
757
00:26:41,640 --> 00:26:44,120
than attention does and they are often more efficient
758
00:26:44,120 --> 00:26:45,440
for certain kinds of context.
759
00:26:45,440 --> 00:26:48,000
By combining the two, Microsoft is testing
760
00:26:48,000 --> 00:26:49,960
whether the next efficiency gain comes
761
00:26:49,960 --> 00:26:51,760
from a smarter core architecture
762
00:26:51,760 --> 00:26:54,240
instead of just smarter training data.
763
00:26:54,240 --> 00:26:55,800
When you put all of this together,
764
00:26:55,800 --> 00:26:58,560
you get a model built to run at the edge
765
00:26:58,560 --> 00:27:00,720
on modest hardware without the cloud.
766
00:27:00,720 --> 00:27:03,080
But a runtime that lives at the edge only matters
767
00:27:03,080 --> 00:27:04,840
if it can connect to something bigger
768
00:27:04,840 --> 00:27:06,200
when a request outgrows it.
769
00:27:06,200 --> 00:27:09,000
That is where MI1's architecture comes into the picture.
770
00:27:09,000 --> 00:27:11,160
Inside MI1's actual architecture.
771
00:27:11,160 --> 00:27:13,320
Now we flip to the other side of the stack.
772
00:27:13,320 --> 00:27:15,920
MI1's architecture is solving a completely different problem
773
00:27:15,920 --> 00:27:17,520
than the one 5.4 just solved.
774
00:27:17,520 --> 00:27:19,720
It uses a sparse mixture of expert system.
775
00:27:19,720 --> 00:27:21,560
Imagine a roster of specialists.
776
00:27:21,560 --> 00:27:24,040
Way more than you would ever need in one room at once.
777
00:27:24,040 --> 00:27:26,520
Where each one is trained on a different slice of knowledge.
778
00:27:26,520 --> 00:27:28,600
When a request comes in, the system does not wake up
779
00:27:28,600 --> 00:27:29,600
the whole roster.
780
00:27:29,600 --> 00:27:31,680
It picks a small handful of experts relevant
781
00:27:31,680 --> 00:27:35,000
to that specific question and only those experts do any work.
782
00:27:35,000 --> 00:27:37,080
Everyone else on the roster stays idle.
783
00:27:37,080 --> 00:27:39,800
That is the many experts, few active structure.
784
00:27:39,800 --> 00:27:42,280
You have massive total capacity sitting on the shelf
785
00:27:42,280 --> 00:27:45,120
but only a sliver of it is switched on for any single token.
786
00:27:45,120 --> 00:27:46,920
This is the key to controlling cost.
787
00:27:46,920 --> 00:27:49,240
A dense model of a similar size would have to activate
788
00:27:49,240 --> 00:27:51,200
everything for every single request.
789
00:27:51,200 --> 00:27:53,040
That is fine when the total size is small
790
00:27:53,040 --> 00:27:55,480
but it stops being fine once the model gets frontier large.
791
00:27:55,480 --> 00:27:57,480
At that scale you would be paying full compute
792
00:27:57,480 --> 00:27:59,720
for the entire roster on every query
793
00:27:59,720 --> 00:28:02,800
whether the question needed three specialists or 300.
794
00:28:02,800 --> 00:28:05,160
Sparse activation breaks that link.
795
00:28:05,160 --> 00:28:07,680
MI1 can carry an enormous total parameter count
796
00:28:07,680 --> 00:28:10,520
without paying full price for that capacity on every request.
797
00:28:10,520 --> 00:28:13,480
You get the scale without inheriting the normal price tag.
798
00:28:13,480 --> 00:28:15,240
Then you have to layer in the hardware side.
799
00:28:15,240 --> 00:28:18,680
Microsoft co-designed MI1 with its own Maya 200 chips
800
00:28:18,680 --> 00:28:21,600
and the reported payoff is a 1.4 times gain in performance
801
00:28:21,600 --> 00:28:23,520
per what compared to generic GPUs.
802
00:28:23,520 --> 00:28:25,200
This is not just a small tuning win.
803
00:28:25,200 --> 00:28:27,960
It is the difference between squeezing gains out of hardware
804
00:28:27,960 --> 00:28:29,640
you did not build for this job
805
00:28:29,640 --> 00:28:31,480
and using hardware that was shaped around
806
00:28:31,480 --> 00:28:34,200
the model's actual activation pattern from the start.
807
00:28:34,200 --> 00:28:36,120
This co-design is a strategic move.
808
00:28:36,120 --> 00:28:37,400
Not just an engineering flex,
809
00:28:37,400 --> 00:28:38,920
anyone can rent GPU capacity
810
00:28:38,920 --> 00:28:41,680
because that is a commodity available to whoever has the budget.
811
00:28:41,680 --> 00:28:43,480
But designing your chip and your model together
812
00:28:43,480 --> 00:28:46,360
so the activation pattern maps onto the silicon underneath it
813
00:28:46,360 --> 00:28:48,640
is not something a competitor can just buy.
814
00:28:48,640 --> 00:28:51,080
It is a mode built out of years of coordinated engineering
815
00:28:51,080 --> 00:28:53,640
between teams that do not normally sit in the same building.
816
00:28:53,640 --> 00:28:55,640
Copying the model architecture is possible
817
00:28:55,640 --> 00:28:58,040
but copying the chip model relationship behind it
818
00:28:58,040 --> 00:28:59,240
is a much longer road.
819
00:28:59,240 --> 00:29:02,320
And that is exactly the point Suleiman has been making.
820
00:29:02,320 --> 00:29:04,000
The goal is not just building good models.
821
00:29:04,000 --> 00:29:05,640
It is AI self-sufficiency.
822
00:29:05,640 --> 00:29:07,160
Microsoft wants to reach a point
823
00:29:07,160 --> 00:29:09,800
where they are not permanently renting someone else's compute
824
00:29:09,800 --> 00:29:11,640
to run their own intelligence layer.
825
00:29:11,640 --> 00:29:14,200
Owning the chip, the model and the relationship between them
826
00:29:14,200 --> 00:29:17,240
is what makes self-sufficiency more than a talking point.
827
00:29:17,240 --> 00:29:19,840
On paper, this is a genuinely elegant design.
828
00:29:19,840 --> 00:29:22,040
You have sparse activation for cost control,
829
00:29:22,040 --> 00:29:23,760
custom silicon for efficiency
830
00:29:23,760 --> 00:29:26,000
and a strategic goal tying it all together.
831
00:29:26,000 --> 00:29:28,800
But an architecture diagram is not proof of anything.
832
00:29:28,800 --> 00:29:29,920
None of this actually matters
833
00:29:29,920 --> 00:29:32,280
unless it holds up outside a keynote slide.
834
00:29:32,280 --> 00:29:34,280
Under real workloads with real deadlines
835
00:29:34,280 --> 00:29:37,680
and real budgets on the line, the Excel proof point.
836
00:29:37,680 --> 00:29:40,200
So here is what all of that architecture actually produced
837
00:29:40,200 --> 00:29:42,360
when Microsoft pointed it at a real problem
838
00:29:42,360 --> 00:29:44,120
instead of a benchmark chart.
839
00:29:44,120 --> 00:29:46,160
Microsoft wanted agentic Excel tasks
840
00:29:46,160 --> 00:29:47,960
that actually held up under real use.
841
00:29:47,960 --> 00:29:50,440
They didn't want a demo where the model reads a spreadsheet
842
00:29:50,440 --> 00:29:53,720
and summarizes it once for an audience that claps and moves on.
843
00:29:53,720 --> 00:29:55,560
They needed real agentic work.
844
00:29:55,560 --> 00:29:57,600
The kind where a model has to look at a workbook
845
00:29:57,600 --> 00:29:59,600
understand what is actually being asked,
846
00:29:59,600 --> 00:30:02,400
take multiple steps and get it right every single time.
847
00:30:02,400 --> 00:30:04,400
And it has to do that across the massive volume
848
00:30:04,400 --> 00:30:07,720
of requests an actual Excel user-based generates every day.
849
00:30:07,720 --> 00:30:08,800
But here is the problem.
850
00:30:08,800 --> 00:30:11,440
General purpose frontier models could technically do this.
851
00:30:11,440 --> 00:30:14,120
But at the volume Excel operates at hundreds of millions
852
00:30:14,120 --> 00:30:17,000
of users and an enormous number of daily requests.
853
00:30:17,000 --> 00:30:19,240
Those models were either too slow to feel usable
854
00:30:19,240 --> 00:30:21,000
or too expensive to run at scale
855
00:30:21,000 --> 00:30:22,640
without the economics falling apart.
856
00:30:22,640 --> 00:30:25,000
You can build an impressive demo with a frontier model
857
00:30:25,000 --> 00:30:26,840
answering one spreadsheet question.
858
00:30:26,840 --> 00:30:28,720
But you cannot afford to run that same model
859
00:30:28,720 --> 00:30:30,920
at that same depth for every formula question
860
00:30:30,920 --> 00:30:33,080
and data cleanup task hitting Excel in production.
861
00:30:33,080 --> 00:30:35,360
So this is the exact gap we have been building toward.
862
00:30:35,360 --> 00:30:36,680
This isn't a hypothetical.
863
00:30:36,680 --> 00:30:39,440
It is an actual product decision Microsoft had to make.
864
00:30:39,440 --> 00:30:40,520
And here is the shift.
865
00:30:40,520 --> 00:30:44,600
M.A.I. tuned models ended up matching GPT 5.4 on the relevant bench
866
00:30:44,600 --> 00:30:47,200
marks while running a 10 times greater cost efficiency.
867
00:30:47,200 --> 00:30:48,200
Read that again.
868
00:30:48,200 --> 00:30:50,200
Because it is the whole thesis of this episode
869
00:30:50,200 --> 00:30:53,200
compressed into one result, it wasn't close enough performance
870
00:30:53,200 --> 00:30:54,440
for a discount.
871
00:30:54,440 --> 00:30:56,680
It was matching performance at a 10th of the cost.
872
00:30:56,680 --> 00:30:57,760
That is not a trade-off.
873
00:30:57,760 --> 00:31:00,200
That is what happens when a model gets tuned specifically
874
00:31:00,200 --> 00:31:03,120
for a task instead of being asked to be brilliant at everything.
875
00:31:03,120 --> 00:31:04,360
And this was not a one-off.
876
00:31:04,360 --> 00:31:07,480
The same pattern repeated with McKinsey's task-specific tuning.
877
00:31:07,480 --> 00:31:10,160
When M.A.I. got tuned on McKinsey's actual workflows,
878
00:31:10,160 --> 00:31:13,000
it delivered the highest win rate of any model tested and out
879
00:31:13,000 --> 00:31:16,320
performed GPT 5.5 on their specific tasks.
880
00:31:16,320 --> 00:31:18,560
It landed at that same 10 times efficiency gain
881
00:31:18,560 --> 00:31:20,280
to completely different organizations,
882
00:31:20,280 --> 00:31:21,720
two completely different workloads.
883
00:31:21,720 --> 00:31:23,800
Spreadsheets on one side and consulting workflows
884
00:31:23,800 --> 00:31:24,480
on the other.
885
00:31:24,480 --> 00:31:26,080
The same result shows up both times.
886
00:31:26,080 --> 00:31:28,240
So what is actually happening is a pattern.
887
00:31:28,240 --> 00:31:30,320
One win could be a fluke or a bench mark
888
00:31:30,320 --> 00:31:32,480
that happened to favor M.A.I.I's training data.
889
00:31:32,480 --> 00:31:34,240
But two wins in unrelated domains
890
00:31:34,240 --> 00:31:36,520
landing on the same 10 times efficiency number
891
00:31:36,520 --> 00:31:37,800
means something.
892
00:31:37,800 --> 00:31:40,960
It is a sign the gain isn't specific to spreadsheets or consulting.
893
00:31:40,960 --> 00:31:42,920
It is what shows up whenever reasoning and runtime
894
00:31:42,920 --> 00:31:45,200
get tuned together against a real narrow workflow
895
00:31:45,200 --> 00:31:47,280
instead of being asked to perform generally.
896
00:31:47,280 --> 00:31:48,600
And that is the real point here.
897
00:31:48,600 --> 00:31:50,120
This is not a demo.
898
00:31:50,120 --> 00:31:52,520
Demo's show you what is possible in ideal conditions.
899
00:31:52,520 --> 00:31:55,080
This is what happens when the architecture we have described
900
00:31:55,080 --> 00:31:57,120
gets pointed at actual production workloads
901
00:31:57,120 --> 00:31:59,520
with actual volume and actual cost pressure.
902
00:31:59,520 --> 00:32:02,000
These are real users who would notice if it broke,
903
00:32:02,000 --> 00:32:03,680
which raises a question worth sitting with.
904
00:32:03,680 --> 00:32:05,400
If tuning a model on your own workflows
905
00:32:05,400 --> 00:32:06,920
gets your results like this.
906
00:32:06,920 --> 00:32:08,800
What does that actually mean for who ends up owning
907
00:32:08,800 --> 00:32:11,240
the model that comes out the other side?
908
00:32:11,240 --> 00:32:13,880
Frontier tuning and wired changes who owns the model.
909
00:32:13,880 --> 00:32:15,840
The answer starts with a piece of infrastructure
910
00:32:15,840 --> 00:32:18,320
Microsoft calls reinforcement learning environments.
911
00:32:18,320 --> 00:32:19,520
Or RLEs.
912
00:32:19,520 --> 00:32:21,600
Think of an RLE as a training gym built
913
00:32:21,600 --> 00:32:23,200
for exactly one company's problems.
914
00:32:23,200 --> 00:32:25,040
It is not a general gym with treadmills
915
00:32:25,040 --> 00:32:26,520
for anybody off the street.
916
00:32:26,520 --> 00:32:29,360
It is a facility custom built around the specific movements
917
00:32:29,360 --> 00:32:31,600
one team actually needs to get good at.
918
00:32:31,600 --> 00:32:34,520
Microsoft used its own RLEs combined with M.A.I.I models
919
00:32:34,520 --> 00:32:36,680
to climb to what better performance on Excel tasks
920
00:32:36,680 --> 00:32:37,680
specifically.
921
00:32:37,680 --> 00:32:39,720
McKinsey did the same thing with their own tasks
922
00:32:39,720 --> 00:32:41,040
in their own environment.
923
00:32:41,040 --> 00:32:42,240
The gym isn't shared.
924
00:32:42,240 --> 00:32:44,720
It is built around one organization's actual work.
925
00:32:44,720 --> 00:32:46,480
And the model only gets strong at the things
926
00:32:46,480 --> 00:32:47,680
that gym trains it on.
927
00:32:47,680 --> 00:32:49,440
But here is the differentiator that matters more
928
00:32:49,440 --> 00:32:50,480
than anything else.
929
00:32:50,480 --> 00:32:52,240
You don't rent shared intelligence.
930
00:32:52,240 --> 00:32:53,400
You keep the resulting model.
931
00:32:53,400 --> 00:32:54,520
Sit with that for a second.
932
00:32:54,520 --> 00:32:56,400
Because it is a genuinely different relationship
933
00:32:56,400 --> 00:32:58,480
than most companies have with A.I.I right now.
934
00:32:58,480 --> 00:33:00,640
Most SAS A.I.Tools work the opposite way.
935
00:33:00,640 --> 00:33:01,440
You use the tool.
936
00:33:01,440 --> 00:33:03,320
Your usage feeds a shared system.
937
00:33:03,320 --> 00:33:05,400
And every improvement that comes from your data
938
00:33:05,400 --> 00:33:07,160
gets folded into the same general model
939
00:33:07,160 --> 00:33:09,240
every other customer is also using.
940
00:33:09,240 --> 00:33:11,480
Your competitor uses the same tool and benefits
941
00:33:11,480 --> 00:33:14,240
from the exact same improvements your usage helped create.
942
00:33:14,240 --> 00:33:15,960
You are not building anything proprietary.
943
00:33:15,960 --> 00:33:17,880
You are contributing to somebody else's product
944
00:33:17,880 --> 00:33:19,240
one query at a time.
945
00:33:19,240 --> 00:33:21,840
And you don't get to walk away with what you helped build.
946
00:33:21,840 --> 00:33:23,360
Frontier tuning breaks that arrangement.
947
00:33:23,360 --> 00:33:25,840
When you tune M.A.I.I models inside your own RLE
948
00:33:25,840 --> 00:33:28,280
on your own tasks, the resulting model is yours.
949
00:33:28,280 --> 00:33:30,600
It is not a shared checkpoint everyone draws from.
950
00:33:30,600 --> 00:33:33,000
It is a model shaped specifically by your workflows
951
00:33:33,000 --> 00:33:34,280
and your institutional know-how.
952
00:33:34,280 --> 00:33:36,800
Nobody else gets to touch what came out of that process.
953
00:33:36,800 --> 00:33:39,000
And this is the shift where your workflows become your mode
954
00:33:39,000 --> 00:33:40,240
instead of a vendor's mode.
955
00:33:40,240 --> 00:33:42,800
Every company has some version of tribal knowledge.
956
00:33:42,800 --> 00:33:45,120
The specific way they handle a certain kind of customer
957
00:33:45,120 --> 00:33:47,000
complaint or the particular judgment calls
958
00:33:47,000 --> 00:33:49,600
that separate a senior employee from a junior one.
959
00:33:49,600 --> 00:33:52,240
Normally that knowledge stays locked inside people's heads
960
00:33:52,240 --> 00:33:54,280
or scattered across documentation.
961
00:33:54,280 --> 00:33:55,200
Nobody reads.
962
00:33:55,200 --> 00:33:56,840
Frontier tuning turns that same knowledge
963
00:33:56,840 --> 00:33:58,280
into a model asset.
964
00:33:58,280 --> 00:34:00,560
The workflows that make your company good at what it does
965
00:34:00,560 --> 00:34:02,360
become the thing the model gets tuned on.
966
00:34:02,360 --> 00:34:04,240
The resulting intelligence belongs to you.
967
00:34:04,240 --> 00:34:06,240
Not to whoever built the underlying model.
968
00:34:06,240 --> 00:34:08,200
Now connect this back to the architecture we have been
969
00:34:08,200 --> 00:34:09,000
describing.
970
00:34:09,000 --> 00:34:10,840
Phi 4 becomes the tunable runtime.
971
00:34:10,840 --> 00:34:12,880
The fast and local layer that can be shaped around your
972
00:34:12,880 --> 00:34:14,520
specific execution needs.
973
00:34:14,520 --> 00:34:17,360
M.I.I.I.I sits underneath as the reasoning foundation.
974
00:34:17,360 --> 00:34:20,760
The deep layer custom agents draw on when a request actually
975
00:34:20,760 --> 00:34:22,800
needs planning instead of just action.
976
00:34:22,800 --> 00:34:25,320
Frontier tuning is what lets both layers get shaped around
977
00:34:25,320 --> 00:34:28,600
one company's actual work instead of staying generic and shared
978
00:34:28,600 --> 00:34:31,760
across every customer using the same off-the-shelf model.
979
00:34:31,760 --> 00:34:33,800
And this is where things change for IT decision-makers
980
00:34:33,800 --> 00:34:36,920
up until now choosing an AI vendor mostly meant choosing a tool.
981
00:34:36,920 --> 00:34:39,280
Now it means choosing whether the intelligence your company
982
00:34:39,280 --> 00:34:41,520
builds through daily use stays yours
983
00:34:41,520 --> 00:34:44,360
or quietly becomes part of somebody else's shared product.
984
00:34:44,360 --> 00:34:46,680
That is not a question engineering teams have historically
985
00:34:46,680 --> 00:34:49,000
had to ask, but it is about to become one of the most
986
00:34:49,000 --> 00:34:50,720
consequential calls they make.
987
00:34:50,720 --> 00:34:53,520
If this changed how you think, follow me, Mirko Peters,
988
00:34:53,520 --> 00:34:54,680
on LinkedIn.
989
00:34:54,680 --> 00:34:56,680
And if you want more of this, leave a review.
990
00:34:56,680 --> 00:34:57,880
It helps more people find it.
991
00:34:57,880 --> 00:35:00,400
Share this with your team, especially if you are dealing
992
00:35:00,400 --> 00:35:01,440
with this right now.
993
00:35:01,440 --> 00:35:03,800
What this means for IT architects and admins.
994
00:35:03,800 --> 00:35:05,680
So here is where this lands for the people who actually
995
00:35:05,680 --> 00:35:06,680
have to build it.
996
00:35:06,680 --> 00:35:09,360
For years, the standard IT question was simple.
997
00:35:09,360 --> 00:35:11,000
Which model should we standardize on?
998
00:35:11,000 --> 00:35:12,680
You pick the vendor, you pick the model,
999
00:35:12,680 --> 00:35:15,200
you roll it out across the org, and then everyone
1000
00:35:15,200 --> 00:35:16,480
built against that one choice.
1001
00:35:16,480 --> 00:35:18,880
That question made sense when there was only one kind of model
1002
00:35:18,880 --> 00:35:19,760
to pick from.
1003
00:35:19,760 --> 00:35:21,160
It does not make sense anymore.
1004
00:35:21,160 --> 00:35:23,480
Standardizing on a single model in an architecture
1005
00:35:23,480 --> 00:35:25,560
built around a reason and runtime split
1006
00:35:25,560 --> 00:35:28,600
is like hiring one person to do every job in your company.
1007
00:35:28,600 --> 00:35:30,680
You wouldn't ask the same person to file paperwork
1008
00:35:30,680 --> 00:35:32,000
and negotiate a merger.
1009
00:35:32,000 --> 00:35:34,440
The question was never wrong because the answer was hard.
1010
00:35:34,440 --> 00:35:36,960
It was wrong because it assumed the wrong thing needed
1011
00:35:36,960 --> 00:35:37,760
choosing.
1012
00:35:37,760 --> 00:35:39,160
The actual question is different.
1013
00:35:39,160 --> 00:35:40,400
What is our rooting logic?
1014
00:35:40,400 --> 00:35:41,720
And who owns the decision layer?
1015
00:35:41,720 --> 00:35:43,560
That is the only question that matters now.
1016
00:35:43,560 --> 00:35:46,040
Not which model, but who decides which model?
1017
00:35:46,040 --> 00:35:47,200
And based on what?
1018
00:35:47,200 --> 00:35:48,960
Because as we saw, the routing layer
1019
00:35:48,960 --> 00:35:51,560
is the part nobody wants to build and everybody needs.
1020
00:35:51,560 --> 00:35:54,000
Someone in the organization has to own that decision layer.
1021
00:35:54,000 --> 00:35:56,320
Someone has to be responsible for the classification logic
1022
00:35:56,320 --> 00:35:58,520
that sends a routine request to 5/4
1023
00:35:58,520 --> 00:36:02,320
and a genuinely hard one up to my one right now in most organizations.
1024
00:36:02,320 --> 00:36:03,320
Nobody owns that.
1025
00:36:03,320 --> 00:36:06,080
It is nobody's job, which means it is not getting built,
1026
00:36:06,080 --> 00:36:09,640
which means every request still defaults to whatever single model
1027
00:36:09,640 --> 00:36:11,080
happens to be plugged in.
1028
00:36:11,080 --> 00:36:14,520
That ownership gap points to a skill most IT teams do not have yet.
1029
00:36:14,520 --> 00:36:18,080
Workload classification, not model tuning, not prompt engineering,
1030
00:36:18,080 --> 00:36:19,560
workload classification.
1031
00:36:19,560 --> 00:36:21,880
This is the ability to look at an incoming request
1032
00:36:21,880 --> 00:36:24,560
and judge whether it is simple or reasoning heavy.
1033
00:36:24,560 --> 00:36:27,720
Before it ever reaches a model, that is a genuinely new skill.
1034
00:36:27,720 --> 00:36:29,480
Most architects were never trained to do this,
1035
00:36:29,480 --> 00:36:31,680
because until recently, there was no reason to do it.
1036
00:36:31,680 --> 00:36:34,040
Every request went to the same place regardless.
1037
00:36:34,040 --> 00:36:35,960
Now, that distinction is the whole game.
1038
00:36:35,960 --> 00:36:37,400
And most teams are starting from zero.
1039
00:36:37,400 --> 00:36:40,160
There is a governance layer sitting on top of all this too.
1040
00:36:40,160 --> 00:36:42,200
On-prem 5/4 deployments become the answer
1041
00:36:42,200 --> 00:36:44,240
when data sovereignty is non-negotiable,
1042
00:36:44,240 --> 00:36:46,240
when a request can never leave the building,
1043
00:36:46,240 --> 00:36:47,720
or the endpoint, or the country.
1044
00:36:47,720 --> 00:36:48,920
You go local.
1045
00:36:48,920 --> 00:36:50,800
Cloud-based MI1 becomes the answer
1046
00:36:50,800 --> 00:36:52,200
when you need centralized reasoning,
1047
00:36:52,200 --> 00:36:54,760
when you need shared context across a whole organization,
1048
00:36:54,760 --> 00:36:56,840
or the kind of planning work that benefits from sitting
1049
00:36:56,840 --> 00:36:57,680
in one place.
1050
00:36:57,680 --> 00:36:58,680
Go to the cloud.
1051
00:36:58,680 --> 00:37:00,520
That is not a technology decision anymore.
1052
00:37:00,520 --> 00:37:02,280
That is a governance decision.
1053
00:37:02,280 --> 00:37:04,680
And it has to be made deliberately, not by default.
1054
00:37:04,680 --> 00:37:07,440
And here is the real friction point, stated plainly.
1055
00:37:07,440 --> 00:37:10,400
Most organizations do not have this rooting infrastructure yet.
1056
00:37:10,400 --> 00:37:11,240
It does not exist.
1057
00:37:11,240 --> 00:37:12,040
The models exist.
1058
00:37:12,040 --> 00:37:13,040
5/4 is sitting there.
1059
00:37:13,040 --> 00:37:13,960
MIT licensed.
1060
00:37:13,960 --> 00:37:14,880
Ready to deploy.
1061
00:37:14,880 --> 00:37:16,720
MI1 is available through Foundry.
1062
00:37:16,720 --> 00:37:19,280
But the layer that decides which request goes where?
1063
00:37:19,280 --> 00:37:22,560
The classification logic, the ownership, the governance rules.
1064
00:37:22,560 --> 00:37:24,160
That is the actual work still ahead.
1065
00:37:24,160 --> 00:37:25,520
Nobody sells you that off the shelf.
1066
00:37:25,520 --> 00:37:27,280
You have to build it, which is a very different
1067
00:37:27,280 --> 00:37:29,440
conversation depending on who is sitting in the room
1068
00:37:29,440 --> 00:37:30,760
for architects and admins.
1069
00:37:30,760 --> 00:37:32,320
This is an infrastructure problem.
1070
00:37:32,320 --> 00:37:34,000
For the people signing the budget,
1071
00:37:34,000 --> 00:37:36,480
it looks like something else entirely.
1072
00:37:36,480 --> 00:37:39,000
What this means for decision makers and budget owners.
1073
00:37:39,000 --> 00:37:40,560
Here is how that conversation changes
1074
00:37:40,560 --> 00:37:42,960
once it reaches whoever signs the check.
1075
00:37:42,960 --> 00:37:46,400
Most executives still frame this as, how much does AI cost?
1076
00:37:46,400 --> 00:37:47,280
The wrong question.
1077
00:37:47,280 --> 00:37:49,760
The real question is, how much mis-rooted AI costs?
1078
00:37:49,760 --> 00:37:51,720
Because that is the number nobody is tracking.
1079
00:37:51,720 --> 00:37:53,640
It is not the invoice from the model provider.
1080
00:37:53,640 --> 00:37:56,200
It is the invisible cost sitting underneath it.
1081
00:37:56,200 --> 00:37:59,000
Every request that got sent to an expensive model
1082
00:37:59,000 --> 00:38:01,160
when a cheap one would have handled it just as well.
1083
00:38:01,160 --> 00:38:01,960
Is a loss.
1084
00:38:01,960 --> 00:38:04,000
That is not a line item on any budget report.
1085
00:38:04,000 --> 00:38:05,040
It is a leak.
1086
00:38:05,040 --> 00:38:07,280
And leaks do not show up until someone finally goes looking
1087
00:38:07,280 --> 00:38:07,760
for them.
1088
00:38:07,760 --> 00:38:09,040
Look at what correct rooting actually
1089
00:38:09,040 --> 00:38:10,400
produces in practice.
1090
00:38:10,400 --> 00:38:12,000
In one case, fuel cost reduction
1091
00:38:12,000 --> 00:38:14,320
was tied directly to model routing decisions.
1092
00:38:14,320 --> 00:38:16,840
Those savings showed up because the right model was matched
1093
00:38:16,840 --> 00:38:18,880
to the right task at the right layer.
1094
00:38:18,880 --> 00:38:20,880
We see the same pattern on the support side.
1095
00:38:20,880 --> 00:38:23,800
Latency was cut from 15 seconds down to under two.
1096
00:38:23,800 --> 00:38:26,440
That drop did not come from a faster frontier model.
1097
00:38:26,440 --> 00:38:28,440
It came from rooting routine requests,
1098
00:38:28,440 --> 00:38:31,280
somewhere that never needed frontier depth in the first place.
1099
00:38:31,280 --> 00:38:33,360
That is the whole argument made concrete.
1100
00:38:33,360 --> 00:38:35,000
The game was not capability.
1101
00:38:35,000 --> 00:38:36,400
It was placement.
1102
00:38:36,400 --> 00:38:38,640
Now, here is the number that turns this
1103
00:38:38,640 --> 00:38:40,400
into a real financial decision instead
1104
00:38:40,400 --> 00:38:41,680
of a technical preference.
1105
00:38:41,680 --> 00:38:43,480
Self-hosted small models break even
1106
00:38:43,480 --> 00:38:46,800
against managed frontier APIs inside 18 months.
1107
00:38:46,800 --> 00:38:48,920
Once you are operating at volume, 18 months
1108
00:38:48,920 --> 00:38:51,320
that is not a speculative payback period stretched out
1109
00:38:51,320 --> 00:38:53,880
over a decade to make the math look better on a slide.
1110
00:38:53,880 --> 00:38:56,520
That is a timeline a CFO can actually plan around.
1111
00:38:56,520 --> 00:38:58,640
At real volume, the math does not stay close.
1112
00:38:58,640 --> 00:39:00,600
It tilts hard toward whichever company
1113
00:39:00,600 --> 00:39:02,440
built the rooting infrastructure early,
1114
00:39:02,440 --> 00:39:04,440
which is exactly why this stops being purely
1115
00:39:04,440 --> 00:39:07,200
an engineering decision and becomes a board level conversation.
1116
00:39:07,200 --> 00:39:10,200
Engineering teams can debate architecture patterns all day,
1117
00:39:10,200 --> 00:39:13,240
but an 18 month break even on infrastructure spend.
1118
00:39:13,240 --> 00:39:17,000
Tied to a measurable latency and cost outcome is a different story.
1119
00:39:17,000 --> 00:39:19,760
That is the kind of number that gets a line in a quarterly review.
1120
00:39:19,760 --> 00:39:22,680
Once the payback period is that short and that provable,
1121
00:39:22,680 --> 00:39:24,520
the decision is not, should we let the engineers
1122
00:39:24,520 --> 00:39:25,560
experiment with this?
1123
00:39:25,560 --> 00:39:28,840
A, the decision is, why haven't we already funded it?
1124
00:39:28,840 --> 00:39:30,840
And here are the stakes worth sitting with.
1125
00:39:30,840 --> 00:39:33,800
Because it does not stay contained to one budget cycle.
1126
00:39:33,800 --> 00:39:35,320
The companies that get routing right
1127
00:39:35,320 --> 00:39:37,920
are not just saving money on this quarter's cloud bill.
1128
00:39:37,920 --> 00:39:40,320
They are operating at a structurally lower cost basis
1129
00:39:40,320 --> 00:39:42,080
than competitors who never built the split.
1130
00:39:42,080 --> 00:39:43,520
That is not a temporary advantage
1131
00:39:43,520 --> 00:39:45,400
that erodes once everyone catches up.
1132
00:39:45,400 --> 00:39:46,320
It compounds.
1133
00:39:46,320 --> 00:39:48,680
Every request processed at the correct layer.
1134
00:39:48,680 --> 00:39:51,440
Instead of defaulting to the most expensive option available,
1135
00:39:51,440 --> 00:39:53,200
is margin the other company does not have.
1136
00:39:53,200 --> 00:39:55,120
It stays that way, quarter after quarter,
1137
00:39:55,120 --> 00:39:56,960
at whatever scale they are operating at.
1138
00:39:56,960 --> 00:39:59,040
So the case for budget owners is straightforward.
1139
00:39:59,040 --> 00:40:01,040
This is not a bet on which model wins.
1140
00:40:01,040 --> 00:40:02,880
It is a bet on whether your cost structure
1141
00:40:02,880 --> 00:40:04,400
looks like the company is still routing
1142
00:40:04,400 --> 00:40:06,240
everything through one expensive door,
1143
00:40:06,240 --> 00:40:08,920
or the ones who build the infrastructure to stop doing that.
1144
00:40:08,920 --> 00:40:11,200
But before this turns into an uncritical pitch,
1145
00:40:11,200 --> 00:40:13,200
there is a tension worth being honest about,
1146
00:40:13,200 --> 00:40:16,520
one that sits uncomfortably next to everything just described.
1147
00:40:16,520 --> 00:40:18,800
The trust gap Microsoft hasn't solved yet.
1148
00:40:18,800 --> 00:40:20,000
There is a tension here.
1149
00:40:20,000 --> 00:40:23,280
It is worth saying out loud, instead of skipping past it.
1150
00:40:23,280 --> 00:40:25,920
Microsoft pitches Copilot as a serious enterprise tool.
1151
00:40:25,920 --> 00:40:26,960
They wire it into Excel.
1152
00:40:26,960 --> 00:40:27,920
They put it in Teams.
1153
00:40:27,920 --> 00:40:29,680
They tell you it belongs in the exact workflows
1154
00:40:29,680 --> 00:40:30,880
we have been talking about.
1155
00:40:30,880 --> 00:40:32,480
But then you look at the fine print.
1156
00:40:32,480 --> 00:40:35,520
Copilot's own terms of use still describe the product
1157
00:40:35,520 --> 00:40:38,040
as being for entertainment purposes only.
1158
00:40:38,040 --> 00:40:39,400
That is the actual language.
1159
00:40:39,400 --> 00:40:40,760
It is not a marketing footnote.
1160
00:40:40,760 --> 00:40:42,360
It is a legal disclaimer telling you
1161
00:40:42,360 --> 00:40:44,400
not to treat the outputs as reliable.
1162
00:40:44,400 --> 00:40:45,720
And Microsoft is not alone.
1163
00:40:45,720 --> 00:40:48,480
Open AI and XAI have versions of the same language
1164
00:40:48,480 --> 00:40:49,400
in their own terms.
1165
00:40:49,400 --> 00:40:50,920
This is an industry-wide pattern.
1166
00:40:50,920 --> 00:40:52,520
Corsion is baked into the legal layer
1167
00:40:52,520 --> 00:40:54,360
of every major AI provider.
1168
00:40:54,360 --> 00:40:56,280
It does not matter how confidently they talk
1169
00:40:56,280 --> 00:40:58,040
about the technology in public.
1170
00:40:58,040 --> 00:40:59,320
So here's the problem.
1171
00:40:59,320 --> 00:41:01,880
These companies want AI to be critical infrastructure.
1172
00:41:01,880 --> 00:41:04,160
They wanted to be the engine behind your Excel agents
1173
00:41:04,160 --> 00:41:05,160
and your support systems.
1174
00:41:05,160 --> 00:41:06,400
They wanted to handle decisions
1175
00:41:06,400 --> 00:41:08,520
that affect real budgets and real customers.
1176
00:41:08,520 --> 00:41:10,760
But at the same time, the legal language says,
1177
00:41:10,760 --> 00:41:12,360
don't actually rely on this.
1178
00:41:12,360 --> 00:41:13,840
You have critical infrastructure
1179
00:41:13,840 --> 00:41:16,400
and legally-disclamed unreliability sitting
1180
00:41:16,400 --> 00:41:18,080
in the same product at the same time.
1181
00:41:18,080 --> 00:41:19,560
That gap matters if you are planning
1182
00:41:19,560 --> 00:41:20,880
to actually use this in a business.
1183
00:41:20,880 --> 00:41:23,120
It is not a reason to avoid the architecture we have been
1184
00:41:23,120 --> 00:41:25,880
building, but it is a reason not to build dependencies
1185
00:41:25,880 --> 00:41:27,280
that you cannot reverse.
1186
00:41:27,280 --> 00:41:30,000
Nobody is willing to formally stand behind these outputs yet.
1187
00:41:30,000 --> 00:41:32,800
If a workflow can handle an occasional wrong answer,
1188
00:41:32,800 --> 00:41:33,760
that is one thing.
1189
00:41:33,760 --> 00:41:34,440
You catch it.
1190
00:41:34,440 --> 00:41:35,280
You correct it.
1191
00:41:35,280 --> 00:41:37,720
You route it to a human when confidence is low.
1192
00:41:37,720 --> 00:41:40,640
But if your workflow assumes the model is simply correct,
1193
00:41:40,640 --> 00:41:41,920
you are taking a massive risk.
1194
00:41:41,920 --> 00:41:43,240
That is the risk the terms of use
1195
00:41:43,240 --> 00:41:45,400
are warning you about whether you read them or not.
1196
00:41:45,400 --> 00:41:47,440
Now, we should add some context here.
1197
00:41:47,440 --> 00:41:50,680
Microsoft has acknowledged that this wording is outdated.
1198
00:41:50,680 --> 00:41:51,600
That matters.
1199
00:41:51,600 --> 00:41:53,600
A company saying the language does not reflect
1200
00:41:53,600 --> 00:41:56,200
the product is different from a company defending it
1201
00:41:56,200 --> 00:41:57,120
as accurate.
1202
00:41:57,120 --> 00:41:59,000
It suggests this is just a lagging artifact.
1203
00:41:59,000 --> 00:42:01,960
It was written years before Copilot could do what it does today.
1204
00:42:01,960 --> 00:42:03,520
But acknowledging outdated language
1205
00:42:03,520 --> 00:42:05,720
and actually fixing the gap are two different things.
1206
00:42:05,720 --> 00:42:07,200
Only one of those has happened.
1207
00:42:07,200 --> 00:42:09,280
This does not change the strategy we have covered.
1208
00:42:09,280 --> 00:42:10,960
The reason and runtime split still works.
1209
00:42:10,960 --> 00:42:13,400
The routing logic and the cost math still hold up.
1210
00:42:13,400 --> 00:42:15,040
What it means is that governance has
1211
00:42:15,040 --> 00:42:16,400
to catch up to the technology.
1212
00:42:16,400 --> 00:42:19,160
The architecture is ahead of the paperwork right now.
1213
00:42:19,160 --> 00:42:20,840
And until that paperwork catches up,
1214
00:42:20,840 --> 00:42:23,160
the responsible move is to build verification
1215
00:42:23,160 --> 00:42:24,040
into your workflow.
1216
00:42:24,040 --> 00:42:26,120
Do not assume the legal language will update itself
1217
00:42:26,120 --> 00:42:28,840
before something important breaks.
1218
00:42:28,840 --> 00:42:31,360
Multimodality as the next layer of the runtime.
1219
00:42:31,360 --> 00:42:33,400
Now zoom out from governance for a second.
1220
00:42:33,400 --> 00:42:35,760
There is another layer to this runtime story.
1221
00:42:35,760 --> 00:42:37,240
It lives inside the word multimodal,
1222
00:42:37,240 --> 00:42:39,280
whether when you look at 5 or 4 multimodal,
1223
00:42:39,280 --> 00:42:42,640
it does not run vision, audio and text as three separate products.
1224
00:42:42,640 --> 00:42:45,240
It is not a text model with a camera and a microphone
1225
00:42:45,240 --> 00:42:46,760
bolted on as afterthoughts.
1226
00:42:46,760 --> 00:42:49,640
Those setups have separate pipelines and separate failure points.
1227
00:42:49,640 --> 00:42:50,760
This is one backbone.
1228
00:42:50,760 --> 00:42:53,360
One core handles all three kinds of input.
1229
00:42:53,360 --> 00:42:55,360
Here is what that looks like in the real world.
1230
00:42:55,360 --> 00:42:59,080
You get OCR accuracy that holds up on messy documents.
1231
00:42:59,080 --> 00:43:00,880
You can hand the model an image and ask
1232
00:43:00,880 --> 00:43:02,920
specific questions about what is in it.
1233
00:43:02,920 --> 00:43:05,120
You get answers grounded in what is actually there.
1234
00:43:05,120 --> 00:43:08,160
You get audio transcription and translation in one single pass.
1235
00:43:08,160 --> 00:43:11,200
There is no separate translation step added on afterward.
1236
00:43:11,200 --> 00:43:14,640
It is one model, one forward pass, three different senses
1237
00:43:14,640 --> 00:43:16,280
feeding into the same understanding.
1238
00:43:16,280 --> 00:43:17,800
Why does this matter for developers?
1239
00:43:17,800 --> 00:43:19,640
Because the alternative is a nightmare.
1240
00:43:19,640 --> 00:43:22,160
Separate pipelines mean separate failure points
1241
00:43:22,160 --> 00:43:23,800
and separate latency budgets.
1242
00:43:23,800 --> 00:43:26,040
You have to keep different versions synchronized.
1243
00:43:26,040 --> 00:43:28,040
Things break quietly and you do not notice
1244
00:43:28,040 --> 00:43:29,280
until a user complains.
1245
00:43:29,280 --> 00:43:31,160
A unified backbone collapses all of that.
1246
00:43:31,160 --> 00:43:33,480
It gives you one thing to deploy, one thing to monitor,
1247
00:43:33,480 --> 00:43:34,680
one thing to update.
1248
00:43:34,680 --> 00:43:36,720
That is the difference between a runtime you can actually
1249
00:43:36,720 --> 00:43:40,800
maintain and one that becomes unmanageable as you add more inputs.
1250
00:43:40,800 --> 00:43:42,600
This pattern shows up on the MAI side too.
1251
00:43:42,600 --> 00:43:45,920
They use different names like MAI voice or MI image.
1252
00:43:45,920 --> 00:43:48,240
But they follow the exact same logic we have been describing.
1253
00:43:48,240 --> 00:43:49,840
Reasoning heavy work sits in one place.
1254
00:43:49,840 --> 00:43:52,240
Fast specialized execution sits somewhere else.
1255
00:43:52,240 --> 00:43:54,440
It stays closer to where the input is happening.
1256
00:43:54,440 --> 00:43:57,040
It is easy to think multi-modal is a separate category.
1257
00:43:57,040 --> 00:43:57,680
It isn't.
1258
00:43:57,680 --> 00:44:00,120
Even inside these systems, the same split holds.
1259
00:44:00,120 --> 00:44:02,840
Heavy reasoning needs deep context and multi-step planning
1260
00:44:02,840 --> 00:44:04,480
that stays with the reasoning layer.
1261
00:44:04,480 --> 00:44:07,160
Fast execution needs to transcribe audio or read text
1262
00:44:07,160 --> 00:44:08,560
off a scan in real time.
1263
00:44:08,560 --> 00:44:10,200
That stays with the runtime layer.
1264
00:44:10,200 --> 00:44:12,720
Adding vision and audio did not break the reason and runtime
1265
00:44:12,720 --> 00:44:13,400
pattern.
1266
00:44:13,400 --> 00:44:16,040
It just gave the pattern more kinds of input to sort through.
1267
00:44:16,040 --> 00:44:18,200
The architecture does not get more complicated
1268
00:44:18,200 --> 00:44:19,720
as you add modalities.
1269
00:44:19,720 --> 00:44:23,000
It just gets applied more broadly, which leads to the next question.
1270
00:44:23,000 --> 00:44:25,680
If this split holds true across text, vision, and audio
1271
00:44:25,680 --> 00:44:27,720
today, where does Microsoft take it next?
1272
00:44:27,720 --> 00:44:30,360
And what does that tell you about what is coming?
1273
00:44:30,360 --> 00:44:32,800
The trajectory where this architecture goes next.
1274
00:44:32,800 --> 00:44:35,240
Microsoft is signaling where this is headed.
1275
00:44:35,240 --> 00:44:38,040
And it is much bigger than the seven models we saw at build.
1276
00:44:38,040 --> 00:44:40,520
The plan does not stop at task-specific models.
1277
00:44:40,520 --> 00:44:43,320
Mustafa Saliman has been very direct about the goal here.
1278
00:44:43,320 --> 00:44:45,520
They are building full-frontier LLMs
1279
00:44:45,520 --> 00:44:47,640
to compete with systems like GPT-4.
1280
00:44:47,640 --> 00:44:49,320
They aren't just building narrow tools
1281
00:44:49,320 --> 00:44:50,960
for coding or transcription.
1282
00:44:50,960 --> 00:44:52,280
Everything we have looked at so far,
1283
00:44:52,280 --> 00:44:54,800
the MAI thinking models, the Excel tuning,
1284
00:44:54,800 --> 00:44:57,680
the McKinsey results, that is just the current chapter.
1285
00:44:57,680 --> 00:44:58,600
It is not the whole book.
1286
00:44:58,600 --> 00:45:00,040
Microsoft is moving toward a world
1287
00:45:00,040 --> 00:45:01,440
where general frontier capabilities
1288
00:45:01,440 --> 00:45:03,520
sits right alongside the specialized layer.
1289
00:45:03,520 --> 00:45:05,080
It is not an either/or strategy.
1290
00:45:05,080 --> 00:45:07,680
We see the same expansion happening on the runtime side.
1291
00:45:07,680 --> 00:45:10,800
5.4 is not staying a small line-up of two or three models.
1292
00:45:10,800 --> 00:45:13,440
It is already growing toward 10 different versions.
1293
00:45:13,440 --> 00:45:17,080
These will span from 3.8 billion parameters up to 15 billion.
1294
00:45:17,080 --> 00:45:20,080
And every single one of them is still MIT licensed.
1295
00:45:20,080 --> 00:45:23,240
That licensing detail is just as important now as it was earlier.
1296
00:45:23,240 --> 00:45:26,560
It means this isn't just Microsoft growing its own internal layer.
1297
00:45:26,560 --> 00:45:28,680
It is a runtime layer that any developer can use.
1298
00:45:28,680 --> 00:45:30,280
You can pick the size that fits your hardware
1299
00:45:30,280 --> 00:45:32,560
without needing a legal team to sign off first.
1300
00:45:32,560 --> 00:45:35,160
But here is a signal you should pay close attention to.
1301
00:45:35,160 --> 00:45:36,680
Look at the Mayo Clinic partnership.
1302
00:45:36,680 --> 00:45:38,680
Microsoft is not just licensing a generic model
1303
00:45:38,680 --> 00:45:40,360
to a hospital and walking away.
1304
00:45:40,360 --> 00:45:42,840
They are jointly building a frontier model specifically
1305
00:45:42,840 --> 00:45:43,800
for healthcare.
1306
00:45:43,800 --> 00:45:45,520
They are combining frontier reasoning
1307
00:45:45,520 --> 00:45:47,120
with Mayo's own clinical data
1308
00:45:47,120 --> 00:45:49,120
and deploying it at a massive hospital scale.
1309
00:45:49,120 --> 00:45:49,920
This is not a demo.
1310
00:45:49,920 --> 00:45:51,520
This is the reason and runtime pattern
1311
00:45:51,520 --> 00:45:54,560
applied to one of the highest stakes domains in the world.
1312
00:45:54,560 --> 00:45:57,360
You have a major institution putting its name on the outcome.
1313
00:45:57,360 --> 00:45:58,840
So what does that tell us about the future?
1314
00:45:58,840 --> 00:46:02,480
It points toward vertical specific pairs of reason and runtime.
1315
00:46:02,480 --> 00:46:04,800
We are going to see the show up industry by industry.
1316
00:46:04,800 --> 00:46:06,240
Healthcare gets its own tuned pairing
1317
00:46:06,240 --> 00:46:07,320
built on clinical data.
1318
00:46:07,320 --> 00:46:08,520
Legal gets its own.
1319
00:46:08,520 --> 00:46:11,160
Manufacturing, logistics and financial services
1320
00:46:11,160 --> 00:46:13,040
will likely follow the same path.
1321
00:46:13,040 --> 00:46:15,760
Each industry ends up with a frontier reasoning layer
1322
00:46:15,760 --> 00:46:17,800
shaped around its specific problems.
1323
00:46:17,800 --> 00:46:20,480
And that layer is paired with a fast local runtime
1324
00:46:20,480 --> 00:46:22,080
tuned for daily execution.
1325
00:46:22,080 --> 00:46:23,720
If you push that idea forward,
1326
00:46:23,720 --> 00:46:25,960
you can see what it looks like inside a single company.
1327
00:46:25,960 --> 00:46:27,560
You won't have one shared model bolted
1328
00:46:27,560 --> 00:46:30,200
onto every department just because that is what the license allowed.
1329
00:46:30,200 --> 00:46:32,800
Instead, every department runs its own tuned runtime.
1330
00:46:32,800 --> 00:46:35,040
It is shaped around their specific workflows,
1331
00:46:35,040 --> 00:46:37,280
just like the Excel and McKinsey examples.
1332
00:46:37,280 --> 00:46:39,040
Each of those runtimes connects back
1333
00:46:39,040 --> 00:46:41,480
to a shared reasoning layer underneath.
1334
00:46:41,480 --> 00:46:44,480
That foundation provides the deep planning capability
1335
00:46:44,480 --> 00:46:48,000
that the individual runtimes don't need to carry themselves.
1336
00:46:48,000 --> 00:46:49,920
Finance gets a runtime for finance.
1337
00:46:49,920 --> 00:46:51,440
Support gets one for support.
1338
00:46:51,440 --> 00:46:52,480
Underneath all of them,
1339
00:46:52,480 --> 00:46:54,480
the whole company draws from one reasoning foundation
1340
00:46:54,480 --> 00:46:55,960
instead of rebuilding it every time.
1341
00:46:55,960 --> 00:46:57,320
That is the direction of travel.
1342
00:46:57,320 --> 00:46:59,600
It is not one giant model absorbing everything.
1343
00:46:59,600 --> 00:47:02,320
It is also not a scattered pile of small disconnected tools.
1344
00:47:02,320 --> 00:47:03,120
It is a structure.
1345
00:47:03,120 --> 00:47:04,920
You have many tuned runtimes on top
1346
00:47:04,920 --> 00:47:06,840
and one shared reasoning layer underneath.
1347
00:47:06,840 --> 00:47:09,480
It happens industry by industry and department by department.
1348
00:47:09,480 --> 00:47:10,680
Now we need to step back.
1349
00:47:10,680 --> 00:47:13,120
We need to look at how the models, the routing,
1350
00:47:13,120 --> 00:47:15,920
the governance and the trust gap all fit together.
1351
00:47:15,920 --> 00:47:17,160
They aren't just parts.
1352
00:47:17,160 --> 00:47:19,120
They are one single architecture.
1353
00:47:19,120 --> 00:47:20,680
Tying the architecture together.
1354
00:47:20,680 --> 00:47:22,920
Here is the whole stack laid out in one place.
1355
00:47:22,920 --> 00:47:24,960
You have models that split by design.
1356
00:47:24,960 --> 00:47:28,560
You have a runtime layer built to execute as close to the work as possible.
1357
00:47:28,560 --> 00:47:30,760
There is an orchestration layer that decides
1358
00:47:30,760 --> 00:47:32,640
where each request needs to go.
1359
00:47:32,640 --> 00:47:36,000
You have a governance layer trying to solve sovereignty questions and trust gaps.
1360
00:47:36,000 --> 00:47:39,480
And at the very bottom, you have silicon built in lockstep with the models.
1361
00:47:39,480 --> 00:47:42,760
The models aren't trying to fit into whatever hardware was available.
1362
00:47:42,760 --> 00:47:44,000
The hardware was built for them.
1363
00:47:44,000 --> 00:47:46,840
These are five layers and none of them work in isolation.
1364
00:47:46,840 --> 00:47:48,600
If you skipped to this part of the video,
1365
00:47:48,600 --> 00:47:50,560
here is the one sentence you need to hear.
1366
00:47:50,560 --> 00:47:53,000
Five four executes and MI1 decides.
1367
00:47:53,000 --> 00:47:54,120
That is the entire reframe.
1368
00:47:54,120 --> 00:47:55,280
They are not competitors.
1369
00:47:55,280 --> 00:47:57,120
This is not a contest of big versus small.
1370
00:47:57,120 --> 00:47:59,840
It is not a leaderboard with a winner and a loser.
1371
00:47:59,840 --> 00:48:02,760
These are two different jobs based on two different philosophies.
1372
00:48:02,760 --> 00:48:04,280
They are working together on purpose,
1373
00:48:04,280 --> 00:48:06,320
but we should be very clear about something before we finish.
1374
00:48:06,320 --> 00:48:07,680
This is not unique to Microsoft.
1375
00:48:07,680 --> 00:48:10,920
Every major AI provider is moving towards some version of this same split.
1376
00:48:10,920 --> 00:48:13,520
You have a heavy reasoning layer for the hard problems
1377
00:48:13,520 --> 00:48:16,040
and a fast specialized layer for everything else.
1378
00:48:16,040 --> 00:48:18,880
The economics of AI force everyone toward this answer eventually.
1379
00:48:18,880 --> 00:48:21,680
Microsoft just happened to turn it into a product first.
1380
00:48:21,680 --> 00:48:24,400
They did it at their own scale and on their own chips.
1381
00:48:24,400 --> 00:48:26,400
The pattern belongs to the whole industry,
1382
00:48:26,400 --> 00:48:27,800
not just one company road map.
1383
00:48:27,800 --> 00:48:30,880
Because of that, the competitive question has changed.
1384
00:48:30,880 --> 00:48:33,680
It is no longer about who has access to the models.
1385
00:48:33,680 --> 00:48:35,320
The organizations that win the next few years
1386
00:48:35,320 --> 00:48:37,040
won't be the ones with the best license.
1387
00:48:37,040 --> 00:48:39,000
Model access is basically a commodity now.
1388
00:48:39,000 --> 00:48:40,920
If you have a budget, you have a model.
1389
00:48:40,920 --> 00:48:42,840
What is not a commodity is your rooting logic.
1390
00:48:42,840 --> 00:48:44,680
The real value is in the classification work
1391
00:48:44,680 --> 00:48:45,960
and the governance decisions.
1392
00:48:45,960 --> 00:48:47,440
It is the willingness to build the layer
1393
00:48:47,440 --> 00:48:49,400
that actually decides where a request goes.
1394
00:48:49,400 --> 00:48:50,840
That part does not come prebuilt.
1395
00:48:50,840 --> 00:48:53,480
That is what separates companies that are structurally faster
1396
00:48:53,480 --> 00:48:55,360
from companies that still send every request
1397
00:48:55,360 --> 00:48:56,880
through one expensive door.
1398
00:48:56,880 --> 00:48:58,600
Now I want to add a quick caveat here.
1399
00:48:58,600 --> 00:49:01,200
This deserves honesty rather than just acting like everything
1400
00:49:01,200 --> 00:49:02,080
is a settled fact.
1401
00:49:02,080 --> 00:49:03,800
Some of what we talked about is reported.
1402
00:49:03,800 --> 00:49:05,280
It comes from post-built coverage,
1403
00:49:05,280 --> 00:49:08,040
analyst reports and public statements from Microsoft.
1404
00:49:08,040 --> 00:49:11,160
It is not all confirmed line by line in technical docs yet.
1405
00:49:11,160 --> 00:49:13,040
The 5.4 architecture is well documented,
1406
00:49:13,040 --> 00:49:14,800
but some of the newer MEI details
1407
00:49:14,800 --> 00:49:16,440
and the exact benchmark numbers
1408
00:49:16,440 --> 00:49:18,760
are still coming out as the system matures.
1409
00:49:18,760 --> 00:49:20,000
It is worth watching,
1410
00:49:20,000 --> 00:49:21,600
but don't treat it as final truth
1411
00:49:21,600 --> 00:49:24,000
just because the numbers sound precise.
1412
00:49:24,000 --> 00:49:25,360
That distinction is important,
1413
00:49:25,360 --> 00:49:27,040
but it does not change the core point.
1414
00:49:27,040 --> 00:49:29,560
The architecture pattern is real, it stays real,
1415
00:49:29,560 --> 00:49:31,800
regardless of what the final numbers look like.
1416
00:49:31,800 --> 00:49:34,240
And that leaves one big question on the table.
1417
00:49:34,240 --> 00:49:36,080
If model access is a solved problem
1418
00:49:36,080 --> 00:49:39,000
and rooting is the real battleground, what does that change?
1419
00:49:39,000 --> 00:49:40,800
It changes how organizations have to think.
1420
00:49:40,800 --> 00:49:42,240
It isn't just about AI anymore,
1421
00:49:42,240 --> 00:49:45,480
it is about how you build anything from here on out.
1422
00:49:45,480 --> 00:49:48,440
The real shift, nobody's naming, zoom all the way out.
1423
00:49:48,440 --> 00:49:51,720
There is a bigger shift buried underneath everything we just covered.
1424
00:49:51,720 --> 00:49:53,240
And it is not really about Microsoft,
1425
00:49:53,240 --> 00:49:56,440
this is the end of the pick one AI vendor era.
1426
00:49:56,440 --> 00:49:58,160
That instinct made sense a few years ago
1427
00:49:58,160 --> 00:49:59,840
when you were just picking a horse.
1428
00:49:59,840 --> 00:50:02,520
One contract, one model, one API key and you were done.
1429
00:50:02,520 --> 00:50:03,640
That world is over.
1430
00:50:03,640 --> 00:50:06,040
It is over because the winning architecture
1431
00:50:06,040 --> 00:50:07,680
was never going to be a single model.
1432
00:50:07,680 --> 00:50:09,080
It was always going to be a system,
1433
00:50:09,080 --> 00:50:10,440
layers doing different jobs.
1434
00:50:10,440 --> 00:50:13,360
And the system does not come from a single vendor relationship.
1435
00:50:13,360 --> 00:50:16,920
It comes from design, which points to the actual shift worth naming.
1436
00:50:16,920 --> 00:50:19,200
Architectural literacy is becoming a business skill,
1437
00:50:19,200 --> 00:50:20,600
not just a technical one.
1438
00:50:20,600 --> 00:50:25,560
For years, understanding AI meant knowing which model scored higher on a benchmark.
1439
00:50:25,560 --> 00:50:28,160
That conversation lived entirely inside engineering,
1440
00:50:28,160 --> 00:50:30,600
but that is not where the conversation lives anymore.
1441
00:50:30,600 --> 00:50:34,400
Knowing why a request should go to a fast local model instead of a frontier one,
1442
00:50:34,400 --> 00:50:36,600
knowing what rooting logic costs to skip,
1443
00:50:36,600 --> 00:50:39,760
knowing where governance sits relative to where reasoning happens.
1444
00:50:39,760 --> 00:50:41,120
That is business judgment now.
1445
00:50:41,120 --> 00:50:43,080
It shows up in budget meetings and board decks.
1446
00:50:43,080 --> 00:50:44,920
The people who need to understand this split
1447
00:50:44,920 --> 00:50:47,000
are no longer just the ones writing the code.
1448
00:50:47,000 --> 00:50:49,120
Go back to how this whole episode opened.
1449
00:50:49,120 --> 00:50:51,840
The question everyone kept asking was which model wins?
1450
00:50:51,840 --> 00:50:54,560
M-A-I-1 or F-I-4?
1451
00:50:54,560 --> 00:50:55,880
Bigger or smaller?
1452
00:50:55,880 --> 00:50:57,320
Faster or smarter?
1453
00:50:57,320 --> 00:50:59,800
That question was never the real one.
1454
00:50:59,800 --> 00:51:02,480
The real question, the one worth asking from the very first minute,
1455
00:51:02,480 --> 00:51:04,760
was how these systems divide labor.
1456
00:51:04,760 --> 00:51:05,720
It is not a contest.
1457
00:51:05,720 --> 00:51:07,000
It is a design decision.
1458
00:51:07,000 --> 00:51:08,920
Every section since has just been unpacking
1459
00:51:08,920 --> 00:51:10,960
what that division looks like in practice.
1460
00:51:10,960 --> 00:51:14,040
What it costs to get wrong and what it is worth to get right.
1461
00:51:14,040 --> 00:51:16,960
So here is the belief breaking statement worth sitting with.
1462
00:51:16,960 --> 00:51:19,440
The companies still asking which model is smarter
1463
00:51:19,440 --> 00:51:22,360
are already behind the ones asking how to root intelligently.
1464
00:51:22,360 --> 00:51:23,720
That gap is not going to close.
1465
00:51:23,720 --> 00:51:26,680
It is going to widen quietly, one quarter at a time.
1466
00:51:26,680 --> 00:51:29,520
Eventually the company is still fighting over leaderboard positions
1467
00:51:29,520 --> 00:51:32,680
will look up and realize the race was never about the leaderboard.
1468
00:51:32,680 --> 00:51:34,280
So here is what to actually do with this.
1469
00:51:34,280 --> 00:51:37,200
Take your current A-I workloads and sort them into two piles.
1470
00:51:37,200 --> 00:51:40,440
Needs reasoning, needs fast execution.
1471
00:51:40,440 --> 00:51:42,480
Most of what is running through a frontier model
1472
00:51:42,480 --> 00:51:44,280
right now belongs in that second pile.
1473
00:51:44,280 --> 00:51:47,280
Pick one high volume low complexity task this month.
1474
00:51:47,280 --> 00:51:49,960
Test it against a small model before you default
1475
00:51:49,960 --> 00:51:51,800
to the expensive option out of habit.
1476
00:51:51,800 --> 00:51:54,040
If this changed how you think about your own A-I stack,
1477
00:51:54,040 --> 00:51:56,120
follow me, M��ko Peters on LinkedIn.
1478
00:51:56,120 --> 00:51:58,320
And if you want more of this, leave a review.
1479
00:51:58,320 --> 00:51:59,600
It helps more people find it.
1480
00:51:59,600 --> 00:52:01,320
The architecture is already being built.
1481
00:52:01,320 --> 00:52:04,520
The only open question left is who builds the routing layer first?