Your Copilot Has No Memory — Why RAG Fails and LLM Wiki Fixes It
Microsoft Copilot can feel remarkably intelligent in a demo. It can find a project document, summarize a policy, extract information from SharePoint, and connect pieces of enterprise content into a convincing answer. But there is a fundamental architectural limitation hiding behind that experience: retrieval is not memory. Copilot can search your organization, but that does not mean it has built a persistent understanding of your organization. In this deep dive, we explore the difference between retrieving information and actually compiling organizational knowledge. We examine how Retrieval-Augmented Generation works, why traditional RAG architectures become unreliable when questions require continuity and context, and why an emerging LLM Wiki approach could provide a fundamentally different knowledge layer for enterprise AI.
COPILOT DOESN'T REMEMBER — IT RETRIEVES
The experience of using Copilot creates an important illusion. When it successfully connects information from Microsoft 365, it can appear as though the system has learned something about your organization. Ask a similar question later, however, and the system may produce a different result because it performs another retrieval operation rather than simply recalling the understanding it established previously. That distinction becomes critical when organizations move beyond basic summarization and start expecting AI to support real decisions. If essentially identical questions can produce inconsistent answers depending on which information was retrieved, users quickly become reluctant to rely on AI for business-critical work. The result can become an adoption problem rather than merely a technical problem.
WHAT RETRIEVAL-AUGMENTED GENERATION ACTUALLY DOES RAG
stands for Retrieval-Augmented Generation. At a simplified level, enterprise documents are divided into chunks, those chunks are represented through embeddings, a user's question is transformed into a comparable representation, and the system searches for semantically relevant chunks. The highest-ranking pieces of information are then supplied to the language model as context for generating its response. This architecture is extremely useful. It allows an LLM to answer questions using information that was never part of its original training data and provides a practical way to ground AI responses in enterprise content. But it also creates an important architectural constraint: the model receives fragments selected for the current query rather than maintaining a complete persistent representation of the organization's knowledge.
THE CHUNKING PROBLEM
Enterprise knowledge rarely exists as isolated paragraphs. A project plan might contain dependencies distributed across dozens of pages. A policy might reference another policy. A technical architecture could depend on decisions documented months earlier in meeting notes, Teams conversations, SharePoint pages, and design documents. Traditional RAG breaks those sources into smaller units and determines which fragments appear relevant to the current question. The model therefore sees selected pieces rather than necessarily understanding the complete document and all of its relationships. For straightforward information retrieval, that can work extremely well. For questions requiring relationships, historical context, dependencies, accumulated decisions, or reasoning across many sources, the limitations become much more visible.
THE STATELESS RAG TRAP
A conventional retrieval workflow has no inherent concept of something being "already figured out." A question is received, information is retrieved, an answer is generated, and the process effectively starts again for the next retrieval operation. The script describes this as the point where RAG's statelessness becomes a problem. That means knowledge discovered during one interaction does not automatically become durable organizational knowledge available to every future interaction. For enterprise AI, this is a major distinction. Organizations do not simply need better search. They increasingly need systems capable of maintaining a structured understanding of projects, processes, policies, technologies, people, decisions, dependencies, and the relationships connecting them.
SEARCHING FOR KNOWLEDGE VS. COMPILING KNOWLEDGE
This leads to the central architectural idea of the episode: instead of repeatedly reconstructing organizational knowledge at query time, what if AI compiled that knowledge beforehand? An LLM Wiki represents that shift in thinking. Rather than treating every enterprise document as another collection of fragments waiting for retrieval, AI can synthesize information into structured knowledge artifacts that represent what the organization currently understands. The important change is not simply another user interface. It is moving intelligence from query-time reconstruction toward persistent knowledge synthesis.
WHY AN LLM WIKI CHANGES THE MODEL
Imagine thousands of documents describing the same product, customer, project, policy, or architecture. Traditional RAG waits for a question and then tries to locate the fragments most likely to answer it. An LLM Wiki approach instead attempts to continuously transform those fragmented sources into coherent knowledge pages. Relationships, decisions, definitions, dependencies, historical context, and supporting sources can become part of a maintained knowledge representation. Copilot or another AI agent can then retrieve from a layer containing synthesized organizational understanding instead of repeatedly attempting to reconstruct that understanding from raw documents.
FROM DOCUMENT REPOSITORY TO KNOWLEDGE LAYER
This changes the role of systems such as SharePoint. SharePoint can continue serving as the authoritative repository for documents, pages, policies, presentations, meeting artifacts, and collaboration content. But raw enterprise content does not automatically constitute usable organizational knowledge. An AI-generated knowledge layer can sit above those source systems and transform scattered information into something closer to an organizational map: projects connected to decisions, policies connected to processes, systems connected to owners, and concepts connected to their supporting evidence. The goal is not to eliminate source documents. It is to make the relationships hidden inside them explicit.
WHY BETTER PROMPTS DON'T SOLVE THE ARCHITECTURE
Prompt engineering can improve how an LLM interprets retrieved context, but it cannot guarantee that the correct context was retrieved in the first place. If the retrieval layer returns incomplete fragments, misses an important dependency, or surfaces an outdated document, even an excellent model is reasoning over an incomplete information set. This is why improving the model alone cannot solve every enterprise AI problem. The quality and structure of the knowledge supplied to the model remain fundamental.
THE GOVERNANCE PROBLEM GETS BIGGER
There is also a significant warning attached to this architecture. If an organization's SharePoint environment contains obsolete policies, duplicate documentation, abandoned processes, contradictory instructions, or documents without clear ownership, an LLM Wiki can synthesize that bad information just as efficiently as it synthesizes good information. The danger is that synthesized knowledge can look considerably cleaner and more authoritative than the underlying content deserves. AI therefore makes traditional information governance more important rather than less important. Content ownership, lifecycle management, versioning, retention, archival processes, authoritative sources, metadata, and clearly defined systems of record become foundational components of AI architecture.
AI READINESS STARTS WITH CONTENT QUALITY
Organizations frequently approach Copilot readiness as a licensing, security, deployment, or training project. Those elements matter, but the underlying knowledge environment matters just as much. If nobody knows which document represents the current process, the AI cannot magically resolve the organizational ambiguity. If three departments maintain contradictory versions of a policy, AI has inherited three versions of the truth. If obsolete documentation remains searchable indefinitely, it remains potential grounding material. Enterprise AI therefore exposes knowledge-management debt that organizations could previously ignore.
WHY TRUST DETERMINES COPILOT ADOPTION
The technical consequences quickly become business consequences. Users may tolerate occasional inconsistencies when AI is used for drafting an email or summarizing a meeting. They become much less tolerant when AI is expected to explain policy, support customer decisions, interpret project status, provide compliance information, or guide operational processes. Once users experience inconsistent answers to important questions, they frequently return to trusted human experts and established manual processes. The script identifies this as a major reason why Copilot adoption can flatten even after an apparently successful rollout. Trust therefore becomes an architectural requirement, not merely an adoption metric.
Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-modern-work-security-and-productivity-with-microsoft-365--6704921/support.
🚀 Want to be part of m365.fm?
Then stop just listening… and start showing up.
👉 Connect with me on LinkedIn and let’s make something happen:
- 🎙️ Be a podcast guest and share your story
- 🎧 Host your own episode (yes, seriously)
- 💡 Pitch topics the community actually wants to hear
- 🌍 Build your personal brand in the Microsoft 365 space
This isn’t just a podcast — it’s a platform for people who take action.
🔥 Most people wait. The best ones don’t.
👉 Connect with me on LinkedIn and send me a message:
"I want in"
Let’s build something awesome 👊
00:00:00,000 --> 00:00:02,360
Last week, your copilot answered a question perfectly.
2
00:00:02,360 --> 00:00:04,200
Someone asked about a project timeline,
3
00:00:04,200 --> 00:00:06,440
and it pulled the right details connected the right dots
4
00:00:06,440 --> 00:00:08,320
gave you an answer that felt almost human.
5
00:00:08,320 --> 00:00:10,160
Today, someone asks the same question,
6
00:00:10,160 --> 00:00:11,880
slightly different wording.
7
00:00:11,880 --> 00:00:14,960
And copilot fumbles the exact connection it made yesterday.
8
00:00:14,960 --> 00:00:17,600
Here's the assumption everyone walks around with.
9
00:00:17,600 --> 00:00:19,600
Copilot remembers your organization.
10
00:00:19,600 --> 00:00:20,120
It doesn't.
11
00:00:20,120 --> 00:00:20,920
It searches.
12
00:00:20,920 --> 00:00:22,920
Every single time, from zero, that's the tension
13
00:00:22,920 --> 00:00:23,920
nobody talks about.
14
00:00:23,920 --> 00:00:25,760
Retrieval and memory are not the same thing.
15
00:00:25,760 --> 00:00:28,800
One finds information, the other builds on what it already knows.
16
00:00:28,800 --> 00:00:31,160
Microsoft built copilot on the first one
17
00:00:31,160 --> 00:00:33,040
and let the marketing imply the second.
18
00:00:33,040 --> 00:00:35,240
There's a pattern circulating right now.
19
00:00:35,240 --> 00:00:37,160
Born out of AI research circles,
20
00:00:37,160 --> 00:00:39,200
that treats knowledge as something you compile,
21
00:00:39,200 --> 00:00:41,160
not something you search for over and over.
22
00:00:41,160 --> 00:00:42,560
We'll name it in a few minutes.
23
00:00:42,560 --> 00:00:44,360
By the end of this episode, you'll understand
24
00:00:44,360 --> 00:00:46,560
why your copilot deployment feels sharp in demos
25
00:00:46,560 --> 00:00:48,600
and shallow in production and what structural change
26
00:00:48,600 --> 00:00:49,880
actually fixes it.
27
00:00:49,880 --> 00:00:51,240
Quick thing before we get into it.
28
00:00:51,240 --> 00:00:53,240
If you're into this channel because you want to understand
29
00:00:53,240 --> 00:00:55,480
the systems behind Microsoft 365,
30
00:00:55,480 --> 00:00:59,480
not just the button clicking, hit subscribe on M365FM podcast.
31
00:00:59,480 --> 00:01:02,040
This episode is exactly the kind of thing we do here.
32
00:01:02,040 --> 00:01:03,720
And I want to set the expectation right now.
33
00:01:03,720 --> 00:01:05,560
This isn't a copilot feature tour.
34
00:01:05,560 --> 00:01:07,360
We're not walking through settings menus.
35
00:01:07,360 --> 00:01:08,840
This is a structural diagnosis,
36
00:01:08,840 --> 00:01:11,320
the kind you do on a system that keeps almost working
37
00:01:11,320 --> 00:01:12,800
but never quite getting there.
38
00:01:12,800 --> 00:01:14,040
So if you came for prompt tips,
39
00:01:14,040 --> 00:01:15,480
this one's going to disappoint you.
40
00:01:15,480 --> 00:01:18,400
We're going under the interface into the architecture.
41
00:01:18,400 --> 00:01:20,000
That's where the real answer lives.
42
00:01:20,000 --> 00:01:22,840
The day one copilot promise.
43
00:01:22,840 --> 00:01:25,000
Think back to how copilot got pitched to you
44
00:01:25,000 --> 00:01:26,920
or to your leadership or in whatever deck
45
00:01:26,920 --> 00:01:28,800
convinced someone to write the check.
46
00:01:28,800 --> 00:01:30,360
Ask questions in plain English.
47
00:01:30,360 --> 00:01:33,040
Get answers grounded in your organization's own data.
48
00:01:33,040 --> 00:01:34,480
No more digging through SharePoint.
49
00:01:34,480 --> 00:01:37,120
No more pinging three people on teams to find one document.
50
00:01:37,120 --> 00:01:38,720
Just ask and copilot knows.
51
00:01:38,720 --> 00:01:40,760
Microsoft backed that pitch with numbers.
52
00:01:40,760 --> 00:01:42,960
11 minutes saved per day per user.
53
00:01:42,960 --> 00:01:46,080
Do the math across a year and that's close to a full work week
54
00:01:46,080 --> 00:01:47,400
handed back to every employee.
55
00:01:47,400 --> 00:01:49,400
That's a real number and for a lot of organizations
56
00:01:49,400 --> 00:01:51,720
it's the number that justified the license cost.
57
00:01:51,720 --> 00:01:53,960
But here's what nobody set out loud during that pitch.
58
00:01:53,960 --> 00:01:56,600
Baked into asked questions and get grounded answers
59
00:01:56,600 --> 00:01:58,440
is an implicit promise.
60
00:01:58,440 --> 00:02:00,560
This thing gets better as you use it.
61
00:02:00,560 --> 00:02:03,280
You feed it more meetings, more documents, more decisions
62
00:02:03,280 --> 00:02:06,120
and it becomes sharper, more attuned to how your organization
63
00:02:06,120 --> 00:02:06,840
actually works.
64
00:02:06,840 --> 00:02:09,200
That's the story everyone heard, even if nobody wrote it
65
00:02:09,200 --> 00:02:10,000
on a slide.
66
00:02:10,000 --> 00:02:11,800
What actually ships is something different.
67
00:02:11,800 --> 00:02:14,280
It's a retrieval system wrapped in a chat interface.
68
00:02:14,280 --> 00:02:16,280
Dressed up just enough to feel like it's learning.
69
00:02:16,280 --> 00:02:17,080
It isn't.
70
00:02:17,080 --> 00:02:19,840
Every time someone types a question, copilot goes and looks,
71
00:02:19,840 --> 00:02:20,760
it doesn't recall.
72
00:02:20,760 --> 00:02:22,720
It searches, finds what looks relevant
73
00:02:22,720 --> 00:02:24,360
and generates a response from that.
74
00:02:24,360 --> 00:02:26,240
Ask again tomorrow and it searches again
75
00:02:26,240 --> 00:02:28,000
as if yesterday never happened.
76
00:02:28,000 --> 00:02:30,560
Now you might be thinking this sounds like a UX complaint,
77
00:02:30,560 --> 00:02:31,720
a minor annoyance.
78
00:02:31,720 --> 00:02:32,560
It's not.
79
00:02:32,560 --> 00:02:34,320
This is a governance and adoption problem
80
00:02:34,320 --> 00:02:36,360
and those are expensive in a completely different way
81
00:02:36,360 --> 00:02:37,800
than a clunky interface.
82
00:02:37,800 --> 00:02:39,800
When an assistant gives inconsistent answers
83
00:02:39,800 --> 00:02:41,480
to the same underlying question,
84
00:02:41,480 --> 00:02:43,440
people stop trusting it for anything that matters.
85
00:02:43,440 --> 00:02:44,760
They keep it around for quick summaries
86
00:02:44,760 --> 00:02:46,160
and calendar questions, sure.
87
00:02:46,160 --> 00:02:48,520
But the moment a decision actually rides on the answer,
88
00:02:48,520 --> 00:02:49,720
they go back to asking a person.
89
00:02:49,720 --> 00:02:50,680
That's not a bug report.
90
00:02:50,680 --> 00:02:52,520
That's an organization quietly deciding
91
00:02:52,520 --> 00:02:54,800
the tool isn't reliable enough to lean on.
92
00:02:54,800 --> 00:02:56,200
And once that decision gets made,
93
00:02:56,200 --> 00:02:58,000
even informally, adoption stores.
94
00:02:58,000 --> 00:02:59,640
Not because people didn't try copilot
95
00:02:59,640 --> 00:03:01,680
because they tried it, got burned once or twice
96
00:03:01,680 --> 00:03:03,040
by an inconsistent answer
97
00:03:03,040 --> 00:03:04,560
and adjusted their behavior around it.
98
00:03:04,560 --> 00:03:05,840
Nobody files a ticket for that.
99
00:03:05,840 --> 00:03:08,240
It just shows up months later as flat usage numbers
100
00:03:08,240 --> 00:03:11,000
and a leadership team asking why the tool they paid for
101
00:03:11,000 --> 00:03:13,200
isn't delivering the return they were promised.
102
00:03:13,200 --> 00:03:16,120
So the gap between the pitch and the product isn't cosmetic.
103
00:03:16,120 --> 00:03:18,840
It's the difference between a tool that compounds in value
104
00:03:18,840 --> 00:03:20,760
and one that just sits there, static,
105
00:03:20,760 --> 00:03:22,400
no matter how long you've had it deployed,
106
00:03:22,400 --> 00:03:24,160
to actually see where that gap comes from.
107
00:03:24,160 --> 00:03:26,280
You have to look at what's happening mechanically
108
00:03:26,280 --> 00:03:28,840
every single time someone types a question into copilot
109
00:03:28,840 --> 00:03:31,160
because the behavior we just described isn't random.
110
00:03:31,160 --> 00:03:32,880
It's the direct, predictable result
111
00:03:32,880 --> 00:03:35,720
of how the system is built underneath the chat window.
112
00:03:35,720 --> 00:03:37,000
What drag actually does?
113
00:03:37,000 --> 00:03:38,200
So let's name it directly.
114
00:03:38,200 --> 00:03:40,320
Our Ag stands for retrieval augmented generation.
115
00:03:40,320 --> 00:03:42,640
Strip the jargon and here's what that actually means.
116
00:03:42,640 --> 00:03:44,720
The system retrieves some information,
117
00:03:44,720 --> 00:03:47,320
then uses that information to help generate an answer.
118
00:03:47,320 --> 00:03:48,920
That's the whole concept.
119
00:03:48,920 --> 00:03:51,480
Retrieval, then generation, glued together.
120
00:03:51,480 --> 00:03:53,400
Here's how it works mechanically, step by step.
121
00:03:53,400 --> 00:03:54,800
First, your documents get chunked.
122
00:03:54,800 --> 00:03:57,720
Every SharePoint file, every Teams transcript,
123
00:03:57,720 --> 00:04:00,400
every policy PDF gets sliced into smaller pieces,
124
00:04:00,400 --> 00:04:02,240
usually a few hundred words at a time.
125
00:04:02,240 --> 00:04:04,000
Second, each of those chunks gets converted
126
00:04:04,000 --> 00:04:05,440
into something called an embedding,
127
00:04:05,440 --> 00:04:07,240
which is just a mathematical fingerprint
128
00:04:07,240 --> 00:04:08,600
of what that chunk means.
129
00:04:08,600 --> 00:04:10,400
Third, when someone types a question,
130
00:04:10,400 --> 00:04:12,640
that question also gets turned into a fingerprint
131
00:04:12,640 --> 00:04:14,600
and the system searches for chunks
132
00:04:14,600 --> 00:04:16,160
whose fingerprints look similar.
133
00:04:16,160 --> 00:04:18,440
Fourth, whichever chunk score highest gets stuffed
134
00:04:18,440 --> 00:04:20,880
into a prompt alongside the original question.
135
00:04:20,880 --> 00:04:22,920
Fifth, the language model reads all of that
136
00:04:22,920 --> 00:04:23,920
and generates an answer.
137
00:04:23,920 --> 00:04:25,280
Now pay attention to that word chunk
138
00:04:25,280 --> 00:04:27,960
because it's doing more work than people give it credit for.
139
00:04:27,960 --> 00:04:29,720
Documents get sliced into fragments
140
00:04:29,720 --> 00:04:31,560
with no awareness of the whole.
141
00:04:31,560 --> 00:04:34,400
A 40-page project plan doesn't get read as a project plan.
142
00:04:34,400 --> 00:04:37,200
It gets cut into 15 or 20 disconnected pieces.
143
00:04:37,200 --> 00:04:39,360
Each one judged purely on whether it resembles
144
00:04:39,360 --> 00:04:40,800
the question being asked.
145
00:04:40,800 --> 00:04:43,000
The system never holds the full document in mind.
146
00:04:43,000 --> 00:04:45,120
It holds fragments and it bets that the right fragments
147
00:04:45,120 --> 00:04:47,680
to stick together will look enough like a coherent answer.
148
00:04:47,680 --> 00:04:49,680
This is a search engine with a language model
149
00:04:49,680 --> 00:04:50,760
bolted on the end.
150
00:04:50,760 --> 00:04:53,440
It's not a thinking system and it was never built to be one.
151
00:04:53,440 --> 00:04:55,240
The intelligence you're seeing when co-pilot
152
00:04:55,240 --> 00:04:58,080
gives you a good answer isn't co-pilot understanding
153
00:04:58,080 --> 00:04:59,120
your organization.
154
00:04:59,120 --> 00:05:01,560
It's a language model doing a genuinely impressive job
155
00:05:01,560 --> 00:05:04,000
of stitching together whatever fragments got handed to it.
156
00:05:04,000 --> 00:05:05,120
That's a real skill.
157
00:05:05,120 --> 00:05:08,080
It's just not memory and it's not comprehension of the whole.
158
00:05:08,080 --> 00:05:09,640
If you want a simple way to picture this,
159
00:05:09,640 --> 00:05:12,280
think about typing a question into SharePoint Search.
160
00:05:12,280 --> 00:05:14,880
You type in a phrase, SharePoint goes and finds documents
161
00:05:14,880 --> 00:05:16,800
that contain similar words or concepts
162
00:05:16,800 --> 00:05:18,360
and it hands you a list of results.
163
00:05:18,360 --> 00:05:19,840
Ragn does almost exactly that.
164
00:05:19,840 --> 00:05:21,560
Except instead of handing you the list,
165
00:05:21,560 --> 00:05:24,360
it reads the top few results and writes you a paragraph summarizing
166
00:05:24,360 --> 00:05:24,760
them.
167
00:05:24,760 --> 00:05:27,880
Same search, same fragments, just a nicer wrapper on the output.
168
00:05:27,880 --> 00:05:29,960
And look, for a single look up, this works fine.
169
00:05:29,960 --> 00:05:31,000
Genuinely fine.
170
00:05:31,000 --> 00:05:33,640
If someone asks, what's our current PTO policy?
171
00:05:33,640 --> 00:05:35,800
Ragn finds the chunk that talks about PTO,
172
00:05:35,800 --> 00:05:37,520
generates a clean summary, and everyone
173
00:05:37,520 --> 00:05:38,480
moves on with their day.
174
00:05:38,480 --> 00:05:42,760
One question, one search, one answer, no complaints.
175
00:05:42,760 --> 00:05:44,440
The failure doesn't show up there.
176
00:05:44,440 --> 00:05:46,240
It shows up the moment you need continuity.
177
00:05:46,240 --> 00:05:47,960
The moment a question depends on something
178
00:05:47,960 --> 00:05:50,000
the system already figured out five minutes ago
179
00:05:50,000 --> 00:05:51,840
or yesterday or last month.
180
00:05:51,840 --> 00:05:54,880
Because Ragn has no concept of already figured out.
181
00:05:54,880 --> 00:05:56,400
It only has search again.
182
00:05:56,400 --> 00:05:59,160
And that's exactly where the cracks start to show.
183
00:05:59,160 --> 00:06:00,520
The stateless trap.
184
00:06:00,520 --> 00:06:02,320
Let's put a real word on what we just described
185
00:06:02,320 --> 00:06:03,400
because it matters.
186
00:06:03,400 --> 00:06:05,440
That behavior is called statelessness.
187
00:06:05,440 --> 00:06:07,760
Every query starts from zero, no matter how many times
188
00:06:07,760 --> 00:06:09,080
you've asked something related before,
189
00:06:09,080 --> 00:06:11,080
no matter how obviously connected today's question
190
00:06:11,080 --> 00:06:12,400
is to yesterday's.
191
00:06:12,400 --> 00:06:14,400
The system has no concept of before.
192
00:06:14,400 --> 00:06:16,360
There's only this query right now treated
193
00:06:16,360 --> 00:06:18,040
like the first one it's ever seen.
194
00:06:18,040 --> 00:06:20,720
Walk through what that actually looks like inside co-pilot.
195
00:06:20,720 --> 00:06:22,440
Someone on a project team asks,
196
00:06:22,440 --> 00:06:24,960
what's the status of the client onboarding for Meridian?
197
00:06:24,960 --> 00:06:27,160
Co-pilot searches, finds the right team's thread,
198
00:06:27,160 --> 00:06:30,080
the right sharepoint update, maybe a status email from last week
199
00:06:30,080 --> 00:06:31,600
and gives a genuinely good answer.
200
00:06:31,600 --> 00:06:33,480
Clean, accurate, useful.
201
00:06:33,480 --> 00:06:35,960
The person moves on with their day, impressed.
202
00:06:35,960 --> 00:06:38,240
Tomorrow, that same person asks a follow-up, something
203
00:06:38,240 --> 00:06:41,240
like, what's blocking us from finishing that onboarding?
204
00:06:41,240 --> 00:06:43,320
Notice what just happened in that sentence.
205
00:06:43,320 --> 00:06:44,960
That onboarding.
206
00:06:44,960 --> 00:06:46,360
A human would hear this and instantly
207
00:06:46,360 --> 00:06:48,040
connected to yesterday's conversation,
208
00:06:48,040 --> 00:06:50,320
no clarification needed, co-pilot doesn't.
209
00:06:50,320 --> 00:06:52,800
It has no idea that onboarding refers to Meridian
210
00:06:52,800 --> 00:06:55,440
because it never stored the fact that Meridian came up yesterday.
211
00:06:55,440 --> 00:06:58,800
So it searches again from scratch, hoping the word onboarding alone
212
00:06:58,800 --> 00:07:00,640
points it back to the right documents.
213
00:07:00,640 --> 00:07:02,480
Sometimes it gets lucky, sometimes it grabs
214
00:07:02,480 --> 00:07:04,680
the different clients onboarding thread entirely
215
00:07:04,680 --> 00:07:06,800
and now you've got a confidently wrong answer
216
00:07:06,800 --> 00:07:09,360
standing right next to yesterday's confidently right one.
217
00:07:09,360 --> 00:07:10,960
Here's the part that actually costs you.
218
00:07:10,960 --> 00:07:11,920
Nothing compounds.
219
00:07:11,920 --> 00:07:14,480
There's no accumulated understanding of the Meridian project
220
00:07:14,480 --> 00:07:16,600
building up over these two conversations,
221
00:07:16,600 --> 00:07:19,520
no growing sense of who's involved, what decisions got made,
222
00:07:19,520 --> 00:07:21,560
what the blockers have historically been.
223
00:07:21,560 --> 00:07:23,080
Each question is an island.
224
00:07:23,080 --> 00:07:24,960
You could ask co-pilot about the same project
225
00:07:24,960 --> 00:07:27,960
50 times over three months and it would know exactly
226
00:07:27,960 --> 00:07:31,080
as much on question 50 as it did on question one
227
00:07:31,080 --> 00:07:34,160
because it never carried anything forward between them.
228
00:07:34,160 --> 00:07:36,880
Now compare that to how an actual colleague would handle it,
229
00:07:36,880 --> 00:07:39,560
say the person who handled Meridian's onboarding last month,
230
00:07:39,560 --> 00:07:42,440
the answer to what's blocking us comes wrapped in memory.
231
00:07:42,440 --> 00:07:43,760
They remember the client pushed back
232
00:07:43,760 --> 00:07:45,920
on the integration timeline two weeks ago.
233
00:07:45,920 --> 00:07:47,640
They remember a decision got made in a meeting
234
00:07:47,640 --> 00:07:48,920
to reprioritize.
235
00:07:48,920 --> 00:07:50,520
They don't need to redrive any of that
236
00:07:50,520 --> 00:07:52,240
because they were there and they kept it.
237
00:07:52,240 --> 00:07:55,280
Every follow-up question they answer builds on the last one,
238
00:07:55,280 --> 00:07:57,760
gets sharper, gets more useful over time.
239
00:07:57,760 --> 00:07:59,800
That's what a second conversation with a person actually
240
00:07:59,800 --> 00:08:02,280
feels like and it's exactly what a second conversation
241
00:08:02,280 --> 00:08:03,440
with co-pilot doesn't.
242
00:08:03,440 --> 00:08:05,240
So name the real cost here plainly
243
00:08:05,240 --> 00:08:06,960
because it's not what people assume.
244
00:08:06,960 --> 00:08:09,240
The cost isn't wrong answers, not usually.
245
00:08:09,240 --> 00:08:10,720
Rags often technically accurate
246
00:08:10,720 --> 00:08:12,520
on any single question you throw at it.
247
00:08:12,520 --> 00:08:15,120
The cost is shallow answers, ones that never deep in no matter
248
00:08:15,120 --> 00:08:16,800
how long you've been using the tool,
249
00:08:16,800 --> 00:08:20,000
no matter how much history exists between you and the system.
250
00:08:20,000 --> 00:08:21,560
You get competence without growth,
251
00:08:21,560 --> 00:08:24,240
fluency without accumulation.
252
00:08:24,240 --> 00:08:26,880
And here's the thing worth sitting with before we go further.
253
00:08:26,880 --> 00:08:29,120
This isn't a training problem, it's not a prompt problem
254
00:08:29,120 --> 00:08:31,240
and it's not a rollout problem you can fix
255
00:08:31,240 --> 00:08:32,640
with better change management.
256
00:08:32,640 --> 00:08:34,960
It's what happens structurally every single time
257
00:08:34,960 --> 00:08:37,680
you take a search mechanism and ask it to do memory's job.
258
00:08:37,680 --> 00:08:40,480
Search was never built to remember, it was built to find.
259
00:08:40,480 --> 00:08:42,080
And no amount of clever phrasing changes
260
00:08:42,080 --> 00:08:44,960
what the system was designed to do underneath.
261
00:08:44,960 --> 00:08:46,920
Why this isn't a prompting problem?
262
00:08:46,920 --> 00:08:50,120
At this point, someone in IT is going to have a reflex reaction
263
00:08:50,120 --> 00:08:51,600
to everything we just said.
264
00:08:51,600 --> 00:08:53,040
They're going to blame the prompts.
265
00:08:53,040 --> 00:08:55,880
Maybe the rollout was rushed, maybe users just don't know how
266
00:08:55,880 --> 00:08:57,400
to ask questions the right way yet
267
00:08:57,400 --> 00:08:59,640
and a few training sessions will close the gap.
268
00:08:59,640 --> 00:09:00,720
That instinct makes sense.
269
00:09:00,720 --> 00:09:02,280
It's usually where the blame lands first
270
00:09:02,280 --> 00:09:03,640
because it's the easiest thing to fix
271
00:09:03,640 --> 00:09:05,520
without touching the underlying system.
272
00:09:05,520 --> 00:09:08,120
Rewrite the prompt library, run a lunch and learn,
273
00:09:08,120 --> 00:09:09,640
tell people to be more specific
274
00:09:09,640 --> 00:09:12,160
and show a better phrasing helps at the margins.
275
00:09:12,160 --> 00:09:13,640
But it doesn't touch the actual problem
276
00:09:13,640 --> 00:09:16,480
and there's a number that makes this hard to ignore.
277
00:09:16,480 --> 00:09:18,000
Research on structured knowledge systems
278
00:09:18,000 --> 00:09:20,840
shows organizations using layered knowledge with LLMs
279
00:09:20,840 --> 00:09:23,080
cut AI error rates by more than 60%
280
00:09:23,080 --> 00:09:25,280
compared to standard rag setups, 60%.
281
00:09:25,280 --> 00:09:26,600
That's not a rounding error
282
00:09:26,600 --> 00:09:28,200
and it's not the kind of gap you close
283
00:09:28,200 --> 00:09:30,320
by teaching people to phrase their questions better.
284
00:09:30,320 --> 00:09:33,000
Think about what that number actually implies.
285
00:09:33,000 --> 00:09:35,240
It's not measuring how well people ask questions.
286
00:09:35,240 --> 00:09:37,240
It's measuring what the system does before a question
287
00:09:37,240 --> 00:09:38,320
ever gets typed.
288
00:09:38,320 --> 00:09:41,360
60% fewer errors didn't come from smarter users.
289
00:09:41,360 --> 00:09:43,440
It came from a completely different relationship
290
00:09:43,440 --> 00:09:45,800
between the AI and the knowledge underneath it,
291
00:09:45,800 --> 00:09:47,600
one where the knowledge had already been organized,
292
00:09:47,600 --> 00:09:50,800
connected and reconciled before anyone opened a chat window.
293
00:09:50,800 --> 00:09:52,920
So here's the reframe and it's worth sitting with
294
00:09:52,920 --> 00:09:54,920
because it changes where you point your effort.
295
00:09:54,920 --> 00:09:56,920
You can't prompt your way out of an architecture
296
00:09:56,920 --> 00:09:59,280
that has no place to store what it learns.
297
00:09:59,280 --> 00:10:01,760
It doesn't matter how carefully worded the question is.
298
00:10:01,760 --> 00:10:03,800
If the system searches fresh every time,
299
00:10:03,800 --> 00:10:06,000
then even a perfect prompt just retrieves a fresh set
300
00:10:06,000 --> 00:10:07,920
of chunks and generates a fresh answer,
301
00:10:07,920 --> 00:10:10,240
disconnected from anything that came before.
302
00:10:10,240 --> 00:10:12,360
A better prompt gets you a better single search.
303
00:10:12,360 --> 00:10:13,880
It doesn't get you memory because memory
304
00:10:13,880 --> 00:10:15,400
was never something a prompt controls.
305
00:10:15,400 --> 00:10:17,200
That's the part worth naming plainly.
306
00:10:17,200 --> 00:10:18,680
The fix doesn't live in the interface.
307
00:10:18,680 --> 00:10:20,400
It doesn't live in how questions get asked.
308
00:10:20,400 --> 00:10:23,080
It lives below all of that in how knowledge gets prepared
309
00:10:23,080 --> 00:10:25,440
and represented before Copilot ever touches it.
310
00:10:25,440 --> 00:10:27,240
If the knowledge sitting behind the assistant
311
00:10:27,240 --> 00:10:29,960
is a pile of raw documents waiting to be chunked on demand,
312
00:10:29,960 --> 00:10:31,640
no amount of prompt engineering changes
313
00:10:31,640 --> 00:10:32,640
what happens next.
314
00:10:32,640 --> 00:10:35,720
The system is built to search and search is what it will keep doing
315
00:10:35,720 --> 00:10:37,600
no matter how the question gets phrased,
316
00:10:37,600 --> 00:10:38,960
which raises the actual question worth
317
00:10:38,960 --> 00:10:40,360
spending the rest of this episode on.
318
00:10:40,360 --> 00:10:42,720
Not how do we ask Copilot better questions.
319
00:10:42,720 --> 00:10:44,840
But what does the system actually look like
320
00:10:44,840 --> 00:10:46,960
when it's built for memory instead of search
321
00:10:46,960 --> 00:10:48,960
because that system exists and it works
322
00:10:48,960 --> 00:10:50,240
on a completely different principle
323
00:10:50,240 --> 00:10:52,560
than the one copilot ships with today?
324
00:10:52,560 --> 00:10:53,880
Enter the LLM Wiki.
325
00:10:53,880 --> 00:10:55,040
Here's that system.
326
00:10:55,040 --> 00:10:56,400
It didn't come out of Redmond
327
00:10:56,400 --> 00:10:58,120
and it didn't come out of a product road map.
328
00:10:58,120 --> 00:11:00,280
It surfaced from AI research circles
329
00:11:00,280 --> 00:11:02,840
and the person most associated with articulating it clearly
330
00:11:02,840 --> 00:11:05,680
is Andre Kapati, one of the more influential voices
331
00:11:05,680 --> 00:11:07,120
in how people think about working
332
00:11:07,120 --> 00:11:08,600
with these models day to day.
333
00:11:08,600 --> 00:11:10,480
He wasn't building an enterprise product.
334
00:11:10,480 --> 00:11:12,160
He was solving his own problem,
335
00:11:12,160 --> 00:11:14,800
building a personal knowledge base for research he cared about
336
00:11:14,800 --> 00:11:17,520
and the pattern he landed on is the one worth stealing.
337
00:11:17,520 --> 00:11:18,680
The core move is simple to say
338
00:11:18,680 --> 00:11:20,160
and genuinely different in practice.
339
00:11:20,160 --> 00:11:21,560
Instead of searching raw documents
340
00:11:21,560 --> 00:11:23,440
every single time a question comes in
341
00:11:23,440 --> 00:11:25,320
and AI reads the sources once
342
00:11:25,320 --> 00:11:28,360
and builds a structured interlinked knowledge base out of them.
343
00:11:28,360 --> 00:11:29,960
Once, not once per question.
344
00:11:29,960 --> 00:11:32,320
Once, period, then it keeps that structure around
345
00:11:32,320 --> 00:11:33,320
and refers back to it.
346
00:11:33,320 --> 00:11:35,360
Picture three layers because this is where it stops
347
00:11:35,360 --> 00:11:37,160
being an abstract idea and starts
348
00:11:37,160 --> 00:11:38,760
being something you could actually build.
349
00:11:38,760 --> 00:11:40,240
Layer one is raw sources.
350
00:11:40,240 --> 00:11:42,400
These are your original documents whatever they are
351
00:11:42,400 --> 00:11:43,640
and they stay read only.
352
00:11:43,640 --> 00:11:45,400
Nobody touches them, nothing re-rides them.
353
00:11:45,400 --> 00:11:47,360
They're ground truth sitting there untouched
354
00:11:47,360 --> 00:11:49,160
the same way a source document should.
355
00:11:49,160 --> 00:11:50,560
Layer two is the Wiki itself.
356
00:11:50,560 --> 00:11:51,960
This is where the actual work happens.
357
00:11:51,960 --> 00:11:53,680
It's a set of pages, the AI writes.
358
00:11:53,680 --> 00:11:55,480
One per concept, one per entity,
359
00:11:55,480 --> 00:11:57,560
one per recurring topic, all cross linked
360
00:11:57,560 --> 00:11:59,200
to each other the way a real Wiki works.
361
00:11:59,200 --> 00:12:00,440
Not a database of chunks.
362
00:12:00,440 --> 00:12:02,800
Pages written in something like plain language
363
00:12:02,800 --> 00:12:04,280
meant to be read as a whole.
364
00:12:04,280 --> 00:12:05,680
Layer three is the schema.
365
00:12:05,680 --> 00:12:07,080
Think of it as the rule book.
366
00:12:07,080 --> 00:12:09,440
It tells the AI how to structure new pages,
367
00:12:09,440 --> 00:12:11,080
how to decide when something updates
368
00:12:11,080 --> 00:12:13,520
an existing page versus when it deserves a new one,
369
00:12:13,520 --> 00:12:15,480
how to format things consistently
370
00:12:15,480 --> 00:12:17,840
so the Wiki doesn't turn into chaos as it grows.
371
00:12:17,840 --> 00:12:21,680
Now here's the part people get wrong on first hearing this.
372
00:12:21,680 --> 00:12:23,960
They assume the Wiki is just static storage,
373
00:12:23,960 --> 00:12:26,640
a folder where documents get copied and organized once.
374
00:12:26,640 --> 00:12:28,160
It's not, it's maintained.
375
00:12:28,160 --> 00:12:30,680
Every time a new source comes in, the AI reads it
376
00:12:30,680 --> 00:12:32,000
and does real work with it.
377
00:12:32,000 --> 00:12:34,600
It updates existing pages if the new source adds detail
378
00:12:34,600 --> 00:12:36,040
to something already compiled.
379
00:12:36,040 --> 00:12:38,280
It creates new pages if the source introduces
380
00:12:38,280 --> 00:12:41,280
a genuinely new concept that didn't exist in the Wiki yet.
381
00:12:41,280 --> 00:12:43,320
And if the new source says something that contradicts
382
00:12:43,320 --> 00:12:46,240
what's already there, it doesn't just override quietly.
383
00:12:46,240 --> 00:12:48,120
It flags it, that flag is the whole point
384
00:12:48,120 --> 00:12:49,920
and we'll come back to why in a bit.
385
00:12:49,920 --> 00:12:52,440
So say the line plainly, because it's the whole argument
386
00:12:52,440 --> 00:12:53,440
in one sentence.
387
00:12:53,440 --> 00:12:55,480
Knowledge gets compiled once then kept current.
388
00:12:55,480 --> 00:12:57,440
That's the exact opposite of what Ragn does.
389
00:12:57,440 --> 00:13:00,560
Ragn rebuilds understanding from scratch on every query.
390
00:13:00,560 --> 00:13:02,840
A Wiki builds it once and then just maintains it,
391
00:13:02,840 --> 00:13:04,800
the way you'd maintain anything that's actually alive
392
00:13:04,800 --> 00:13:07,680
and growing instead of something you keep rebuilding from nothing.
393
00:13:07,680 --> 00:13:10,280
To see why this distinction actually matters for co-pilot
394
00:13:10,280 --> 00:13:13,120
and not just as a nice idea for personal research projects,
395
00:13:13,120 --> 00:13:15,160
you need to put the two systems side by side
396
00:13:15,160 --> 00:13:17,800
and watch where the actual work happens in each one.
397
00:13:17,800 --> 00:13:20,040
Ragn vs LLM Wiki side by side.
398
00:13:20,040 --> 00:13:22,720
Put the two systems next to each other and ask one question,
399
00:13:22,720 --> 00:13:24,480
where does the thinking actually happen?
400
00:13:24,480 --> 00:13:26,280
Not the answering, the thinking.
401
00:13:26,280 --> 00:13:28,640
The part where raw information turns into something
402
00:13:28,640 --> 00:13:30,240
organized enough to be useful.
403
00:13:30,240 --> 00:13:32,160
That question tells you almost everything
404
00:13:32,160 --> 00:13:34,240
you need to know about why one system compounds
405
00:13:34,240 --> 00:13:35,320
and the other doesn't.
406
00:13:35,320 --> 00:13:37,440
In Ragn, the thinking happens at query time.
407
00:13:37,440 --> 00:13:39,440
Every single question triggers the whole process
408
00:13:39,440 --> 00:13:40,360
from the beginning.
409
00:13:40,360 --> 00:13:42,120
Fresh search across the embeddings,
410
00:13:42,120 --> 00:13:44,880
fresh chunking logic deciding which fragments matter.
411
00:13:44,880 --> 00:13:47,840
Fresh synthesis, where the model reads whatever got pulled
412
00:13:47,840 --> 00:13:49,920
and stitches it into an answer on the spot.
413
00:13:49,920 --> 00:13:52,600
None of that work exists until the question shows up.
414
00:13:52,600 --> 00:13:55,240
The system has zero understanding sitting in reserve.
415
00:13:55,240 --> 00:13:58,080
It manufactures understanding, on demand every time
416
00:13:58,080 --> 00:14:00,880
and then throws it away the moment the answer gets delivered.
417
00:14:00,880 --> 00:14:04,000
In the LLM Wiki, the thinking happens at ingestion time.
418
00:14:04,000 --> 00:14:05,520
That's the entire structural difference
419
00:14:05,520 --> 00:14:07,560
and it's easy to miss because it sounds small.
420
00:14:07,560 --> 00:14:10,840
The synthesis is already done before anyone asks the question.
421
00:14:10,840 --> 00:14:14,520
When a new source comes in, the AI does the hard part right then.
422
00:14:14,520 --> 00:14:16,560
Reading it, deciding what it means,
423
00:14:16,560 --> 00:14:18,880
connecting it to what's already compiled,
424
00:14:18,880 --> 00:14:21,240
updating pages, flagging conflicts.
425
00:14:21,240 --> 00:14:24,320
By the time someone actually types a question into the system,
426
00:14:24,320 --> 00:14:25,600
there's nothing left to figure out.
427
00:14:25,600 --> 00:14:27,600
The system isn't generating understanding.
428
00:14:27,600 --> 00:14:30,000
It's retrieving understanding that already exists,
429
00:14:30,000 --> 00:14:31,920
fully formed, sitting on a page.
430
00:14:31,920 --> 00:14:34,680
That difference shows up in a number worth paying attention to.
431
00:14:34,680 --> 00:14:37,400
For focused knowledge basis, LLM Wiki can cut token usage
432
00:14:37,400 --> 00:14:40,480
by roughly 95% compared to naive rag document loading.
433
00:14:40,480 --> 00:14:42,960
95% that's not a marginal efficiency gain
434
00:14:42,960 --> 00:14:44,760
that's a different order of magnitude.
435
00:14:44,760 --> 00:14:46,960
Here's why that number exists and it's not magic.
436
00:14:46,960 --> 00:14:49,920
Rag spends tokens researching and restitching every single time.
437
00:14:49,920 --> 00:14:52,400
It has to pull multiple chunks, feed all of them
438
00:14:52,400 --> 00:14:54,520
into the prompt alongside the question,
439
00:14:54,520 --> 00:14:57,480
and let the model do the work of reconciling fragments
440
00:14:57,480 --> 00:15:00,040
that were never written to sit next to each other.
441
00:15:00,040 --> 00:15:02,800
All of that costs tokens every query, forever.
442
00:15:02,800 --> 00:15:04,120
A Wiki skips that entirely.
443
00:15:04,120 --> 00:15:05,960
You're not researching and restitching.
444
00:15:05,960 --> 00:15:08,080
You're reading a page that was already organized,
445
00:15:08,080 --> 00:15:10,880
already coherent, already written, to be read as a whole.
446
00:15:10,880 --> 00:15:13,640
Less raw material has to get shoved into the model's context window
447
00:15:13,640 --> 00:15:15,880
because someone or something already
448
00:15:15,880 --> 00:15:17,720
did the organizing work in advance.
449
00:15:17,720 --> 00:15:19,960
Now, let's be honest about where this doesn't apply
450
00:15:19,960 --> 00:15:22,720
because overselling it here would undercut the whole argument.
451
00:15:22,720 --> 00:15:26,080
Rag stays the default for large, relatively static repositories
452
00:15:26,080 --> 00:15:27,240
and that's completely fine.
453
00:15:27,240 --> 00:15:29,440
If you're sitting on hundreds of thousands of documents
454
00:15:29,440 --> 00:15:31,600
that don't change much and get searched broadly
455
00:15:31,600 --> 00:15:33,240
across a huge range of topics,
456
00:15:33,240 --> 00:15:37,040
Rag's search first design is doing exactly what it's built for.
457
00:15:37,040 --> 00:15:38,920
The Wiki pattern isn't trying to replace that.
458
00:15:38,920 --> 00:15:40,560
Where the Wiki pattern actually shines
459
00:15:40,560 --> 00:15:44,280
is below roughly 50 to 100,000 tokens of compiled knowledge.
460
00:15:44,280 --> 00:15:45,760
That's a bounded focused domain.
461
00:15:45,760 --> 00:15:47,840
Not everything your organization has ever produced,
462
00:15:47,840 --> 00:15:50,760
a specific slice of it, the stuff people ask about again and again
463
00:15:50,760 --> 00:15:52,320
compiled once in Keptrap.
464
00:15:52,320 --> 00:15:54,400
None of this stays theoretical for long though
465
00:15:54,400 --> 00:15:57,000
because there's a version of exactly this tension
466
00:15:57,000 --> 00:15:59,720
already sitting inside Microsoft 365 right now,
467
00:15:59,720 --> 00:16:01,680
quietly shaping how co-pilot behaves
468
00:16:01,680 --> 00:16:03,760
every time someone opens a chat window.
469
00:16:03,760 --> 00:16:05,960
Where co-pilot's retrieval actually lives,
470
00:16:05,960 --> 00:16:08,960
let's get concrete about where all of this actually plays out
471
00:16:08,960 --> 00:16:11,800
because up to now we've been talking architecture in the abstract.
472
00:16:11,800 --> 00:16:15,120
Inside Microsoft 365, co-pilot indexes SharePoint OneDrive,
473
00:16:15,120 --> 00:16:17,440
Teams Exchange, and whatever other connected sources
474
00:16:17,440 --> 00:16:18,920
your organization has plugged in,
475
00:16:18,920 --> 00:16:21,320
it respects your permission structure while it does this.
476
00:16:21,320 --> 00:16:23,800
So someone only sees what they're already allowed to see.
477
00:16:23,800 --> 00:16:25,440
That parts real and it matters.
478
00:16:25,440 --> 00:16:27,160
But look at what that list actually is.
479
00:16:27,160 --> 00:16:29,320
It's a description of Rag at Enterprise Scale.
480
00:16:29,320 --> 00:16:31,520
Microsoft calls it Graph Grounding.
481
00:16:31,520 --> 00:16:33,240
And the name makes it sound like something new,
482
00:16:33,240 --> 00:16:35,120
something built specifically for this moment.
483
00:16:35,120 --> 00:16:35,920
It isn't.
484
00:16:35,920 --> 00:16:37,520
Strip the branding off and its retrieval
485
00:16:37,520 --> 00:16:39,480
same as we've been describing this whole episode,
486
00:16:39,480 --> 00:16:41,240
just dressed in Enterprise clothing.
487
00:16:41,240 --> 00:16:44,280
Co-pilot reaches into Graph, pulls whatever content looks relevant
488
00:16:44,280 --> 00:16:47,120
to the question and generates an answer from those fragments.
489
00:16:47,120 --> 00:16:49,920
That's the identical mechanism running against a much bigger,
490
00:16:49,920 --> 00:16:52,400
much messier pile of source material.
491
00:16:52,400 --> 00:16:54,440
Now, the permission model deserves real credit here
492
00:16:54,440 --> 00:16:56,400
because it solves a genuinely hard problem.
493
00:16:56,400 --> 00:16:58,560
Making sure someone in finance doesn't accidentally
494
00:16:58,560 --> 00:17:01,240
see HR's compensation planning docs through a chat window
495
00:17:01,240 --> 00:17:03,640
that's not trivial and Microsoft built real infrastructure
496
00:17:03,640 --> 00:17:04,560
to handle it.
497
00:17:04,560 --> 00:17:06,920
But notice exactly what that infrastructure solves.
498
00:17:06,920 --> 00:17:09,520
It solves the access problem, who's allowed to see what.
499
00:17:09,520 --> 00:17:11,840
It does absolutely nothing about the memory problem
500
00:17:11,840 --> 00:17:13,680
we spent the last several sections on.
501
00:17:13,680 --> 00:17:16,920
Permissions decide whether a document is eligible to be retrieved.
502
00:17:16,920 --> 00:17:19,680
They say nothing about whether the system remembers retrieving it
503
00:17:19,680 --> 00:17:22,240
or builds on it or connects it to the conversation
504
00:17:22,240 --> 00:17:23,240
from yesterday.
505
00:17:23,240 --> 00:17:25,760
Access and memory are two entirely separate questions
506
00:17:25,760 --> 00:17:27,520
and solving one doesn't touch the other.
507
00:17:27,520 --> 00:17:29,240
Here's where it gets interesting though.
508
00:17:29,240 --> 00:17:32,240
Microsoft's own guidance quietly admits something worth pausing
509
00:17:32,240 --> 00:17:32,720
on.
510
00:17:32,720 --> 00:17:35,000
Admins can now register specific SharePoint sites
511
00:17:35,000 --> 00:17:36,800
as knowledge sources for Co-pilot.
512
00:17:36,800 --> 00:17:38,000
Read that carefully.
513
00:17:38,000 --> 00:17:40,240
That feature only makes sense if not all content
514
00:17:40,240 --> 00:17:42,360
is equally trustworthy or equally structured
515
00:17:42,360 --> 00:17:44,920
or equally worth trusting Co-pilot to pull from.
516
00:17:44,920 --> 00:17:47,040
If every SharePoint site were equally reliable,
517
00:17:47,040 --> 00:17:49,800
there'd be no reason to let admins flag some of them as special.
518
00:17:49,800 --> 00:17:52,400
The existence of that toggle is Microsoft tacitly admitting
519
00:17:52,400 --> 00:17:53,440
the obvious.
520
00:17:53,440 --> 00:17:55,720
Some of what's sitting in your tenant is good, organized,
521
00:17:55,720 --> 00:17:59,080
current, and some of it is stale, duplicated, or just noise.
522
00:17:59,080 --> 00:18:01,000
And Co-pilot, left to its own devices,
523
00:18:01,000 --> 00:18:02,360
treats all of it the same.
524
00:18:02,360 --> 00:18:04,960
That single feature is Microsoft nudging your organization
525
00:18:04,960 --> 00:18:08,200
toward curation without ever fully naming the underlying issue
526
00:18:08,200 --> 00:18:09,160
out loud.
527
00:18:09,160 --> 00:18:11,600
It's a quiet acknowledgement wrapped in an admin setting.
528
00:18:11,600 --> 00:18:13,120
Nobody in the documentation says,
529
00:18:13,120 --> 00:18:16,400
our retrieval system can't tell good content from bad content,
530
00:18:16,400 --> 00:18:18,000
so we built your workaround.
531
00:18:18,000 --> 00:18:20,440
But that's functionally what registering a knowledge source
532
00:18:20,440 --> 00:18:21,040
does.
533
00:18:21,040 --> 00:18:23,160
It's damage control for a structural limitation
534
00:18:23,160 --> 00:18:24,600
presented as a feature.
535
00:18:24,600 --> 00:18:28,040
And curation, as a starting point, is genuinely useful.
536
00:18:28,040 --> 00:18:30,040
But it's not the same thing as compilation,
537
00:18:30,040 --> 00:18:31,560
and the difference between those two words
538
00:18:31,560 --> 00:18:33,880
is exactly where this conversation needs to go next.
539
00:18:33,880 --> 00:18:35,680
Curation isn't compilation.
540
00:18:35,680 --> 00:18:37,600
So let's draw the line clearly, because it's
541
00:18:37,600 --> 00:18:39,360
easy to conflate these two ideas.
542
00:18:39,360 --> 00:18:42,080
And conflating them is exactly how organizations end up
543
00:18:42,080 --> 00:18:43,920
thinking they've solved the memory problem
544
00:18:43,920 --> 00:18:46,520
when they've only solved the access problem twice.
545
00:18:46,520 --> 00:18:48,760
Marking a SharePoint site as a trusted knowledge source
546
00:18:48,760 --> 00:18:50,120
tells Co-pilot where to look.
547
00:18:50,120 --> 00:18:50,800
That's all it does.
548
00:18:50,800 --> 00:18:51,480
It's a pointer.
549
00:18:51,480 --> 00:18:53,840
It says, when you're searching, wait this location more heavily
550
00:18:53,840 --> 00:18:54,880
than that one.
551
00:18:54,880 --> 00:18:56,760
What it doesn't do, what it can't do,
552
00:18:56,760 --> 00:18:59,040
is tell Co-pilot what it learned from that site.
553
00:18:59,040 --> 00:19:01,040
There's no learning happening in the act
554
00:19:01,040 --> 00:19:02,560
of registering a knowledge source.
555
00:19:02,560 --> 00:19:04,680
There's just a narrower, more confident search.
556
00:19:04,680 --> 00:19:06,160
And that's the part we're sitting with.
557
00:19:06,160 --> 00:19:08,640
Curated sources still get retrieved fresh, chunked fresh,
558
00:19:08,640 --> 00:19:10,440
synthesized fresh every single time someone
559
00:19:10,440 --> 00:19:11,400
asks a question.
560
00:19:11,400 --> 00:19:13,520
Registring a site as trusted doesn't change
561
00:19:13,520 --> 00:19:15,520
the mechanism underneath it at all.
562
00:19:15,520 --> 00:19:17,560
It changes which pile of documents get searched.
563
00:19:17,560 --> 00:19:19,800
It doesn't change the fact that it's still a search starting
564
00:19:19,800 --> 00:19:21,280
from nothing every time.
565
00:19:21,280 --> 00:19:23,520
You've made the haste-axe more and more reliable.
566
00:19:23,520 --> 00:19:26,400
You haven't given the system a memory of what's in it.
567
00:19:26,400 --> 00:19:27,920
The wiki pattern goes one step further,
568
00:19:27,920 --> 00:19:30,000
and this is the actual structural difference.
569
00:19:30,000 --> 00:19:31,280
It doesn't just point at good sources
570
00:19:31,280 --> 00:19:33,200
and say, look here first.
571
00:19:33,200 --> 00:19:35,960
It processes those sources into standing knowledge.
572
00:19:35,960 --> 00:19:38,160
It reads them once and turns them into something
573
00:19:38,160 --> 00:19:40,800
that already exists before the next question arrives.
574
00:19:40,800 --> 00:19:41,920
That's not a better pointer.
575
00:19:41,920 --> 00:19:43,720
That's a different kind of object entirely.
576
00:19:43,720 --> 00:19:45,080
Here's the planist way to say it.
577
00:19:45,080 --> 00:19:46,840
Curation is picking better ingredients.
578
00:19:46,840 --> 00:19:48,440
You've gone through the pantry, thrown out,
579
00:19:48,440 --> 00:19:50,240
what's expired, kept what's fresh,
580
00:19:50,240 --> 00:19:53,440
and told the kitchen, cook from the shelf, not that one.
581
00:19:53,440 --> 00:19:54,520
That's real and it matters.
582
00:19:54,520 --> 00:19:56,760
But compilation is actually cooking the dish once
583
00:19:56,760 --> 00:19:59,040
and keeping it in the fridge ready to serve.
584
00:19:59,040 --> 00:20:00,560
Nobody's back in the kitchen chopping
585
00:20:00,560 --> 00:20:03,240
and searing from scratch every time someone's hungry.
586
00:20:03,240 --> 00:20:04,360
The work happened once.
587
00:20:04,360 --> 00:20:05,800
What's left is just serving it.
588
00:20:05,800 --> 00:20:08,840
Without that second step, even perfectly curated content
589
00:20:08,840 --> 00:20:11,520
still forces co-pilot to start over on every question.
590
00:20:11,520 --> 00:20:13,840
You can register every sharepoint site in your tenant
591
00:20:13,840 --> 00:20:15,760
as a trusted knowledge source, get the pantry
592
00:20:15,760 --> 00:20:17,000
as clean as it's ever going to be
593
00:20:17,000 --> 00:20:19,400
and co-pilot will still chunk those documents fresh,
594
00:20:19,400 --> 00:20:21,400
still search them fresh, still synthesize
595
00:20:21,400 --> 00:20:23,560
an answer fresh every single time.
596
00:20:23,560 --> 00:20:25,560
Curation narrows what gets searched.
597
00:20:25,560 --> 00:20:27,400
It doesn't touch how the searching works
598
00:20:27,400 --> 00:20:28,800
or what happens to the understanding
599
00:20:28,800 --> 00:20:30,280
once the answer's been delivered.
600
00:20:30,280 --> 00:20:33,560
It evaporates, same as always, waiting for the next question
601
00:20:33,560 --> 00:20:34,880
to rebuild it from nothing.
602
00:20:34,880 --> 00:20:36,480
So curation is a real improvement
603
00:20:36,480 --> 00:20:38,400
and worth doing regardless of anything else
604
00:20:38,400 --> 00:20:40,200
in this episode, but it's not the fix.
605
00:20:40,200 --> 00:20:42,600
It's step one of a process that most organizations
606
00:20:42,600 --> 00:20:45,200
stop halfway through mistaking a cleaner haystack
607
00:20:45,200 --> 00:20:46,520
for an actual memory.
608
00:20:46,520 --> 00:20:48,720
This is exactly where the week is three-layer structure
609
00:20:48,720 --> 00:20:51,840
stops being an abstract idea from AI research circles
610
00:20:51,840 --> 00:20:53,680
and starts being something you could genuinely build
611
00:20:53,680 --> 00:20:55,760
inside your own tenant.
612
00:20:55,760 --> 00:20:57,920
The three layers applied to co-pilot.
613
00:20:57,920 --> 00:21:00,720
So let's put actual Microsoft 365 nouns
614
00:21:00,720 --> 00:21:02,200
onto those three layers because that's
615
00:21:02,200 --> 00:21:04,120
where this stops sounding like a research idea
616
00:21:04,120 --> 00:21:05,800
and starts sounding like something you could sketch
617
00:21:05,800 --> 00:21:08,080
on a whiteboard this week.
618
00:21:08,080 --> 00:21:09,800
Layer one raw sources.
619
00:21:09,800 --> 00:21:11,920
In your tenant, that's your SharePoint libraries,
620
00:21:11,920 --> 00:21:14,080
your Teams meeting transcripts, exchange threads
621
00:21:14,080 --> 00:21:16,920
that carry real decisions buried in reply chains.
622
00:21:16,920 --> 00:21:19,520
Viva engage posts where someone announced a policy change
623
00:21:19,520 --> 00:21:21,520
nobody bothered to document properly.
624
00:21:21,520 --> 00:21:23,160
All of it stays exactly where it is.
625
00:21:23,160 --> 00:21:24,520
Nobody touches it, nobody edits it,
626
00:21:24,520 --> 00:21:27,120
nobody deletes it to clean things up.
627
00:21:27,120 --> 00:21:28,880
It's ground truth, sitting untouched,
628
00:21:28,880 --> 00:21:30,720
the same way it should sit, whether or not
629
00:21:30,720 --> 00:21:32,520
you ever build anything on top of it.
630
00:21:32,520 --> 00:21:34,200
Layer two, the wiki itself.
631
00:21:34,200 --> 00:21:36,360
This is the part worth picturing concretely,
632
00:21:36,360 --> 00:21:37,800
not a folder of documents.
633
00:21:37,800 --> 00:21:40,000
A maintained set of pages, one per project,
634
00:21:40,000 --> 00:21:41,920
one per client, one per recurring decision,
635
00:21:41,920 --> 00:21:45,400
your organization keeps having to re-explain to somebody new.
636
00:21:45,400 --> 00:21:47,080
Cross-linked, so the page about a client
637
00:21:47,080 --> 00:21:48,520
connects to the page about the project,
638
00:21:48,520 --> 00:21:50,760
connects to the page about the compliance interpretation
639
00:21:50,760 --> 00:21:52,560
that shaped how that project got scripted.
640
00:21:52,560 --> 00:21:54,800
This is the layer that didn't exist an hour ago
641
00:21:54,800 --> 00:21:57,560
in your tenant and does exist now once someone builds it.
642
00:21:57,560 --> 00:21:59,160
Layer three, the schema.
643
00:21:59,160 --> 00:22:01,480
Think of this as the rule book, the thing that
644
00:22:01,480 --> 00:22:04,040
decides how any of this actually happens
645
00:22:04,040 --> 00:22:05,680
when a new Teams transcript comes in,
646
00:22:05,680 --> 00:22:07,480
does it update an existing client page
647
00:22:07,480 --> 00:22:08,800
or does it spin up a new one?
648
00:22:08,800 --> 00:22:10,480
How does a page get formatted?
649
00:22:10,480 --> 00:22:12,920
So it's consistent with every other page in the wiki.
650
00:22:12,920 --> 00:22:16,000
And critically, what happens when something in that transcript
651
00:22:16,000 --> 00:22:17,680
contradicts what's already written down?
652
00:22:17,680 --> 00:22:19,160
The schema is what answers those questions
653
00:22:19,160 --> 00:22:21,760
before anyone has to answer them manually every single time.
654
00:22:21,760 --> 00:22:24,040
Now here's the thing, none of this should feel foreign
655
00:22:24,040 --> 00:22:27,160
to anyone who spent time in Microsoft 365 governance.
656
00:22:27,160 --> 00:22:28,880
Map it onto what you already understand.
657
00:22:28,880 --> 00:22:31,800
Retention labels already decide what happens to content
658
00:22:31,800 --> 00:22:34,200
based on rules you defined in advance.
659
00:22:34,200 --> 00:22:35,840
Content types already give structure
660
00:22:35,840 --> 00:22:38,760
to what would otherwise be undifferentiated files.
661
00:22:38,760 --> 00:22:40,960
Metadata schemas already tell you what fields
662
00:22:40,960 --> 00:22:43,240
a document needs before it counts as complete.
663
00:22:43,240 --> 00:22:46,000
This isn't a new discipline landing on your desk out of nowhere.
664
00:22:46,000 --> 00:22:48,360
It's an extension of information architecture work.
665
00:22:48,360 --> 00:22:50,720
Your organization has probably already been doing
666
00:22:50,720 --> 00:22:52,160
just aimed at a different output.
667
00:22:52,160 --> 00:22:54,760
Instead of the schema governing how a document gets filed,
668
00:22:54,760 --> 00:22:57,640
it's governing how a wiki page gets written and maintained.
669
00:22:57,640 --> 00:22:59,440
But here's the part that doesn't happen automatically
670
00:22:59,440 --> 00:23:01,120
and it's worth being blunt about it.
671
00:23:01,120 --> 00:23:04,680
Someone or something has to actually do the compiling.
672
00:23:04,680 --> 00:23:06,240
The three layers we just walked through
673
00:23:06,240 --> 00:23:09,640
don't assemble themselves the moment you understand the concept.
674
00:23:09,640 --> 00:23:12,520
Current copilot deployments don't ship with a background process
675
00:23:12,520 --> 00:23:13,960
that reads your SharePoint libraries
676
00:23:13,960 --> 00:23:15,840
and writes structured pages out of them.
677
00:23:15,840 --> 00:23:16,840
That work is real work.
678
00:23:16,840 --> 00:23:18,960
It has to be built or run or triggered
679
00:23:18,960 --> 00:23:21,120
by something that treats it as its job.
680
00:23:21,120 --> 00:23:24,080
Right now, in a standard Microsoft 365 environment,
681
00:23:24,080 --> 00:23:25,440
nothing is doing that job.
682
00:23:25,440 --> 00:23:28,120
Copilot searches it doesn't compile the wiki layer
683
00:23:28,120 --> 00:23:31,160
if it exists at all exists because somebody decided to build it,
684
00:23:31,160 --> 00:23:33,160
not because it came in the box.
685
00:23:33,160 --> 00:23:35,160
Which puts a fairly obvious question on the table
686
00:23:35,160 --> 00:23:36,640
and it's the one worth answering next.
687
00:23:36,640 --> 00:23:38,560
If this layer doesn't ship by default
688
00:23:38,560 --> 00:23:42,000
and it doesn't build itself, who or what is actually supposed
689
00:23:42,000 --> 00:23:44,000
to build it inside a Microsoft ecosystem?
690
00:23:44,000 --> 00:23:45,760
Because the three layer structure only matters
691
00:23:45,760 --> 00:23:47,320
if there's a real answer to that question,
692
00:23:47,320 --> 00:23:49,920
not just a diagram that looks clean on a slide.
693
00:23:49,920 --> 00:23:51,400
Who builds the wiki layer?
694
00:23:51,400 --> 00:23:53,480
Let's be honest about something before we go any further
695
00:23:53,480 --> 00:23:56,160
because glossing over it would make the rest of this episode
696
00:23:56,160 --> 00:23:57,840
sound like a feature announcement.
697
00:23:57,840 --> 00:24:00,280
This isn't a Microsoft ship capability today.
698
00:24:00,280 --> 00:24:02,320
There's no toggle in the admin center labeled
699
00:24:02,320 --> 00:24:04,760
turn on compiled knowledge.
700
00:24:04,760 --> 00:24:06,440
What we've been describing is a pattern,
701
00:24:06,440 --> 00:24:08,920
something you'd implement on top of your existing knowledge
702
00:24:08,920 --> 00:24:10,320
state, not something that arrives
703
00:24:10,320 --> 00:24:11,960
with your next licensing renewal.
704
00:24:11,960 --> 00:24:13,960
So what does implementing it actually look like?
705
00:24:13,960 --> 00:24:16,160
Picture an AI agent and if you've spent any time
706
00:24:16,160 --> 00:24:19,040
in Copilot Studio, the shape of this will feel familiar.
707
00:24:19,040 --> 00:24:20,880
Not unlike how a Copilot Studio agent
708
00:24:20,880 --> 00:24:23,120
gets built to watch a trigger and take an action
709
00:24:23,120 --> 00:24:24,800
or how a custom graph connected agent
710
00:24:24,800 --> 00:24:28,160
gets set up to reach into specific data sources on a schedule.
711
00:24:28,160 --> 00:24:31,240
Same underlying capability pointed at a different job.
712
00:24:31,240 --> 00:24:33,760
Instead of answering a user's question in the moment,
713
00:24:33,760 --> 00:24:36,200
this agent's entire purpose is reading sources
714
00:24:36,200 --> 00:24:38,200
and writing structured pages out of them.
715
00:24:38,200 --> 00:24:40,720
It goes into a SharePoint library, reads what's there,
716
00:24:40,720 --> 00:24:42,480
decides what belongs on an existing page
717
00:24:42,480 --> 00:24:44,640
and what deserves a new one and writes accordingly.
718
00:24:44,640 --> 00:24:45,680
Nobody's chatting with it.
719
00:24:45,680 --> 00:24:47,040
It's not waiting for a prompt.
720
00:24:47,040 --> 00:24:49,360
It's just working quietly in the background
721
00:24:49,360 --> 00:24:51,080
whether or not anyone's logged in that day.
722
00:24:51,080 --> 00:24:52,520
That distinction is worth sitting with
723
00:24:52,520 --> 00:24:55,040
because it's where agente orchestration actually earns
724
00:24:55,040 --> 00:24:56,880
its name instead of just being a buzzword.
725
00:24:56,880 --> 00:24:59,120
This isn't a chatbot answering questions on demand.
726
00:24:59,120 --> 00:25:01,680
It's a background process running on its own schedule,
727
00:25:01,680 --> 00:25:03,160
compiling and maintaining knowledge,
728
00:25:03,160 --> 00:25:05,880
whether or not a single person asks it anything that day.
729
00:25:05,880 --> 00:25:07,360
The value isn't in the conversation.
730
00:25:07,360 --> 00:25:09,160
It's in the fact that the conversation never has
731
00:25:09,160 --> 00:25:11,240
to start from zero because something already did
732
00:25:11,240 --> 00:25:12,960
the work before the question showed up.
733
00:25:12,960 --> 00:25:14,960
And here's the part that should feel encouraging
734
00:25:14,960 --> 00:25:15,800
rather than daunting.
735
00:25:15,800 --> 00:25:18,240
None of the infrastructure to build this is missing.
736
00:25:18,240 --> 00:25:20,840
Copilot Studio already supports custom connectors,
737
00:25:20,840 --> 00:25:23,000
reaching into whatever data source you pointed at.
738
00:25:23,000 --> 00:25:24,720
It already supports scheduled flows,
739
00:25:24,720 --> 00:25:26,480
running processes on a timer without a human
740
00:25:26,480 --> 00:25:27,680
kicking them off manually.
741
00:25:27,680 --> 00:25:29,960
Those two capabilities connectors and schedules
742
00:25:29,960 --> 00:25:31,360
are most of what you need to build
743
00:25:31,360 --> 00:25:33,160
the compiling half of the system.
744
00:25:33,160 --> 00:25:35,000
The pieces exist in your tenant right now.
745
00:25:35,000 --> 00:25:36,840
They're just not assembled this way by default
746
00:25:36,840 --> 00:25:39,040
because nobody's shipped a template that says,
747
00:25:39,040 --> 00:25:40,680
"Here's how you wire a connector
748
00:25:40,680 --> 00:25:42,680
and a schedule together to build a wiki
749
00:25:42,680 --> 00:25:45,160
instead of just automating a workflow."
750
00:25:45,160 --> 00:25:47,960
Which means the real barrier here isn't technical capability.
751
00:25:47,960 --> 00:25:51,200
It's a shift in how IT thinks about what Copilot even is.
752
00:25:51,200 --> 00:25:54,400
Right now, most organizations think of Copilot as one thing,
753
00:25:54,400 --> 00:25:55,680
a chat interface.
754
00:25:55,680 --> 00:25:56,720
You type, it answers.
755
00:25:56,720 --> 00:25:58,080
That's the whole mental model.
756
00:25:58,080 --> 00:26:00,240
But once you see the wiki pattern clearly,
757
00:26:00,240 --> 00:26:02,840
Copilot stops being one system and starts being two.
758
00:26:02,840 --> 00:26:04,720
There's the compiler, the background agent quietly
759
00:26:04,720 --> 00:26:06,680
reading, writing and maintaining structured pages.
760
00:26:06,680 --> 00:26:09,080
And there's the responder, the familiar chat interface
761
00:26:09,080 --> 00:26:11,320
that answers questions, except now it's answering
762
00:26:11,320 --> 00:26:13,800
from compiled pages instead of raw fragments.
763
00:26:13,800 --> 00:26:16,440
Two systems doing two different jobs working together.
764
00:26:16,440 --> 00:26:19,800
Most IT teams today have only built or bought the second one.
765
00:26:19,800 --> 00:26:22,760
Once you see it that way, as two systems instead of one,
766
00:26:22,760 --> 00:26:24,200
a different question starts to matter
767
00:26:24,200 --> 00:26:25,760
more than efficiency ever did.
768
00:26:25,760 --> 00:26:28,040
Because splitting the compiler from the responder
769
00:26:28,040 --> 00:26:30,440
doesn't just make answers cheaper or faster.
770
00:26:30,440 --> 00:26:31,880
It changes something more fundamental
771
00:26:31,880 --> 00:26:34,040
about whether people trust what Copilot tells them
772
00:26:34,040 --> 00:26:35,280
in the first place.
773
00:26:35,280 --> 00:26:37,760
The trust problem, RAAG, can't solve.
774
00:26:37,760 --> 00:26:39,560
There's data behind this, and it's worth putting
775
00:26:39,560 --> 00:26:41,560
a real number on it before we go further.
776
00:26:41,560 --> 00:26:43,640
Over 60% of enterprise IT leaders
777
00:26:43,640 --> 00:26:46,000
said they intended to expand generative AI
778
00:26:46,000 --> 00:26:47,200
for knowledge discovery.
779
00:26:47,200 --> 00:26:48,320
That's not a small commitment.
780
00:26:48,320 --> 00:26:50,520
That's leadership teams planning to lean harder
781
00:26:50,520 --> 00:26:52,960
into this budgeting forward, telling their boards
782
00:26:52,960 --> 00:26:54,520
it's the next phase of the rollout.
783
00:26:54,520 --> 00:26:56,680
And then in a lot of those same organizations adoption
784
00:26:56,680 --> 00:26:59,160
stalls, not because people forgot to use the tool,
785
00:26:59,160 --> 00:27:00,520
because somewhere along the way,
786
00:27:00,520 --> 00:27:03,000
the answer started feeling inconsistent.
787
00:27:03,000 --> 00:27:05,160
And inconsistency is quietly fatal
788
00:27:05,160 --> 00:27:06,920
to whether people keep coming back.
789
00:27:06,920 --> 00:27:09,480
Here's what's worth understanding about that inconsistency,
790
00:27:09,480 --> 00:27:11,960
because it isn't random noise you can chalk up to bad luck
791
00:27:11,960 --> 00:27:13,480
or a model having an off day.
792
00:27:13,480 --> 00:27:14,600
It's structural.
793
00:27:14,600 --> 00:27:17,240
Ask the same underlying question two different ways,
794
00:27:17,240 --> 00:27:19,400
and RAAG doesn't just phrase the answer differently.
795
00:27:19,400 --> 00:27:21,640
It retrieves an entirely different set of chunks,
796
00:27:21,640 --> 00:27:23,320
different search, different fragments pulled
797
00:27:23,320 --> 00:27:25,080
from the index different material handed
798
00:27:25,080 --> 00:27:26,800
to the model to synthesize from.
799
00:27:26,800 --> 00:27:28,280
Two people asking about the same policy,
800
00:27:28,280 --> 00:27:31,240
one saying, what's our remote work policy and the other saying,
801
00:27:31,240 --> 00:27:33,080
can I work from home three days a week?
802
00:27:33,080 --> 00:27:35,480
Might get two answers that don't even agree with each other.
803
00:27:35,480 --> 00:27:38,280
Because under the hood, they triggered two separate searches
804
00:27:38,280 --> 00:27:40,640
that happened to land on different documents.
805
00:27:40,640 --> 00:27:42,240
Users notice this fast.
806
00:27:42,240 --> 00:27:44,040
And here's the part that should worry you more
807
00:27:44,040 --> 00:27:45,360
than a formal complaint would.
808
00:27:45,360 --> 00:27:46,320
They don't complain loudly.
809
00:27:46,320 --> 00:27:49,080
Nobody files a ticket that says co-pilot gave me inconsistent
810
00:27:49,080 --> 00:27:49,880
answers.
811
00:27:49,880 --> 00:27:52,400
What actually happens is quieter and harder to track.
812
00:27:52,400 --> 00:27:55,600
Someone gets burned once, maybe twice, on a question that mattered.
813
00:27:55,600 --> 00:27:57,960
And they just stop asking co-pilot the harder questions.
814
00:27:57,960 --> 00:27:59,520
They still use it for quick summaries,
815
00:27:59,520 --> 00:28:02,480
for drafting an email, for things with no real stakes attached.
816
00:28:02,480 --> 00:28:04,600
But the moment a real decision rides on the answer,
817
00:28:04,600 --> 00:28:06,160
they go find a person instead.
818
00:28:06,160 --> 00:28:08,720
That shift happens silently one person at a time.
819
00:28:08,720 --> 00:28:10,360
And it never shows up as a complaint.
820
00:28:10,360 --> 00:28:12,480
It shows up months later as usage numbers
821
00:28:12,480 --> 00:28:15,200
that look fine on paper, but never touch the questions
822
00:28:15,200 --> 00:28:16,280
that actually mattered.
823
00:28:16,280 --> 00:28:19,120
Now compare that to what a compiled wiki page does.
824
00:28:19,120 --> 00:28:21,680
Ask it the same underlying question five different ways,
825
00:28:21,680 --> 00:28:24,240
phrased five different ways, worded by five different people.
826
00:28:24,240 --> 00:28:25,800
You get the same answer every time.
827
00:28:25,800 --> 00:28:28,160
Because the synthesis already happened once carefully
828
00:28:28,160 --> 00:28:30,680
before any of those five people ever typed anything.
829
00:28:30,680 --> 00:28:32,440
There's no fresh search rolling the dice
830
00:28:32,440 --> 00:28:33,920
on which fragments get pulled.
831
00:28:33,920 --> 00:28:36,320
There's one page already reconciled, already settled,
832
00:28:36,320 --> 00:28:38,840
and every version of the question just gets pointed at it.
833
00:28:38,840 --> 00:28:40,840
This is the part worth naming directly,
834
00:28:40,840 --> 00:28:44,480
because it's easy to file under nice to have and move past.
835
00:28:44,480 --> 00:28:46,200
Consistency isn't a polish feature.
836
00:28:46,200 --> 00:28:47,280
It's a trust mechanism.
837
00:28:47,280 --> 00:28:49,400
It's the actual thing that decides whether someone
838
00:28:49,400 --> 00:28:52,880
is willing to let a tools answer carry weight in a real decision,
839
00:28:52,880 --> 00:28:55,080
or whether they keep it around for the easy stuff
840
00:28:55,080 --> 00:28:58,520
and quietly root anything important through a person instead.
841
00:28:58,520 --> 00:29:00,400
Once trust breaks on the hard questions,
842
00:29:00,400 --> 00:29:02,240
it doesn't come back just because the interface
843
00:29:02,240 --> 00:29:03,640
gets a new code of paint.
844
00:29:03,640 --> 00:29:05,280
And trust is only half the payoff here.
845
00:29:05,280 --> 00:29:07,960
The other half shows up somewhere you might not expect.
846
00:29:07,960 --> 00:29:10,520
And how much redundant work is happening across your organization
847
00:29:10,520 --> 00:29:11,560
right now.
848
00:29:11,560 --> 00:29:14,040
Work that exists purely because nobody trusted search
849
00:29:14,040 --> 00:29:17,120
to surface the answer that already existed.
850
00:29:17,120 --> 00:29:19,080
Killing redundant knowledge creation?
851
00:29:19,080 --> 00:29:20,880
There's another number worth putting on the table
852
00:29:20,880 --> 00:29:22,600
and it points at something most organizations
853
00:29:22,600 --> 00:29:24,280
never think to measure.
854
00:29:24,280 --> 00:29:26,880
Early adopters of AI-based knowledge discovery report
855
00:29:26,880 --> 00:29:30,360
a 25% to 40% reduction in redundant knowledge creation.
856
00:29:30,360 --> 00:29:32,440
That's a strange phrase, redundant knowledge creation.
857
00:29:32,440 --> 00:29:34,000
So let's unpack what it actually means.
858
00:29:34,000 --> 00:29:36,480
Because once you see it, you start noticing it everywhere.
859
00:29:36,480 --> 00:29:38,960
Here's the mechanism, and it's simpler than it sounds.
860
00:29:38,960 --> 00:29:41,840
When nobody trusts search to surface the right answer,
861
00:29:41,840 --> 00:29:43,000
people don't just give up.
862
00:29:43,000 --> 00:29:43,600
They rebuild.
863
00:29:43,600 --> 00:29:46,000
Someone needs a document explaining how a process works.
864
00:29:46,000 --> 00:29:47,680
They run a search, get nothing solid,
865
00:29:47,680 --> 00:29:49,920
and instead of digging further, they just write it again
866
00:29:49,920 --> 00:29:51,080
from scratch.
867
00:29:51,080 --> 00:29:53,120
Except that document already exists.
868
00:29:53,120 --> 00:29:55,840
It's sitting in a team's thread from four months ago
869
00:29:55,840 --> 00:29:58,280
or buried on page three of an old SharePoint site.
870
00:29:58,280 --> 00:29:59,840
Nobody thinks to check anymore.
871
00:29:59,840 --> 00:30:01,080
The information isn't missing.
872
00:30:01,080 --> 00:30:04,040
It's unfindable, which functionally amounts to the same thing.
873
00:30:04,040 --> 00:30:07,360
And unfindable information gets recreated over and over
874
00:30:07,360 --> 00:30:09,240
by different people who have no idea.
875
00:30:09,240 --> 00:30:11,400
They're duplicating work that already happened.
876
00:30:11,400 --> 00:30:13,640
This is where a compiled weeky changes the math.
877
00:30:13,640 --> 00:30:15,240
Instead of a search turning up nothing
878
00:30:15,240 --> 00:30:18,200
useful or five slightly different fragments that half answer
879
00:30:18,200 --> 00:30:20,520
the question, it surfaces the existing answer
880
00:30:20,520 --> 00:30:24,120
directly with a citation pointing back to where it came from.
881
00:30:24,120 --> 00:30:26,520
Someone doesn't spend an afternoon reconstructing a document
882
00:30:26,520 --> 00:30:28,680
that already lives in the company's collective memory.
883
00:30:28,680 --> 00:30:30,560
They find it in minutes because the weeky
884
00:30:30,560 --> 00:30:33,120
already did the work of deciding that this concept exists
885
00:30:33,120 --> 00:30:34,320
and here's where it lives.
886
00:30:34,320 --> 00:30:36,240
And this connects back to something we touched on earlier,
887
00:30:36,240 --> 00:30:38,600
but it's worth stating in its sharpest form here.
888
00:30:38,600 --> 00:30:40,680
Ragtreeze every piece of content is equally hard
889
00:30:40,680 --> 00:30:41,880
to find every single time.
890
00:30:41,880 --> 00:30:44,040
It doesn't matter if someone asked the exact same question
891
00:30:44,040 --> 00:30:45,960
last week and got a great answer.
892
00:30:45,960 --> 00:30:48,280
The system starts over searching the same haystack
893
00:30:48,280 --> 00:30:50,280
with the same odds of missing what it's looking for,
894
00:30:50,280 --> 00:30:51,840
a wiki flips that completely.
895
00:30:51,840 --> 00:30:54,480
It treats concepts as already found permanently.
896
00:30:54,480 --> 00:30:57,440
Once something's compiled onto a page, it stays found.
897
00:30:57,440 --> 00:31:00,280
Nobody has to get lucky with their search terms ever again.
898
00:31:00,280 --> 00:31:02,920
Frame the business case here in the planist terms possible
899
00:31:02,920 --> 00:31:05,320
because this is where a lot of the argument in this episode
900
00:31:05,320 --> 00:31:08,080
stops being architectural and starts being financial.
901
00:31:08,080 --> 00:31:10,160
Less duplicated effort isn't a technical
902
00:31:10,160 --> 00:31:12,120
when you put on a slide about system design.
903
00:31:12,120 --> 00:31:16,000
It's hours, actual hours given back to people doing actual work
904
00:31:16,000 --> 00:31:17,560
instead of quietly rebuilding something
905
00:31:17,560 --> 00:31:19,600
that already existed somewhere in the tenant,
906
00:31:19,600 --> 00:31:22,800
multiply that across a few hundred employees over a year
907
00:31:22,800 --> 00:31:24,680
and you're not talking about a nice efficiency gain.
908
00:31:24,680 --> 00:31:26,480
You're talking about a meaningful chunk
909
00:31:26,480 --> 00:31:28,400
of your organization's collective time
910
00:31:28,400 --> 00:31:30,800
spent solving a problem that had already been solved.
911
00:31:30,800 --> 00:31:32,480
But none of this holds up if the material
912
00:31:32,480 --> 00:31:34,600
feeding the wiki in the first place is a mess
913
00:31:34,600 --> 00:31:36,040
and that's the next real constraint
914
00:31:36,040 --> 00:31:37,600
worth being honest about.
915
00:31:37,600 --> 00:31:39,440
Garbage in, still garbage out.
916
00:31:39,440 --> 00:31:40,640
Let's be direct about something
917
00:31:40,640 --> 00:31:42,800
the whole wiki pattern depends on because it's the part
918
00:31:42,800 --> 00:31:45,640
that's easiest to skip past when you're excited about the architecture.
919
00:31:45,640 --> 00:31:47,680
The wiki doesn't fix bad source material.
920
00:31:47,680 --> 00:31:49,040
It organizes what you give it.
921
00:31:49,040 --> 00:31:51,480
It doesn't audit whether what you gave it is true, current
922
00:31:51,480 --> 00:31:53,040
or even still relevant.
923
00:31:53,040 --> 00:31:54,840
That job never belonged to the compiler
924
00:31:54,840 --> 00:31:55,720
and it never will.
925
00:31:55,720 --> 00:31:57,440
Here's what that means in practice.
926
00:31:57,440 --> 00:32:00,280
If your share point is full of outdated policy documents,
927
00:32:00,280 --> 00:32:02,200
documents nobody archived properly,
928
00:32:02,200 --> 00:32:04,520
documents that describe a process your organization
929
00:32:04,520 --> 00:32:06,480
stopped using two reogs ago,
930
00:32:06,480 --> 00:32:08,480
the wiki will read every one of them faithfully
931
00:32:08,480 --> 00:32:12,000
and compile them into clean, confident looking pages.
932
00:32:12,000 --> 00:32:13,360
That's the uncomfortable part.
933
00:32:13,360 --> 00:32:15,080
The output won't look messy or uncertain.
934
00:32:15,080 --> 00:32:17,280
It'll look organized, it'll look authoritative
935
00:32:17,280 --> 00:32:19,440
and it'll be wrong presented with the same polish
936
00:32:19,440 --> 00:32:21,680
as everything else in the wiki that happens to be right.
937
00:32:21,680 --> 00:32:24,120
This is exactly where content lifecycle management
938
00:32:24,120 --> 00:32:26,120
stops being optional and it's worth pointing out
939
00:32:26,120 --> 00:32:27,440
that Microsoft has been saying this
940
00:32:27,440 --> 00:32:30,720
since its own January 2024 Copilot readiness guidance
941
00:32:30,720 --> 00:32:32,240
that guidance wasn't subtle about it.
942
00:32:32,240 --> 00:32:34,640
Content needs owners, it needs version control,
943
00:32:34,640 --> 00:32:36,440
it needs a process for retiring things
944
00:32:36,440 --> 00:32:38,840
that are no longer accurate instead of letting them sit
945
00:32:38,840 --> 00:32:40,880
indefinitely next to whatever replaced them.
946
00:32:40,880 --> 00:32:43,400
None of that is new advice invented for this episode.
947
00:32:43,400 --> 00:32:45,880
It's been sitting in Microsoft's own documentation
948
00:32:45,880 --> 00:32:47,440
for a while mostly ignored
949
00:32:47,440 --> 00:32:49,120
because it sounded like housekeeping
950
00:32:49,120 --> 00:32:50,680
rather than something that determines
951
00:32:50,680 --> 00:32:53,920
whether your AI investment actually works.
952
00:32:53,920 --> 00:32:56,200
So say it plainly because this is the part organizations
953
00:32:56,200 --> 00:32:57,160
keep skipping.
954
00:32:57,160 --> 00:32:59,800
Structured knowledge work starts before any AI
955
00:32:59,800 --> 00:33:02,080
touches your content retention labels still matter,
956
00:33:02,080 --> 00:33:04,000
content type still matter, clear ownership.
957
00:33:04,000 --> 00:33:05,280
So someone's actually responsible
958
00:33:05,280 --> 00:33:08,400
for knowing whether a document is still true, still matters.
959
00:33:08,400 --> 00:33:09,840
None of that becomes less important
960
00:33:09,840 --> 00:33:11,360
once you add a wiki layer on top.
961
00:33:11,360 --> 00:33:13,400
If anything, it becomes the entire foundation
962
00:33:13,400 --> 00:33:14,520
the wiki is standing on.
963
00:33:14,520 --> 00:33:16,120
And here's the sharper version of the risk
964
00:33:16,120 --> 00:33:17,280
the one worth sitting with.
965
00:33:17,280 --> 00:33:19,520
The wiki pattern raises the stakes on curation
966
00:33:19,520 --> 00:33:21,840
because it makes bad content look more authoritative,
967
00:33:21,840 --> 00:33:22,680
not less.
968
00:33:22,680 --> 00:33:25,240
A messy pile of raw documents, at least looks messy.
969
00:33:25,240 --> 00:33:26,960
People approach it with some skepticism.
970
00:33:26,960 --> 00:33:28,520
A beautifully compiled wiki page
971
00:33:28,520 --> 00:33:30,120
doesn't invite that same skepticism.
972
00:33:30,120 --> 00:33:32,040
It reads like something's already been verified,
973
00:33:32,040 --> 00:33:34,440
organized and settled even when the underlying source
974
00:33:34,440 --> 00:33:35,760
was garbage the whole time.
975
00:33:35,760 --> 00:33:38,120
So assuming your source material actually is solid,
976
00:33:38,120 --> 00:33:39,120
here's what happens next.
977
00:33:39,120 --> 00:33:40,760
Once new content starts arriving
978
00:33:40,760 --> 00:33:43,640
and the system has to decide what to do with it.
979
00:33:43,640 --> 00:33:45,560
When bad knowledge looks authoritative,
980
00:33:45,560 --> 00:33:47,280
here's what that means in practice.
981
00:33:47,280 --> 00:33:50,240
If your share point is full of outdated policy documents,
982
00:33:50,240 --> 00:33:51,880
documents nobody archived properly,
983
00:33:51,880 --> 00:33:54,160
documents that describe a process your organization
984
00:33:54,160 --> 00:33:56,000
stopped using two reoggs ago,
985
00:33:56,000 --> 00:33:58,240
the wiki will read every one of them faithfully
986
00:33:58,240 --> 00:34:01,160
and compile them into clean, confident looking pages.
987
00:34:01,160 --> 00:34:02,520
That's the uncomfortable part.
988
00:34:02,520 --> 00:34:04,640
The output won't look messy or uncertain.
989
00:34:04,640 --> 00:34:06,720
It'll look organized, it'll look authoritative
990
00:34:06,720 --> 00:34:09,040
and it'll be wrong, presented with the same polish
991
00:34:09,040 --> 00:34:11,440
as everything else in the wiki that happens to be right.
992
00:34:11,440 --> 00:34:13,880
This is exactly where content lifecycle management
993
00:34:13,880 --> 00:34:15,120
stops being optional.
994
00:34:15,120 --> 00:34:16,840
And it's worth pointing out that Microsoft
995
00:34:16,840 --> 00:34:20,080
has been saying this since its own January 2024
996
00:34:20,080 --> 00:34:21,760
co-pilot readiness guidance.
997
00:34:21,760 --> 00:34:23,480
That guidance wasn't subtle about it.
998
00:34:23,480 --> 00:34:25,000
Content needs owners.
999
00:34:25,000 --> 00:34:26,200
It needs version controlled.
1000
00:34:26,200 --> 00:34:28,000
It needs a process for retiring things
1001
00:34:28,000 --> 00:34:30,160
that are no longer accurate instead of letting them sit
1002
00:34:30,160 --> 00:34:32,440
indefinitely next to whatever replaced them.
1003
00:34:32,440 --> 00:34:34,800
None of that is new advice invented for this episode.
1004
00:34:34,800 --> 00:34:36,840
It's been sitting in Microsoft's own documentation
1005
00:34:36,840 --> 00:34:40,440
for a while, mostly ignored because it sounded like housekeeping
1006
00:34:40,440 --> 00:34:42,680
rather than something that determines whether your AI
1007
00:34:42,680 --> 00:34:44,560
investment actually works.
1008
00:34:44,560 --> 00:34:46,480
So say it plainly, because this is the part
1009
00:34:46,480 --> 00:34:47,960
organizations keep skipping.
1010
00:34:47,960 --> 00:34:51,680
Structured knowledge work starts before any AI touches your content.
1011
00:34:51,680 --> 00:34:54,480
Retention labels still matter, content types still matter.
1012
00:34:54,480 --> 00:34:56,480
Clear ownership, so someone's actually responsible
1013
00:34:56,480 --> 00:34:58,760
for knowing whether a document is still true still matters.
1014
00:34:58,760 --> 00:35:00,200
None of that becomes less important
1015
00:35:00,200 --> 00:35:01,960
once you add a wiki layer on top.
1016
00:35:01,960 --> 00:35:03,760
If anything, it becomes the entire foundation
1017
00:35:03,760 --> 00:35:05,160
the wiki is standing on.
1018
00:35:05,160 --> 00:35:06,840
And here's the sharper version of the risk
1019
00:35:06,840 --> 00:35:08,080
the one worth sitting with.
1020
00:35:08,080 --> 00:35:09,960
The wiki pattern raises the stakes on curation
1021
00:35:09,960 --> 00:35:12,320
because it makes bad content look more authoritative,
1022
00:35:12,320 --> 00:35:13,000
not less.
1023
00:35:13,000 --> 00:35:15,320
A messy pile of raw documents at least looks messy.
1024
00:35:15,320 --> 00:35:17,320
People approach it with some skepticism.
1025
00:35:17,320 --> 00:35:18,520
A beautifully compiled wiki page
1026
00:35:18,520 --> 00:35:20,160
doesn't invite that same skepticism.
1027
00:35:20,160 --> 00:35:22,560
It reads like something's already been verified, organized
1028
00:35:22,560 --> 00:35:24,640
and settled, even when the underlying source
1029
00:35:24,640 --> 00:35:25,920
was garbage the whole time.
1030
00:35:25,920 --> 00:35:28,320
So assuming your source material actually is solid,
1031
00:35:28,320 --> 00:35:29,960
here's what happens next.
1032
00:35:29,960 --> 00:35:33,680
Once new content starts arriving and the system has to decide
1033
00:35:33,680 --> 00:35:34,560
what to do with it.
1034
00:35:34,560 --> 00:35:37,280
So picture what actually happens when a new source lands
1035
00:35:37,280 --> 00:35:40,080
because this is where the wiki pattern stops being a diagram
1036
00:35:40,080 --> 00:35:42,240
and starts being a repeatable process.
1037
00:35:42,240 --> 00:35:44,240
A document shows up, maybe a policy update,
1038
00:35:44,240 --> 00:35:46,400
maybe a new team's transcript from a client call.
1039
00:35:46,400 --> 00:35:48,520
The system reads it once, not skims it,
1040
00:35:48,520 --> 00:35:50,960
not chunks it into fragments for later retrieval.
1041
00:35:50,960 --> 00:35:52,560
Reads it, the way you'd read something
1042
00:35:52,560 --> 00:35:54,440
if your job was to understand it well enough
1043
00:35:54,440 --> 00:35:56,040
to explain it to someone else.
1044
00:35:56,040 --> 00:35:58,280
Key concepts get pulled out during that read.
1045
00:35:58,280 --> 00:35:59,680
What's this document actually about?
1046
00:35:59,680 --> 00:36:00,520
What does it claim?
1047
00:36:00,520 --> 00:36:01,600
What does it connect to?
1048
00:36:01,600 --> 00:36:03,560
From there, one of two things happens.
1049
00:36:03,560 --> 00:36:05,360
And this is worth being precise about because it's
1050
00:36:05,360 --> 00:36:06,400
the whole mechanism.
1051
00:36:06,400 --> 00:36:08,520
If the new source adds detail to something already
1052
00:36:08,520 --> 00:36:11,040
compiled, an existing page gets updated.
1053
00:36:11,040 --> 00:36:13,200
Say there's already a page about a client's onboarding
1054
00:36:13,200 --> 00:36:15,800
process, and this new transcript adds a wrinkle nobody
1055
00:36:15,800 --> 00:36:17,040
had documented yet.
1056
00:36:17,040 --> 00:36:18,000
That page gets richer.
1057
00:36:18,000 --> 00:36:19,120
It doesn't get duplicated.
1058
00:36:19,120 --> 00:36:21,840
It doesn't sit next to an old version creating confusion.
1059
00:36:21,840 --> 00:36:23,520
It gets updated in place.
1060
00:36:23,520 --> 00:36:25,680
But if the source introduces something genuinely new,
1061
00:36:25,680 --> 00:36:27,640
something that doesn't map onto any page that already
1062
00:36:27,640 --> 00:36:29,400
exists, a new page gets created.
1063
00:36:29,400 --> 00:36:31,520
And whichever of those two things happens,
1064
00:36:31,520 --> 00:36:33,120
related ideas get linked.
1065
00:36:33,120 --> 00:36:35,080
The new page connects to the client it belongs to,
1066
00:36:35,080 --> 00:36:37,720
the project it touches, the policy that shaped it, nothing
1067
00:36:37,720 --> 00:36:40,600
sits alone.
1068
00:36:40,600 --> 00:36:42,360
Here's the part that matters most, though.
1069
00:36:42,360 --> 00:36:45,080
And it's the part, rag has no equivalent for it all.
1070
00:36:45,080 --> 00:36:47,480
What happens when the new information contradicts something
1071
00:36:47,480 --> 00:36:48,920
already sitting in the wiki?
1072
00:36:48,920 --> 00:36:51,160
Say the old page claims a process works one way,
1073
00:36:51,160 --> 00:36:52,960
and the new source says it changed.
1074
00:36:52,960 --> 00:36:54,960
The system doesn't just override the old claim
1075
00:36:54,960 --> 00:36:56,240
and move on like nothing happened.
1076
00:36:56,240 --> 00:36:57,120
It flags it.
1077
00:36:57,120 --> 00:36:58,760
Both versions get preserved with a note
1078
00:36:58,760 --> 00:37:00,440
that something shifted, and a signal
1079
00:37:00,440 --> 00:37:02,480
that someone should look at this before treating either
1080
00:37:02,480 --> 00:37:03,800
version as settled.
1081
00:37:03,800 --> 00:37:05,640
This flagging behavior isn't a nice to have
1082
00:37:05,640 --> 00:37:07,240
detailed buried in the mechanics.
1083
00:37:07,240 --> 00:37:09,720
It matters enormously for actual enterprise use,
1084
00:37:09,720 --> 00:37:11,840
because businesses change constantly.
1085
00:37:11,840 --> 00:37:14,000
And most of that change happens quietly.
1086
00:37:14,000 --> 00:37:17,040
In a policy update here, a strategy pivot there,
1087
00:37:17,040 --> 00:37:19,840
a decision that got reversed without anyone sending out
1088
00:37:19,840 --> 00:37:21,240
a formal memo about it.
1089
00:37:21,240 --> 00:37:23,400
When a policy update contradicts an old assumption
1090
00:37:23,400 --> 00:37:25,000
that's been floating around for months,
1091
00:37:25,000 --> 00:37:27,080
that contradiction doesn't get buried under a pile
1092
00:37:27,080 --> 00:37:29,280
of documents that all look equally current.
1093
00:37:29,280 --> 00:37:32,760
It gets surfaced directly to whoever's supposed to notice it.
1094
00:37:32,760 --> 00:37:34,200
Now hold that against what rag does
1095
00:37:34,200 --> 00:37:36,400
with the exact same situation, because the contrast
1096
00:37:36,400 --> 00:37:37,680
is where this really lands.
1097
00:37:37,680 --> 00:37:39,400
A stale document sits in your share point
1098
00:37:39,400 --> 00:37:41,480
next to the document that replaced it.
1099
00:37:41,480 --> 00:37:44,080
Someone asks a question, and rag retrieves both.
1100
00:37:44,080 --> 00:37:46,680
Not one, both, as two equally weighted chunks,
1101
00:37:46,680 --> 00:37:49,200
stitched together into an answer with absolutely no signal
1102
00:37:49,200 --> 00:37:50,960
about which one is actually current.
1103
00:37:50,960 --> 00:37:52,880
The model doesn't know one superseded the other.
1104
00:37:52,880 --> 00:37:54,880
It just sees two fragments that seem relevant
1105
00:37:54,880 --> 00:37:57,640
and does its best to synthesize something coherent
1106
00:37:57,640 --> 00:37:59,840
out of material that's quietly contradicting itself.
1107
00:37:59,840 --> 00:38:01,040
Nobody flagged anything.
1108
00:38:01,040 --> 00:38:03,320
Nobody got a signal that something needed a second look.
1109
00:38:03,320 --> 00:38:05,800
The contradiction just sits there, silently corrupting
1110
00:38:05,800 --> 00:38:07,640
whatever answer comes out the other end.
1111
00:38:07,640 --> 00:38:10,560
That difference, flagging versus silent averaging,
1112
00:38:10,560 --> 00:38:13,000
points at something bigger than a technical detail
1113
00:38:13,000 --> 00:38:14,440
about how updates get handled.
1114
00:38:14,440 --> 00:38:15,800
It's a concept worth naming directly,
1115
00:38:15,800 --> 00:38:17,480
because once you see it, you can't unsee it
1116
00:38:17,480 --> 00:38:19,600
in every rag deployment you've ever looked at.
1117
00:38:19,600 --> 00:38:22,040
Contradiction as a feature, not a bug.
1118
00:38:22,040 --> 00:38:25,400
Most systems, when they run into contradictory information,
1119
00:38:25,400 --> 00:38:26,720
treat it as noise.
1120
00:38:26,720 --> 00:38:28,680
Something to smooth over, average out,
1121
00:38:28,680 --> 00:38:30,760
resolve into a single, clean answer,
1122
00:38:30,760 --> 00:38:32,560
so nobody has to look at the mess underneath.
1123
00:38:32,560 --> 00:38:35,240
The wiki pattern does something almost nobody expects
1124
00:38:35,240 --> 00:38:36,440
the first time they see it.
1125
00:38:36,440 --> 00:38:38,480
It treats contradiction as a signal,
1126
00:38:38,480 --> 00:38:40,680
worth preserving, not a problem to hide.
1127
00:38:40,680 --> 00:38:42,600
Think about what a contradiction actually represents
1128
00:38:42,600 --> 00:38:43,760
inside a business.
1129
00:38:43,760 --> 00:38:45,120
It's rarely a mistake.
1130
00:38:45,120 --> 00:38:46,800
Most of the time, it's a policy shift somebody
1131
00:38:46,800 --> 00:38:49,040
made three months ago and never fully announced,
1132
00:38:49,040 --> 00:38:52,240
or a strategy pivot that quietly replaced last year's plan,
1133
00:38:52,240 --> 00:38:54,880
or a decision that got reversed after a client pushed back,
1134
00:38:54,880 --> 00:38:57,160
and only half the team ever heard about the reversal.
1135
00:38:57,160 --> 00:38:59,360
Contradictions in an organizational context
1136
00:38:59,360 --> 00:39:02,480
are usually just changed that hasn't finished propagating yet.
1137
00:39:02,480 --> 00:39:05,160
Here's where rag runs into a wall it can't get passed.
1138
00:39:05,160 --> 00:39:08,320
A stateless system has no way to represent this used to be true,
1139
00:39:08,320 --> 00:39:09,720
now it isn't.
1140
00:39:09,720 --> 00:39:11,200
There's no mechanism for that.
1141
00:39:11,200 --> 00:39:13,880
Rag just retrieves, whichever chunk seems more relevant
1142
00:39:13,880 --> 00:39:15,600
at the moment the question gets asked,
1143
00:39:15,600 --> 00:39:18,400
and depending on wording, embedding similarity,
1144
00:39:18,400 --> 00:39:20,400
whatever's closest in the vector space,
1145
00:39:20,400 --> 00:39:23,200
it might hand you the old policy or the new one.
1146
00:39:23,200 --> 00:39:24,840
Neither answer is technically wrong.
1147
00:39:24,840 --> 00:39:26,000
The document exists.
1148
00:39:26,000 --> 00:39:28,520
It's just that rag has no concept of existed,
1149
00:39:28,520 --> 00:39:30,120
but got superseded.
1150
00:39:30,120 --> 00:39:33,120
Every chunk is equally present, equally retrievable,
1151
00:39:33,120 --> 00:39:36,560
equally weighted, forever, regardless of whether it's still true.
1152
00:39:36,560 --> 00:39:39,200
A compiled wiki does something structurally different.
1153
00:39:39,200 --> 00:39:42,520
It can hold both the old state and the new state on the same page,
1154
00:39:42,520 --> 00:39:45,720
with a note attached explaining when the change happened and why,
1155
00:39:45,720 --> 00:39:49,600
if that reason is known, not too competing documents floating in a search index
1156
00:39:49,600 --> 00:39:51,360
with no relationship to each other.
1157
00:39:51,360 --> 00:39:53,600
One page, with history built into it.
1158
00:39:53,600 --> 00:39:58,160
The old assumption is still there, but it's marked as superseded, dated, explained.
1159
00:39:58,160 --> 00:40:00,960
Nothing gets erased, nothing gets hidden, it just gets sequenced.
1160
00:40:00,960 --> 00:40:03,560
This is the actual difference between an assistant that answers
1161
00:40:03,560 --> 00:40:06,360
and an assistant that understands your organization's history.
1162
00:40:06,360 --> 00:40:09,160
Answering means producing something plausible when asked.
1163
00:40:09,160 --> 00:40:10,960
Understanding history means knowing that
1164
00:40:10,960 --> 00:40:12,880
what's true today used to be something else
1165
00:40:12,880 --> 00:40:15,800
and knowing why the shift happened and being able to tell you that story
1166
00:40:15,800 --> 00:40:18,680
instead of just picking aside and hoping it's the right one.
1167
00:40:18,680 --> 00:40:20,600
Understanding history like that is one thing,
1168
00:40:20,600 --> 00:40:23,080
making it usable across an entire enterprise,
1169
00:40:23,080 --> 00:40:25,680
not just one project page or one client file,
1170
00:40:25,680 --> 00:40:27,920
but thousands of them, spanning years,
1171
00:40:27,920 --> 00:40:30,640
across every team, is an entirely different problem.
1172
00:40:30,640 --> 00:40:32,840
And it's the one worth turning to next.
1173
00:40:32,840 --> 00:40:34,080
The scale question.
1174
00:40:34,080 --> 00:40:35,520
So let's be honest about the limit here,
1175
00:40:35,520 --> 00:40:37,840
because this pattern doesn't scale the way you'd wanted to,
1176
00:40:37,840 --> 00:40:39,520
just by wishing hard enough.
1177
00:40:39,520 --> 00:40:42,360
Personal scale wikis, the kind cup he's been describing,
1178
00:40:42,360 --> 00:40:44,920
work well around 100 or so source documents.
1179
00:40:44,920 --> 00:40:48,040
That's the range where one person or one small compiling agent
1180
00:40:48,040 --> 00:40:50,480
can keep everything coherent, cross-linked and current
1181
00:40:50,480 --> 00:40:53,400
without the whole thing turning into its own management problem.
1182
00:40:53,400 --> 00:40:55,040
Your enterprise content estate is not that.
1183
00:40:55,040 --> 00:40:57,320
It's not 100 documents, it's not 1,000.
1184
00:40:57,320 --> 00:41:00,280
Most organizations are sitting on content spread across years,
1185
00:41:00,280 --> 00:41:03,640
teams and systems that nobody's fully audited, let alone compiled.
1186
00:41:03,640 --> 00:41:04,920
So where does that leave the pattern
1187
00:41:04,920 --> 00:41:06,760
once your past personal scale?
1188
00:41:06,760 --> 00:41:09,640
The honest answer is pointing towards 2026 outlooks
1189
00:41:09,640 --> 00:41:11,440
that keep landing on the same conclusion.
1190
00:41:11,440 --> 00:41:13,920
Roughly 75% of enterprise applications
1191
00:41:13,920 --> 00:41:15,960
are moving toward hybrid architectures,
1192
00:41:15,960 --> 00:41:18,200
ones that combine wiki style compilation
1193
00:41:18,200 --> 00:41:21,160
with rag and agentic search, rather than picking one
1194
00:41:21,160 --> 00:41:22,400
and walking away from the other.
1195
00:41:22,400 --> 00:41:24,520
That's not a compromise bone out of indecision.
1196
00:41:24,520 --> 00:41:26,080
It's a recognition that these two systems
1197
00:41:26,080 --> 00:41:29,040
are solving different problems, and neither one covers the other's job.
1198
00:41:29,040 --> 00:41:30,800
Here's why hybrid actually makes sense,
1199
00:41:30,800 --> 00:41:33,320
once you stop thinking of this as a competition.
1200
00:41:33,320 --> 00:41:36,680
Rag remains the default for large, relatively static repositories,
1201
00:41:36,680 --> 00:41:39,200
the sprawling stuff that doesn't get asked about constantly,
1202
00:41:39,200 --> 00:41:42,120
but still needs to be searchable when someone does need it.
1203
00:41:42,120 --> 00:41:44,080
Wiki's excel somewhere else entirely,
1204
00:41:44,080 --> 00:41:47,160
the stable, high value knowledge that gets asked about again and again,
1205
00:41:47,160 --> 00:41:48,680
the stuff worth the effort of compiling
1206
00:41:48,680 --> 00:41:50,440
because people keep coming back to it.
1207
00:41:50,440 --> 00:41:52,920
Frame it as a division of labor because that's really what it is.
1208
00:41:52,920 --> 00:41:54,880
Rag handles breadth, it covers the long tail,
1209
00:41:54,880 --> 00:41:56,480
everything that exists, but doesn't get touched
1210
00:41:56,480 --> 00:41:58,840
often enough to justify compiling it by hand.
1211
00:41:58,840 --> 00:42:01,080
The wiki handles depths on the narrow set of things
1212
00:42:01,080 --> 00:42:04,840
that actually matter enough to deserve careful, maintained synthesis.
1213
00:42:04,840 --> 00:42:06,600
Neither one is trying to do the other's job.
1214
00:42:06,600 --> 00:42:07,640
That's the whole point.
1215
00:42:07,640 --> 00:42:09,440
Worth a cost reality check here, too,
1216
00:42:09,440 --> 00:42:11,600
because none of this is free either direction.
1217
00:42:11,600 --> 00:42:15,440
MVP Rag systems typically run 15 to $40,000 to stand up,
1218
00:42:15,440 --> 00:42:17,080
covering the vector infrastructure,
1219
00:42:17,080 --> 00:42:19,360
the embedding pipeline, the retrieval tooling,
1220
00:42:19,360 --> 00:42:20,720
the wiki layer isn't free,
1221
00:42:20,720 --> 00:42:22,320
but its cost doesn't show up the same way.
1222
00:42:22,320 --> 00:42:24,680
It shows up as ongoing maintenance, not infrastructure.
1223
00:42:24,680 --> 00:42:25,960
You're not buying a bigger system.
1224
00:42:25,960 --> 00:42:27,760
You're paying for the compiling and upkeep
1225
00:42:27,760 --> 00:42:31,360
that keeps the wiki accurate as your organization keeps changing.
1226
00:42:31,360 --> 00:42:32,640
Once you see it that way,
1227
00:42:32,640 --> 00:42:34,960
breadth handled one way, depth handled another,
1228
00:42:34,960 --> 00:42:37,640
the question stops being which architecture do we pick,
1229
00:42:37,640 --> 00:42:39,640
and starts being something much more practical.
1230
00:42:39,640 --> 00:42:41,600
How does this hybrid framing actually change
1231
00:42:41,600 --> 00:42:45,960
what a rollout looks like inside a real Microsoft 365 environment?
1232
00:42:45,960 --> 00:42:47,760
What hybrid looks like for co-pilot?
1233
00:42:47,760 --> 00:42:49,880
So picture what this actually looks like once it's running,
1234
00:42:49,880 --> 00:42:52,280
not as a diagram, but as something sitting inside your tenant
1235
00:42:52,280 --> 00:42:54,040
doing two different jobs at once.
1236
00:42:54,040 --> 00:42:55,360
Co-pilot with two tiers,
1237
00:42:55,360 --> 00:42:57,680
one tier is the broad, graph grounded retrieval
1238
00:42:57,680 --> 00:42:58,800
you already have today,
1239
00:42:58,800 --> 00:43:00,560
reaching across the long tail of everything
1240
00:43:00,560 --> 00:43:02,840
in SharePoint teams exchange all of it,
1241
00:43:02,840 --> 00:43:04,880
searchable the way it's always been searchable.
1242
00:43:04,880 --> 00:43:06,720
The other tier is a compiled wiki layer,
1243
00:43:06,720 --> 00:43:08,200
sitting quietly underneath,
1244
00:43:08,200 --> 00:43:10,920
built specifically for the topics your organization asks about
1245
00:43:10,920 --> 00:43:11,960
again and again,
1246
00:43:11,960 --> 00:43:14,040
and that second tier only earns its keep
1247
00:43:14,040 --> 00:43:15,440
on specific kinds of content.
1248
00:43:15,440 --> 00:43:17,840
So it's worth naming what actually belongs there,
1249
00:43:17,840 --> 00:43:18,840
onboarding knowledge,
1250
00:43:18,840 --> 00:43:21,000
the stuff every new hire needs and every team ends up
1251
00:43:21,000 --> 00:43:24,360
re-explaning verbally because nobody trusts the search results.
1252
00:43:24,360 --> 00:43:26,520
Product positioning, the language your organization
1253
00:43:26,520 --> 00:43:27,720
has already agreed on,
1254
00:43:27,720 --> 00:43:29,640
that keeps drifting because three different people
1255
00:43:29,640 --> 00:43:31,240
describe it three different ways,
1256
00:43:31,240 --> 00:43:32,600
depending on who's asked.
1257
00:43:32,600 --> 00:43:34,040
Recurring client questions,
1258
00:43:34,040 --> 00:43:36,280
the ones that show up every quarter with new phrasing,
1259
00:43:36,280 --> 00:43:38,080
but the same underlying ask.
1260
00:43:38,080 --> 00:43:39,480
Compliance interpretations,
1261
00:43:39,480 --> 00:43:42,400
where getting it slightly wrong isn't a minor inconvenience.
1262
00:43:42,400 --> 00:43:44,240
Project histories, the decisions and reversals
1263
00:43:44,240 --> 00:43:45,960
that shaped how something got built,
1264
00:43:45,960 --> 00:43:48,280
the kind of context a new team member has no way of finding
1265
00:43:48,280 --> 00:43:50,560
without asking someone who's been there for years.
1266
00:43:50,560 --> 00:43:52,000
Notice what all of those have in common.
1267
00:43:52,000 --> 00:43:53,240
They're exactly the categories
1268
00:43:53,240 --> 00:43:55,440
where consistency and trust matter most,
1269
00:43:55,440 --> 00:43:58,280
the ones we spent real time on earlier in this episode,
1270
00:43:58,280 --> 00:43:59,520
and they're exactly the categories
1271
00:43:59,520 --> 00:44:01,800
where redundant recreation waste the most time,
1272
00:44:01,800 --> 00:44:03,600
people rebuilding something that already exists
1273
00:44:03,600 --> 00:44:04,920
because search let them down once
1274
00:44:04,920 --> 00:44:06,000
and they stop trusting it.
1275
00:44:06,000 --> 00:44:07,040
That's not a coincidence.
1276
00:44:07,040 --> 00:44:09,240
Those two lists overlap almost completely,
1277
00:44:09,240 --> 00:44:12,000
which is precisely why this is where compiling earns its cost
1278
00:44:12,000 --> 00:44:14,600
instead of just adding complexity for its own sake.
1279
00:44:14,600 --> 00:44:16,880
Here's the mechanic and it's worth describing plainly,
1280
00:44:16,880 --> 00:44:19,480
instead of dressing it up as something more dramatic than it is.
1281
00:44:19,480 --> 00:44:21,560
An agent process runs periodically,
1282
00:44:21,560 --> 00:44:23,240
not constantly, not on every query,
1283
00:44:23,240 --> 00:44:24,320
just on a schedule,
1284
00:44:24,320 --> 00:44:26,440
reading through relevant, share point libraries
1285
00:44:26,440 --> 00:44:28,560
and teams content tied to whichever domain
1286
00:44:28,560 --> 00:44:29,880
it's responsible for.
1287
00:44:29,880 --> 00:44:32,240
It compiles what it finds into structured pages,
1288
00:44:32,240 --> 00:44:33,640
updating what already exists,
1289
00:44:33,640 --> 00:44:36,760
creating new pages where something genuinely new shows up.
1290
00:44:36,760 --> 00:44:39,320
And then when someone asks Copilot a question
1291
00:44:39,320 --> 00:44:41,440
that touches one of those compiled topics,
1292
00:44:41,440 --> 00:44:44,280
Copilot's studio surfaces those pages preferentially,
1293
00:44:44,280 --> 00:44:46,720
not exclusively, not instead of graph grounding entirely,
1294
00:44:46,720 --> 00:44:48,320
just weighted higher, trusted more,
1295
00:44:48,320 --> 00:44:49,920
pulled first when the topic matches.
1296
00:44:49,920 --> 00:44:51,440
Worth being clear about what this isn't
1297
00:44:51,440 --> 00:44:53,560
because it's easy to oversell a pattern like this
1298
00:44:53,560 --> 00:44:56,080
once you've spent an episode building the case for it.
1299
00:44:56,080 --> 00:44:57,800
This isn't a rip and replace of Copilot.
1300
00:44:57,800 --> 00:44:59,680
You're not tearing out graph grounding
1301
00:44:59,680 --> 00:45:01,000
and swapping in something else.
1302
00:45:01,000 --> 00:45:02,000
It's an added layer.
1303
00:45:02,000 --> 00:45:04,160
Sitting alongside what already exists,
1304
00:45:04,160 --> 00:45:06,920
changing what grounding actually means specifically
1305
00:45:06,920 --> 00:45:09,640
for the topics your organization keeps asking about.
1306
00:45:09,640 --> 00:45:11,840
Everything else still works the way it always has.
1307
00:45:11,840 --> 00:45:13,640
The long tail is still there, still searchable,
1308
00:45:13,640 --> 00:45:15,280
still handle the old way,
1309
00:45:15,280 --> 00:45:17,960
because most of your content doesn't need this kind of treatment
1310
00:45:17,960 --> 00:45:19,040
and never will.
1311
00:45:19,040 --> 00:45:20,120
But none of this runs itself.
1312
00:45:20,120 --> 00:45:21,560
Somebody has to decide what counts
1313
00:45:21,560 --> 00:45:24,400
as a recurring topic worth compiling in the first place.
1314
00:45:24,400 --> 00:45:26,840
Somebody has to define how those pages get formatted,
1315
00:45:26,840 --> 00:45:28,040
what triggers an update,
1316
00:45:28,040 --> 00:45:29,560
when a contradiction gets flagged
1317
00:45:29,560 --> 00:45:31,360
instead of quietly resolved.
1318
00:45:31,360 --> 00:45:32,480
And that's a governance question.
1319
00:45:32,480 --> 00:45:34,480
One IT leader's can't really skip
1320
00:45:34,480 --> 00:45:37,360
once they've decided this layer is worth building at all.
1321
00:45:37,360 --> 00:45:39,360
The governance layer nobody's talking about.
1322
00:45:39,360 --> 00:45:40,600
Somebody has to own this.
1323
00:45:40,600 --> 00:45:42,920
Not in a vague, it will figure it out sense,
1324
00:45:42,920 --> 00:45:45,840
but a real role with a real name attached to it.
1325
00:45:45,840 --> 00:45:47,480
Someone has to own the schema,
1326
00:45:47,480 --> 00:45:49,400
the actual rules for what gets compiled
1327
00:45:49,400 --> 00:45:51,600
and what doesn't, how a page gets formatted,
1328
00:45:51,600 --> 00:45:54,120
so it matches every other page in the wiki.
1329
00:45:54,120 --> 00:45:56,680
And critically, when a contradiction gets escalated
1330
00:45:56,680 --> 00:45:59,440
to an actual human, instead of just sitting flagged
1331
00:45:59,440 --> 00:46:01,080
in a system nobody's watching.
1332
00:46:01,080 --> 00:46:02,800
Without that owner, the schema drifts,
1333
00:46:02,800 --> 00:46:05,480
the pages get inconsistent and the whole thing slowly turns
1334
00:46:05,480 --> 00:46:08,120
into exactly the mess it was supposed to replace.
1335
00:46:08,120 --> 00:46:10,320
Here's the good news, and it's worth saying plainly
1336
00:46:10,320 --> 00:46:12,680
because this part tends to get overcomplicated.
1337
00:46:12,680 --> 00:46:14,440
This role isn't new, it maps directly
1338
00:46:14,440 --> 00:46:16,040
onto jobs that already exist inside
1339
00:46:16,040 --> 00:46:18,320
most Microsoft 365 environments.
1340
00:46:18,320 --> 00:46:19,800
Information architects already think
1341
00:46:19,800 --> 00:46:21,560
about how content gets structured.
1342
00:46:21,560 --> 00:46:24,000
Records managers already think about lifecycle,
1343
00:46:24,000 --> 00:46:25,760
about what stays, what gets retired,
1344
00:46:25,760 --> 00:46:27,160
what needs a review cycle.
1345
00:46:27,160 --> 00:46:29,320
Content owners already carry responsibility
1346
00:46:29,320 --> 00:46:30,520
for whether something's accurate.
1347
00:46:30,520 --> 00:46:32,280
None of these people need a new job title.
1348
00:46:32,280 --> 00:46:34,160
They need this added to a job they're already doing
1349
00:46:34,160 --> 00:46:36,400
because the skill underneath it is identical.
1350
00:46:36,400 --> 00:46:39,640
Deciding what belongs where and who's accountable when it's wrong.
1351
00:46:39,640 --> 00:46:41,400
There's already a preview of where this is heading
1352
00:46:41,400 --> 00:46:43,120
and it's sitting inside Pervue right now.
1353
00:46:43,120 --> 00:46:44,840
The data loss prevention controls,
1354
00:46:44,840 --> 00:46:46,920
Microsoft built for co-pilot's web searches,
1355
00:46:46,920 --> 00:46:49,880
the ones that stop sensitive data from leaking into a prompt,
1356
00:46:49,880 --> 00:46:52,240
while still letting co-pilot ground its answer
1357
00:46:52,240 --> 00:46:54,640
in internal sources, that's the same instinct
1358
00:46:54,640 --> 00:46:56,760
applied to a narrower problem.
1359
00:46:56,760 --> 00:46:59,640
Guard rails around what an AI-driven system is allowed to see
1360
00:46:59,640 --> 00:47:01,800
and by extension, what it's allowed to compile.
1361
00:47:01,800 --> 00:47:03,400
Once you're running an agent that reads
1362
00:47:03,400 --> 00:47:06,240
across your SharePoint libraries and writes structured pages,
1363
00:47:06,240 --> 00:47:08,000
you need the same kind of thinking
1364
00:47:08,000 --> 00:47:10,480
just aimed at compilation instead of search.
1365
00:47:10,480 --> 00:47:13,280
So frame this the right way because how you frame it changes
1366
00:47:13,280 --> 00:47:14,840
whether it actually gets funded.
1367
00:47:14,840 --> 00:47:18,320
This isn't a brand new discipline landing on IT's desk out of nowhere.
1368
00:47:18,320 --> 00:47:19,840
It's an extension of governance work
1369
00:47:19,840 --> 00:47:22,120
that's probably already underway in some form.
1370
00:47:22,120 --> 00:47:24,760
If your organization already has content lifecycle policies,
1371
00:47:24,760 --> 00:47:26,600
retention rules, ownership models,
1372
00:47:26,600 --> 00:47:28,880
you already have the starting point for a Wiki schema.
1373
00:47:28,880 --> 00:47:30,440
You're not building from zero.
1374
00:47:30,440 --> 00:47:32,320
You're extending something that exists.
1375
00:47:32,320 --> 00:47:34,640
And this is really the dividing line between organizations
1376
00:47:34,640 --> 00:47:36,320
that get real value out of this pattern
1377
00:47:36,320 --> 00:47:38,600
and organizations that build something impressive looking
1378
00:47:38,600 --> 00:47:40,880
and then watch it decay within a year.
1379
00:47:40,880 --> 00:47:42,920
The ones that benefit most treat this
1380
00:47:42,920 --> 00:47:44,960
as a knowledge management initiative
1381
00:47:44,960 --> 00:47:46,880
that happens to have an AI component,
1382
00:47:46,880 --> 00:47:49,680
not an AI initiative that happens to touch knowledge,
1383
00:47:49,680 --> 00:47:52,200
that ordering matters more than it sounds like it should.
1384
00:47:52,200 --> 00:47:55,200
One puts governance first and lets the technology serve it.
1385
00:47:55,200 --> 00:47:56,720
The other bolts governance on afterward
1386
00:47:56,720 --> 00:47:58,360
once something's already gone wrong.
1387
00:47:58,360 --> 00:48:00,720
With the mechanics covered and the governance question
1388
00:48:00,720 --> 00:48:02,240
at least named honestly,
1389
00:48:02,240 --> 00:48:04,040
it's worth stepping back from the how
1390
00:48:04,040 --> 00:48:07,760
and naming what's actually different here underneath all of it.
1391
00:48:07,760 --> 00:48:09,080
What actually changes?
1392
00:48:09,080 --> 00:48:12,520
So here's the shift stated as plainly as it deserves to be stated.
1393
00:48:12,520 --> 00:48:15,280
Copilot stops being a search box with better manners,
1394
00:48:15,280 --> 00:48:18,200
one that phrases things politely and summarizes nicely,
1395
00:48:18,200 --> 00:48:19,920
but still starts from zero every time
1396
00:48:19,920 --> 00:48:22,280
and starts becoming a system that actually has
1397
00:48:22,280 --> 00:48:23,720
somewhere to put what it learns.
1398
00:48:23,720 --> 00:48:25,680
That's the whole difference, not a smarter model,
1399
00:48:25,680 --> 00:48:27,880
not a better prompt, a place for knowledge to live
1400
00:48:27,880 --> 00:48:29,760
once it's been figured out instead of a mechanism
1401
00:48:29,760 --> 00:48:31,760
that refigures it out on command forever.
1402
00:48:31,760 --> 00:48:33,560
And that difference doesn't show up on day one.
1403
00:48:33,560 --> 00:48:34,960
It shows up over time,
1404
00:48:34,960 --> 00:48:38,560
which is exactly why so many organizations miss it during a pilot.
1405
00:48:38,560 --> 00:48:40,960
A stateless system is the same tool on day 500
1406
00:48:40,960 --> 00:48:42,400
as it was on day one.
1407
00:48:42,400 --> 00:48:45,480
Ask at the same category of question a year from now
1408
00:48:45,480 --> 00:48:47,320
and it goes through the identical motions
1409
00:48:47,320 --> 00:48:48,800
it went through in week one,
1410
00:48:48,800 --> 00:48:53,160
search, chunk, retrieve, synthesize, forget.
1411
00:48:53,160 --> 00:48:55,160
Nothing about that process improves with age,
1412
00:48:55,160 --> 00:48:57,600
it doesn't get faster, it doesn't get more accurate,
1413
00:48:57,600 --> 00:48:58,600
it just repeats.
1414
00:48:58,600 --> 00:49:00,960
A compiled system by contrast is measurably different
1415
00:49:00,960 --> 00:49:04,120
a year in because every project page, every client history,
1416
00:49:04,120 --> 00:49:06,200
every resolved contradiction from month three
1417
00:49:06,200 --> 00:49:08,320
is still sitting there in month 14,
1418
00:49:08,320 --> 00:49:10,840
ready to be built on instead of rediscovered.
1419
00:49:10,840 --> 00:49:12,120
That's not a theoretical claim.
1420
00:49:12,120 --> 00:49:14,600
It shows up in the numbers we've already touched on earlier
1421
00:49:14,600 --> 00:49:16,640
in this episode around structured knowledge
1422
00:49:16,640 --> 00:49:19,520
and it's worth tying back to the sharpest version of them.
1423
00:49:19,520 --> 00:49:22,120
Enterprise knowledge graphs and structured content approaches
1424
00:49:22,120 --> 00:49:25,720
have delivered results like 287% ROI over three years
1425
00:49:25,720 --> 00:49:28,040
with multi million dollar annual benefits reported
1426
00:49:28,040 --> 00:49:29,440
in some organizations.
1427
00:49:29,440 --> 00:49:30,960
Those aren't numbers from a system
1428
00:49:30,960 --> 00:49:32,280
that got asked more questions,
1429
00:49:32,280 --> 00:49:34,320
they're numbers from a system that got smarter
1430
00:49:34,320 --> 00:49:35,680
at answering the same questions
1431
00:49:35,680 --> 00:49:37,840
because the underlying knowledge kept compounding
1432
00:49:37,840 --> 00:49:39,080
instead of resetting.
1433
00:49:39,080 --> 00:49:41,440
And that's really the mechanism worth naming directly
1434
00:49:41,440 --> 00:49:43,720
because it's easy to miss if you're only looking
1435
00:49:43,720 --> 00:49:45,160
at usage dashboards.
1436
00:49:45,160 --> 00:49:47,120
Compounding knowledge, not repeated search,
1437
00:49:47,120 --> 00:49:49,160
is what actually saves time at scale.
1438
00:49:49,160 --> 00:49:52,080
Repeated search just means people are using the tool more often.
1439
00:49:52,080 --> 00:49:54,600
It says nothing about whether the tool is getting better.
1440
00:49:54,600 --> 00:49:56,600
Compounding knowledge means the thousandth question
1441
00:49:56,600 --> 00:49:58,720
benefits from everything the system learned
1442
00:49:58,720 --> 00:50:00,680
answering the first 999.
1443
00:50:00,680 --> 00:50:02,160
That's a completely different kind of value
1444
00:50:02,160 --> 00:50:04,120
and it's the kind drag structurally can't produce
1445
00:50:04,120 --> 00:50:05,440
no matter how much you use it.
1446
00:50:05,440 --> 00:50:07,000
So say the quiet part plainly
1447
00:50:07,000 --> 00:50:08,800
because it's the sentence this whole episode
1448
00:50:08,800 --> 00:50:10,160
has been building toward.
1449
00:50:10,160 --> 00:50:12,760
An assistant that gets smarter is a fundamentally different
1450
00:50:12,760 --> 00:50:15,080
asset than one that just gets asked more questions.
1451
00:50:15,080 --> 00:50:17,720
One compounds the other just accumulates traffic.
1452
00:50:17,720 --> 00:50:20,240
Confusing the two is exactly how organizations end up
1453
00:50:20,240 --> 00:50:21,920
with impressive adoption metrics
1454
00:50:21,920 --> 00:50:24,320
and a tool that still can't hold a coherent memory
1455
00:50:24,320 --> 00:50:26,640
of its own history, which naturally raises
1456
00:50:26,640 --> 00:50:28,440
the only question left worth answering.
1457
00:50:28,440 --> 00:50:29,760
Not whether this is worth doing
1458
00:50:29,760 --> 00:50:31,680
but where an organization actually starts,
1459
00:50:31,680 --> 00:50:33,680
small and deliberate instead of trying to compile
1460
00:50:33,680 --> 00:50:34,960
everything at once.
1461
00:50:34,960 --> 00:50:36,120
What if you start today?
1462
00:50:36,120 --> 00:50:38,120
So pick one thing, not the whole knowledge estate,
1463
00:50:38,120 --> 00:50:41,120
not every department, one recurring high value domain.
1464
00:50:41,120 --> 00:50:42,480
Maybe it's a single product line
1465
00:50:42,480 --> 00:50:44,120
that support keeps asking about.
1466
00:50:44,120 --> 00:50:46,000
Maybe it's one recurring client type,
1467
00:50:46,000 --> 00:50:47,640
the kind of account whose questions repeat
1468
00:50:47,640 --> 00:50:49,480
every quarter with new phrasing.
1469
00:50:49,480 --> 00:50:51,240
Maybe it's a single compliance area
1470
00:50:51,240 --> 00:50:53,440
where getting the answer wrong actually costs something.
1471
00:50:53,440 --> 00:50:55,320
The domain doesn't matter as much as the discipline
1472
00:50:55,320 --> 00:50:56,360
of picking just one.
1473
00:50:56,360 --> 00:50:58,800
Here's what a first 90 days actually looks like
1474
00:50:58,800 --> 00:51:01,120
without dressing it up as bigger than it is.
1475
00:51:01,120 --> 00:51:03,080
Start by identifying the raw sources,
1476
00:51:03,080 --> 00:51:04,720
whatever SharePoint libraries,
1477
00:51:04,720 --> 00:51:06,600
Teams, Threads and old policy docs
1478
00:51:06,600 --> 00:51:09,320
already touch this domain, then define a schema
1479
00:51:09,320 --> 00:51:11,080
and keep it lightweight on purpose.
1480
00:51:11,080 --> 00:51:13,400
Not a governance framework spanning 50 pages,
1481
00:51:13,400 --> 00:51:15,480
just enough rules to say what a page looks like
1482
00:51:15,480 --> 00:51:17,200
and how contradictions get flagged,
1483
00:51:17,200 --> 00:51:19,480
run an initial compilation pass against that schema.
1484
00:51:19,480 --> 00:51:21,320
Then test it against real questions,
1485
00:51:21,320 --> 00:51:23,000
phrase the way actual people phrase them,
1486
00:51:23,000 --> 00:51:25,240
not the clean version you'd write in a demo script.
1487
00:51:25,240 --> 00:51:27,720
Picture the end of that pilot next to where you started.
1488
00:51:27,720 --> 00:51:30,600
Before, co-pilot retrieves five slightly different chunks
1489
00:51:30,600 --> 00:51:32,640
about the same policy and depending on
1490
00:51:32,640 --> 00:51:34,240
how someone phrased the question,
1491
00:51:34,240 --> 00:51:36,240
they get a different one of those five.
1492
00:51:36,240 --> 00:51:39,000
After, it answers from one compiled current page
1493
00:51:39,000 --> 00:51:41,120
every time, no matter how the question gets asked.
1494
00:51:41,120 --> 00:51:43,920
That's the whole test, not whether the answer sounds smart,
1495
00:51:43,920 --> 00:51:45,560
whether it's the same answer twice,
1496
00:51:45,560 --> 00:51:47,200
worth connecting this back to the research
1497
00:51:47,200 --> 00:51:48,600
we've been leaning on all episode
1498
00:51:48,600 --> 00:51:50,720
because it matters that this isn't theoretical.
1499
00:51:50,720 --> 00:51:53,560
The organization's reporting a 25 to 40% reduction
1500
00:51:53,560 --> 00:51:55,000
in redundant knowledge creation,
1501
00:51:55,000 --> 00:51:57,280
the ones seeing meaningfully lower error rates.
1502
00:51:57,280 --> 00:51:59,440
They didn't get there by rebuilding everything at once.
1503
00:51:59,440 --> 00:52:02,040
They got their one domain at a time, exactly like this.
1504
00:52:02,040 --> 00:52:05,160
Small-bounded, tested against real use before expanding,
1505
00:52:05,160 --> 00:52:06,840
and there's a thread here worth planting
1506
00:52:06,840 --> 00:52:08,680
without pulling on a too hard right now.
1507
00:52:08,680 --> 00:52:09,840
The moment you run this pilot,
1508
00:52:09,840 --> 00:52:12,680
you're going to hit real questions about who owns the schema,
1509
00:52:12,680 --> 00:52:14,840
who decides when a contradiction gets escalated,
1510
00:52:14,840 --> 00:52:17,160
who's accountable when a page goes stale.
1511
00:52:17,160 --> 00:52:18,120
That's governance.
1512
00:52:18,120 --> 00:52:21,240
And it's worth its own conversation rather than a rushed afterthought here.
1513
00:52:21,240 --> 00:52:23,760
So that's the technical case and the strategic case,
1514
00:52:23,760 --> 00:52:26,880
sitting next to each other, time to bring them together.
1515
00:52:26,880 --> 00:52:28,080
Key transformation.
1516
00:52:28,080 --> 00:52:31,320
Here's the transformation, stated as plainly as it deserves.
1517
00:52:31,320 --> 00:52:34,320
Copilot's value ceiling isn't set by the model underneath it.
1518
00:52:34,320 --> 00:52:37,440
It's set by whether your organization's knowledge gets compiled
1519
00:52:37,440 --> 00:52:40,120
or just searched over and over forever.
1520
00:52:40,120 --> 00:52:43,320
Rags will keep giving you fast, plausible, forgettable answers.
1521
00:52:43,320 --> 00:52:44,560
So that's what it's built to do,
1522
00:52:44,560 --> 00:52:45,720
and it'll keep doing it well.
1523
00:52:45,720 --> 00:52:48,480
A wiki layer gives you something structurally different,
1524
00:52:48,480 --> 00:52:51,880
an assistant that remembers what your organization has already figured out
1525
00:52:51,880 --> 00:52:54,800
instead of refiguring it out every time someone asks.
1526
00:52:54,800 --> 00:52:57,400
And here's the part worth sitting with on your way out of this episode.
1527
00:52:57,400 --> 00:53:00,920
The tools inside Microsoft 365 to build this exist today.
1528
00:53:00,920 --> 00:53:01,720
Connectors.
1529
00:53:01,720 --> 00:53:02,960
Scheduled flows.
1530
00:53:02,960 --> 00:53:04,320
Copilot studio agents.
1531
00:53:04,320 --> 00:53:05,800
All of it's already sitting in your tenant.
1532
00:53:05,800 --> 00:53:08,040
They're just not assembled this way by default.
1533
00:53:08,040 --> 00:53:09,720
So here's your challenge for this month.
1534
00:53:09,720 --> 00:53:11,760
Pick one recurring knowledge domain,
1535
00:53:11,760 --> 00:53:14,880
and go find out honestly whether Copilot is searching it
1536
00:53:14,880 --> 00:53:17,600
or actually compiling it, not what the vendor deck says,
1537
00:53:17,600 --> 00:53:19,280
what actually happens when someone asks.
1538
00:53:19,280 --> 00:53:21,920
Here's the test, and it's simple enough to run this week.
1539
00:53:21,920 --> 00:53:24,320
Ask the same question two or three different ways.
1540
00:53:24,320 --> 00:53:26,680
If the answer changes depending on how it's phrased,
1541
00:53:26,680 --> 00:53:29,480
you've found a retrieval problem, not a copilot problem.
1542
00:53:29,480 --> 00:53:31,280
That distinction is the whole episode,
1543
00:53:31,280 --> 00:53:33,240
boiled down to one homework assignment.
1544
00:53:33,240 --> 00:53:35,720
If this changed how you think about what's actually running
1545
00:53:35,720 --> 00:53:37,280
under your copilot deployment,
1546
00:53:37,280 --> 00:53:39,640
subscribe to M365FM podcast.
1547
00:53:39,640 --> 00:53:42,000
Next time we're going deeper into the governance side of this,
1548
00:53:42,000 --> 00:53:44,120
who owns the schema, who owns the wiki,
1549
00:53:44,120 --> 00:53:46,240
and what happens when that ownership isn't clear.
1550
00:53:46,240 --> 00:53:49,360
And if you want more people, more IT pros, more decision makers,
1551
00:53:49,360 --> 00:53:52,520
to run into this conversation about fixing Copilot's memory problem,
1552
00:53:52,520 --> 00:53:53,520
leave a review.
1553
00:53:53,520 --> 00:53:56,200
It's genuinely how this finds the people who need it.
1554
00:53:56,200 --> 00:53:59,000
One last thing, connect with me, Mirko Peters on LinkedIn.
1555
00:53:59,000 --> 00:54:01,800
I want to know what recurring knowledge domain you'd compile first
1556
00:54:01,800 --> 00:54:04,480
and what you think the next episode should cover.