Aug. 10, 2026

Your Copilot Has No Memory — Why RAG Fails and LLM Wiki Fixes It

Your Copilot Has No Memory — Why RAG Fails and LLM Wiki Fixes It
Your Copilot Has No Memory — Why RAG Fails and LLM Wiki Fixes It
M365 FM Podcast
Your Copilot Has No Memory — Why RAG Fails and LLM Wiki Fixes It

Microsoft Copilot can feel remarkably intelligent in a demo. It can find a project document, summarize a policy, extract information from SharePoint, and connect pieces of enterprise content into a convincing answer. But there is a fundamental architectural limitation hiding behind that experience: retrieval is not memory. Copilot can search your organization, but that does not mean it has built a persistent understanding of your organization. In this deep dive, we explore the difference between retrieving information and actually compiling organizational knowledge. We examine how Retrieval-Augmented Generation works, why traditional RAG architectures become unreliable when questions require continuity and context, and why an emerging LLM Wiki approach could provide a fundamentally different knowledge layer for enterprise AI.

COPILOT DOESN'T REMEMBER — IT RETRIEVES
The experience of using Copilot creates an important illusion. When it successfully connects information from Microsoft 365, it can appear as though the system has learned something about your organization. Ask a similar question later, however, and the system may produce a different result because it performs another retrieval operation rather than simply recalling the understanding it established previously. That distinction becomes critical when organizations move beyond basic summarization and start expecting AI to support real decisions. If essentially identical questions can produce inconsistent answers depending on which information was retrieved, users quickly become reluctant to rely on AI for business-critical work. The result can become an adoption problem rather than merely a technical problem.

WHAT RETRIEVAL-AUGMENTED GENERATION ACTUALLY DOES RAG
stands for Retrieval-Augmented Generation. At a simplified level, enterprise documents are divided into chunks, those chunks are represented through embeddings, a user's question is transformed into a comparable representation, and the system searches for semantically relevant chunks. The highest-ranking pieces of information are then supplied to the language model as context for generating its response. This architecture is extremely useful. It allows an LLM to answer questions using information that was never part of its original training data and provides a practical way to ground AI responses in enterprise content. But it also creates an important architectural constraint: the model receives fragments selected for the current query rather than maintaining a complete persistent representation of the organization's knowledge.

THE CHUNKING PROBLEM
Enterprise knowledge rarely exists as isolated paragraphs. A project plan might contain dependencies distributed across dozens of pages. A policy might reference another policy. A technical architecture could depend on decisions documented months earlier in meeting notes, Teams conversations, SharePoint pages, and design documents. Traditional RAG breaks those sources into smaller units and determines which fragments appear relevant to the current question. The model therefore sees selected pieces rather than necessarily understanding the complete document and all of its relationships. For straightforward information retrieval, that can work extremely well. For questions requiring relationships, historical context, dependencies, accumulated decisions, or reasoning across many sources, the limitations become much more visible.

THE STATELESS RAG TRAP
A conventional retrieval workflow has no inherent concept of something being "already figured out." A question is received, information is retrieved, an answer is generated, and the process effectively starts again for the next retrieval operation. The script describes this as the point where RAG's statelessness becomes a problem. That means knowledge discovered during one interaction does not automatically become durable organizational knowledge available to every future interaction. For enterprise AI, this is a major distinction. Organizations do not simply need better search. They increasingly need systems capable of maintaining a structured understanding of projects, processes, policies, technologies, people, decisions, dependencies, and the relationships connecting them.

SEARCHING FOR KNOWLEDGE VS. COMPILING KNOWLEDGE
This leads to the central architectural idea of the episode: instead of repeatedly reconstructing organizational knowledge at query time, what if AI compiled that knowledge beforehand? An LLM Wiki represents that shift in thinking. Rather than treating every enterprise document as another collection of fragments waiting for retrieval, AI can synthesize information into structured knowledge artifacts that represent what the organization currently understands. The important change is not simply another user interface. It is moving intelligence from query-time reconstruction toward persistent knowledge synthesis.

WHY AN LLM WIKI CHANGES THE MODEL
Imagine thousands of documents describing the same product, customer, project, policy, or architecture. Traditional RAG waits for a question and then tries to locate the fragments most likely to answer it. An LLM Wiki approach instead attempts to continuously transform those fragmented sources into coherent knowledge pages. Relationships, decisions, definitions, dependencies, historical context, and supporting sources can become part of a maintained knowledge representation. Copilot or another AI agent can then retrieve from a layer containing synthesized organizational understanding instead of repeatedly attempting to reconstruct that understanding from raw documents.

FROM DOCUMENT REPOSITORY TO KNOWLEDGE LAYER
This changes the role of systems such as SharePoint. SharePoint can continue serving as the authoritative repository for documents, pages, policies, presentations, meeting artifacts, and collaboration content. But raw enterprise content does not automatically constitute usable organizational knowledge. An AI-generated knowledge layer can sit above those source systems and transform scattered information into something closer to an organizational map: projects connected to decisions, policies connected to processes, systems connected to owners, and concepts connected to their supporting evidence. The goal is not to eliminate source documents. It is to make the relationships hidden inside them explicit.

WHY BETTER PROMPTS DON'T SOLVE THE ARCHITECTURE
Prompt engineering can improve how an LLM interprets retrieved context, but it cannot guarantee that the correct context was retrieved in the first place. If the retrieval layer returns incomplete fragments, misses an important dependency, or surfaces an outdated document, even an excellent model is reasoning over an incomplete information set. This is why improving the model alone cannot solve every enterprise AI problem. The quality and structure of the knowledge supplied to the model remain fundamental.

THE GOVERNANCE PROBLEM GETS BIGGER
There is also a significant warning attached to this architecture. If an organization's SharePoint environment contains obsolete policies, duplicate documentation, abandoned processes, contradictory instructions, or documents without clear ownership, an LLM Wiki can synthesize that bad information just as efficiently as it synthesizes good information. The danger is that synthesized knowledge can look considerably cleaner and more authoritative than the underlying content deserves. AI therefore makes traditional information governance more important rather than less important. Content ownership, lifecycle management, versioning, retention, archival processes, authoritative sources, metadata, and clearly defined systems of record become foundational components of AI architecture.

AI READINESS STARTS WITH CONTENT QUALITY
Organizations frequently approach Copilot readiness as a licensing, security, deployment, or training project. Those elements matter, but the underlying knowledge environment matters just as much. If nobody knows which document represents the current process, the AI cannot magically resolve the organizational ambiguity. If three departments maintain contradictory versions of a policy, AI has inherited three versions of the truth. If obsolete documentation remains searchable indefinitely, it remains potential grounding material. Enterprise AI therefore exposes knowledge-management debt that organizations could previously ignore.

WHY TRUST DETERMINES COPILOT ADOPTION
The technical consequences quickly become business consequences. Users may tolerate occasional inconsistencies when AI is used for drafting an email or summarizing a meeting. They become much less tolerant when AI is expected to explain policy, support customer decisions, interpret project status, provide compliance information, or guide operational processes. Once users experience inconsistent answers to important questions, they frequently return to trusted human experts and established manual processes. The script identifies this as a major reason why Copilot adoption can flatten even after an apparently successful rollout. Trust therefore becomes an architectural requirement, not merely an adoption metric.

Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-modern-work-security-and-productivity-with-microsoft-365--6704921/support.

🚀 Want to be part of m365.fm?

Then stop just listening… and start showing up.

👉 Connect with me on LinkedIn and let’s make something happen:

  • 🎙️ Be a podcast guest and share your story
  • 🎧 Host your own episode (yes, seriously)
  • 💡 Pitch topics the community actually wants to hear
  • 🌍 Build your personal brand in the Microsoft 365 space

This isn’t just a podcast — it’s a platform for people who take action.

🔥 Most people wait. The best ones don’t.

👉 Connect with me on LinkedIn and send me a message:
"I want in"

Let’s build something awesome 👊

1
00:00:00,000 --> 00:00:02,360
Last week, your copilot answered a question perfectly.

2
00:00:02,360 --> 00:00:04,200
Someone asked about a project timeline,

3
00:00:04,200 --> 00:00:06,440
and it pulled the right details connected the right dots

4
00:00:06,440 --> 00:00:08,320
gave you an answer that felt almost human.

5
00:00:08,320 --> 00:00:10,160
Today, someone asks the same question,

6
00:00:10,160 --> 00:00:11,880
slightly different wording.

7
00:00:11,880 --> 00:00:14,960
And copilot fumbles the exact connection it made yesterday.

8
00:00:14,960 --> 00:00:17,600
Here's the assumption everyone walks around with.

9
00:00:17,600 --> 00:00:19,600
Copilot remembers your organization.

10
00:00:19,600 --> 00:00:20,120
It doesn't.

11
00:00:20,120 --> 00:00:20,920
It searches.

12
00:00:20,920 --> 00:00:22,920
Every single time, from zero, that's the tension

13
00:00:22,920 --> 00:00:23,920
nobody talks about.

14
00:00:23,920 --> 00:00:25,760
Retrieval and memory are not the same thing.

15
00:00:25,760 --> 00:00:28,800
One finds information, the other builds on what it already knows.

16
00:00:28,800 --> 00:00:31,160
Microsoft built copilot on the first one

17
00:00:31,160 --> 00:00:33,040
and let the marketing imply the second.

18
00:00:33,040 --> 00:00:35,240
There's a pattern circulating right now.

19
00:00:35,240 --> 00:00:37,160
Born out of AI research circles,

20
00:00:37,160 --> 00:00:39,200
that treats knowledge as something you compile,

21
00:00:39,200 --> 00:00:41,160
not something you search for over and over.

22
00:00:41,160 --> 00:00:42,560
We'll name it in a few minutes.

23
00:00:42,560 --> 00:00:44,360
By the end of this episode, you'll understand

24
00:00:44,360 --> 00:00:46,560
why your copilot deployment feels sharp in demos

25
00:00:46,560 --> 00:00:48,600
and shallow in production and what structural change

26
00:00:48,600 --> 00:00:49,880
actually fixes it.

27
00:00:49,880 --> 00:00:51,240
Quick thing before we get into it.

28
00:00:51,240 --> 00:00:53,240
If you're into this channel because you want to understand

29
00:00:53,240 --> 00:00:55,480
the systems behind Microsoft 365,

30
00:00:55,480 --> 00:00:59,480
not just the button clicking, hit subscribe on M365FM podcast.

31
00:00:59,480 --> 00:01:02,040
This episode is exactly the kind of thing we do here.

32
00:01:02,040 --> 00:01:03,720
And I want to set the expectation right now.

33
00:01:03,720 --> 00:01:05,560
This isn't a copilot feature tour.

34
00:01:05,560 --> 00:01:07,360
We're not walking through settings menus.

35
00:01:07,360 --> 00:01:08,840
This is a structural diagnosis,

36
00:01:08,840 --> 00:01:11,320
the kind you do on a system that keeps almost working

37
00:01:11,320 --> 00:01:12,800
but never quite getting there.

38
00:01:12,800 --> 00:01:14,040
So if you came for prompt tips,

39
00:01:14,040 --> 00:01:15,480
this one's going to disappoint you.

40
00:01:15,480 --> 00:01:18,400
We're going under the interface into the architecture.

41
00:01:18,400 --> 00:01:20,000
That's where the real answer lives.

42
00:01:20,000 --> 00:01:22,840
The day one copilot promise.

43
00:01:22,840 --> 00:01:25,000
Think back to how copilot got pitched to you

44
00:01:25,000 --> 00:01:26,920
or to your leadership or in whatever deck

45
00:01:26,920 --> 00:01:28,800
convinced someone to write the check.

46
00:01:28,800 --> 00:01:30,360
Ask questions in plain English.

47
00:01:30,360 --> 00:01:33,040
Get answers grounded in your organization's own data.

48
00:01:33,040 --> 00:01:34,480
No more digging through SharePoint.

49
00:01:34,480 --> 00:01:37,120
No more pinging three people on teams to find one document.

50
00:01:37,120 --> 00:01:38,720
Just ask and copilot knows.

51
00:01:38,720 --> 00:01:40,760
Microsoft backed that pitch with numbers.

52
00:01:40,760 --> 00:01:42,960
11 minutes saved per day per user.

53
00:01:42,960 --> 00:01:46,080
Do the math across a year and that's close to a full work week

54
00:01:46,080 --> 00:01:47,400
handed back to every employee.

55
00:01:47,400 --> 00:01:49,400
That's a real number and for a lot of organizations

56
00:01:49,400 --> 00:01:51,720
it's the number that justified the license cost.

57
00:01:51,720 --> 00:01:53,960
But here's what nobody set out loud during that pitch.

58
00:01:53,960 --> 00:01:56,600
Baked into asked questions and get grounded answers

59
00:01:56,600 --> 00:01:58,440
is an implicit promise.

60
00:01:58,440 --> 00:02:00,560
This thing gets better as you use it.

61
00:02:00,560 --> 00:02:03,280
You feed it more meetings, more documents, more decisions

62
00:02:03,280 --> 00:02:06,120
and it becomes sharper, more attuned to how your organization

63
00:02:06,120 --> 00:02:06,840
actually works.

64
00:02:06,840 --> 00:02:09,200
That's the story everyone heard, even if nobody wrote it

65
00:02:09,200 --> 00:02:10,000
on a slide.

66
00:02:10,000 --> 00:02:11,800
What actually ships is something different.

67
00:02:11,800 --> 00:02:14,280
It's a retrieval system wrapped in a chat interface.

68
00:02:14,280 --> 00:02:16,280
Dressed up just enough to feel like it's learning.

69
00:02:16,280 --> 00:02:17,080
It isn't.

70
00:02:17,080 --> 00:02:19,840
Every time someone types a question, copilot goes and looks,

71
00:02:19,840 --> 00:02:20,760
it doesn't recall.

72
00:02:20,760 --> 00:02:22,720
It searches, finds what looks relevant

73
00:02:22,720 --> 00:02:24,360
and generates a response from that.

74
00:02:24,360 --> 00:02:26,240
Ask again tomorrow and it searches again

75
00:02:26,240 --> 00:02:28,000
as if yesterday never happened.

76
00:02:28,000 --> 00:02:30,560
Now you might be thinking this sounds like a UX complaint,

77
00:02:30,560 --> 00:02:31,720
a minor annoyance.

78
00:02:31,720 --> 00:02:32,560
It's not.

79
00:02:32,560 --> 00:02:34,320
This is a governance and adoption problem

80
00:02:34,320 --> 00:02:36,360
and those are expensive in a completely different way

81
00:02:36,360 --> 00:02:37,800
than a clunky interface.

82
00:02:37,800 --> 00:02:39,800
When an assistant gives inconsistent answers

83
00:02:39,800 --> 00:02:41,480
to the same underlying question,

84
00:02:41,480 --> 00:02:43,440
people stop trusting it for anything that matters.

85
00:02:43,440 --> 00:02:44,760
They keep it around for quick summaries

86
00:02:44,760 --> 00:02:46,160
and calendar questions, sure.

87
00:02:46,160 --> 00:02:48,520
But the moment a decision actually rides on the answer,

88
00:02:48,520 --> 00:02:49,720
they go back to asking a person.

89
00:02:49,720 --> 00:02:50,680
That's not a bug report.

90
00:02:50,680 --> 00:02:52,520
That's an organization quietly deciding

91
00:02:52,520 --> 00:02:54,800
the tool isn't reliable enough to lean on.

92
00:02:54,800 --> 00:02:56,200
And once that decision gets made,

93
00:02:56,200 --> 00:02:58,000
even informally, adoption stores.

94
00:02:58,000 --> 00:02:59,640
Not because people didn't try copilot

95
00:02:59,640 --> 00:03:01,680
because they tried it, got burned once or twice

96
00:03:01,680 --> 00:03:03,040
by an inconsistent answer

97
00:03:03,040 --> 00:03:04,560
and adjusted their behavior around it.

98
00:03:04,560 --> 00:03:05,840
Nobody files a ticket for that.

99
00:03:05,840 --> 00:03:08,240
It just shows up months later as flat usage numbers

100
00:03:08,240 --> 00:03:11,000
and a leadership team asking why the tool they paid for

101
00:03:11,000 --> 00:03:13,200
isn't delivering the return they were promised.

102
00:03:13,200 --> 00:03:16,120
So the gap between the pitch and the product isn't cosmetic.

103
00:03:16,120 --> 00:03:18,840
It's the difference between a tool that compounds in value

104
00:03:18,840 --> 00:03:20,760
and one that just sits there, static,

105
00:03:20,760 --> 00:03:22,400
no matter how long you've had it deployed,

106
00:03:22,400 --> 00:03:24,160
to actually see where that gap comes from.

107
00:03:24,160 --> 00:03:26,280
You have to look at what's happening mechanically

108
00:03:26,280 --> 00:03:28,840
every single time someone types a question into copilot

109
00:03:28,840 --> 00:03:31,160
because the behavior we just described isn't random.

110
00:03:31,160 --> 00:03:32,880
It's the direct, predictable result

111
00:03:32,880 --> 00:03:35,720
of how the system is built underneath the chat window.

112
00:03:35,720 --> 00:03:37,000
What drag actually does?

113
00:03:37,000 --> 00:03:38,200
So let's name it directly.

114
00:03:38,200 --> 00:03:40,320
Our Ag stands for retrieval augmented generation.

115
00:03:40,320 --> 00:03:42,640
Strip the jargon and here's what that actually means.

116
00:03:42,640 --> 00:03:44,720
The system retrieves some information,

117
00:03:44,720 --> 00:03:47,320
then uses that information to help generate an answer.

118
00:03:47,320 --> 00:03:48,920
That's the whole concept.

119
00:03:48,920 --> 00:03:51,480
Retrieval, then generation, glued together.

120
00:03:51,480 --> 00:03:53,400
Here's how it works mechanically, step by step.

121
00:03:53,400 --> 00:03:54,800
First, your documents get chunked.

122
00:03:54,800 --> 00:03:57,720
Every SharePoint file, every Teams transcript,

123
00:03:57,720 --> 00:04:00,400
every policy PDF gets sliced into smaller pieces,

124
00:04:00,400 --> 00:04:02,240
usually a few hundred words at a time.

125
00:04:02,240 --> 00:04:04,000
Second, each of those chunks gets converted

126
00:04:04,000 --> 00:04:05,440
into something called an embedding,

127
00:04:05,440 --> 00:04:07,240
which is just a mathematical fingerprint

128
00:04:07,240 --> 00:04:08,600
of what that chunk means.

129
00:04:08,600 --> 00:04:10,400
Third, when someone types a question,

130
00:04:10,400 --> 00:04:12,640
that question also gets turned into a fingerprint

131
00:04:12,640 --> 00:04:14,600
and the system searches for chunks

132
00:04:14,600 --> 00:04:16,160
whose fingerprints look similar.

133
00:04:16,160 --> 00:04:18,440
Fourth, whichever chunk score highest gets stuffed

134
00:04:18,440 --> 00:04:20,880
into a prompt alongside the original question.

135
00:04:20,880 --> 00:04:22,920
Fifth, the language model reads all of that

136
00:04:22,920 --> 00:04:23,920
and generates an answer.

137
00:04:23,920 --> 00:04:25,280
Now pay attention to that word chunk

138
00:04:25,280 --> 00:04:27,960
because it's doing more work than people give it credit for.

139
00:04:27,960 --> 00:04:29,720
Documents get sliced into fragments

140
00:04:29,720 --> 00:04:31,560
with no awareness of the whole.

141
00:04:31,560 --> 00:04:34,400
A 40-page project plan doesn't get read as a project plan.

142
00:04:34,400 --> 00:04:37,200
It gets cut into 15 or 20 disconnected pieces.

143
00:04:37,200 --> 00:04:39,360
Each one judged purely on whether it resembles

144
00:04:39,360 --> 00:04:40,800
the question being asked.

145
00:04:40,800 --> 00:04:43,000
The system never holds the full document in mind.

146
00:04:43,000 --> 00:04:45,120
It holds fragments and it bets that the right fragments

147
00:04:45,120 --> 00:04:47,680
to stick together will look enough like a coherent answer.

148
00:04:47,680 --> 00:04:49,680
This is a search engine with a language model

149
00:04:49,680 --> 00:04:50,760
bolted on the end.

150
00:04:50,760 --> 00:04:53,440
It's not a thinking system and it was never built to be one.

151
00:04:53,440 --> 00:04:55,240
The intelligence you're seeing when co-pilot

152
00:04:55,240 --> 00:04:58,080
gives you a good answer isn't co-pilot understanding

153
00:04:58,080 --> 00:04:59,120
your organization.

154
00:04:59,120 --> 00:05:01,560
It's a language model doing a genuinely impressive job

155
00:05:01,560 --> 00:05:04,000
of stitching together whatever fragments got handed to it.

156
00:05:04,000 --> 00:05:05,120
That's a real skill.

157
00:05:05,120 --> 00:05:08,080
It's just not memory and it's not comprehension of the whole.

158
00:05:08,080 --> 00:05:09,640
If you want a simple way to picture this,

159
00:05:09,640 --> 00:05:12,280
think about typing a question into SharePoint Search.

160
00:05:12,280 --> 00:05:14,880
You type in a phrase, SharePoint goes and finds documents

161
00:05:14,880 --> 00:05:16,800
that contain similar words or concepts

162
00:05:16,800 --> 00:05:18,360
and it hands you a list of results.

163
00:05:18,360 --> 00:05:19,840
Ragn does almost exactly that.

164
00:05:19,840 --> 00:05:21,560
Except instead of handing you the list,

165
00:05:21,560 --> 00:05:24,360
it reads the top few results and writes you a paragraph summarizing

166
00:05:24,360 --> 00:05:24,760
them.

167
00:05:24,760 --> 00:05:27,880
Same search, same fragments, just a nicer wrapper on the output.

168
00:05:27,880 --> 00:05:29,960
And look, for a single look up, this works fine.

169
00:05:29,960 --> 00:05:31,000
Genuinely fine.

170
00:05:31,000 --> 00:05:33,640
If someone asks, what's our current PTO policy?

171
00:05:33,640 --> 00:05:35,800
Ragn finds the chunk that talks about PTO,

172
00:05:35,800 --> 00:05:37,520
generates a clean summary, and everyone

173
00:05:37,520 --> 00:05:38,480
moves on with their day.

174
00:05:38,480 --> 00:05:42,760
One question, one search, one answer, no complaints.

175
00:05:42,760 --> 00:05:44,440
The failure doesn't show up there.

176
00:05:44,440 --> 00:05:46,240
It shows up the moment you need continuity.

177
00:05:46,240 --> 00:05:47,960
The moment a question depends on something

178
00:05:47,960 --> 00:05:50,000
the system already figured out five minutes ago

179
00:05:50,000 --> 00:05:51,840
or yesterday or last month.

180
00:05:51,840 --> 00:05:54,880
Because Ragn has no concept of already figured out.

181
00:05:54,880 --> 00:05:56,400
It only has search again.

182
00:05:56,400 --> 00:05:59,160
And that's exactly where the cracks start to show.

183
00:05:59,160 --> 00:06:00,520
The stateless trap.

184
00:06:00,520 --> 00:06:02,320
Let's put a real word on what we just described

185
00:06:02,320 --> 00:06:03,400
because it matters.

186
00:06:03,400 --> 00:06:05,440
That behavior is called statelessness.

187
00:06:05,440 --> 00:06:07,760
Every query starts from zero, no matter how many times

188
00:06:07,760 --> 00:06:09,080
you've asked something related before,

189
00:06:09,080 --> 00:06:11,080
no matter how obviously connected today's question

190
00:06:11,080 --> 00:06:12,400
is to yesterday's.

191
00:06:12,400 --> 00:06:14,400
The system has no concept of before.

192
00:06:14,400 --> 00:06:16,360
There's only this query right now treated

193
00:06:16,360 --> 00:06:18,040
like the first one it's ever seen.

194
00:06:18,040 --> 00:06:20,720
Walk through what that actually looks like inside co-pilot.

195
00:06:20,720 --> 00:06:22,440
Someone on a project team asks,

196
00:06:22,440 --> 00:06:24,960
what's the status of the client onboarding for Meridian?

197
00:06:24,960 --> 00:06:27,160
Co-pilot searches, finds the right team's thread,

198
00:06:27,160 --> 00:06:30,080
the right sharepoint update, maybe a status email from last week

199
00:06:30,080 --> 00:06:31,600
and gives a genuinely good answer.

200
00:06:31,600 --> 00:06:33,480
Clean, accurate, useful.

201
00:06:33,480 --> 00:06:35,960
The person moves on with their day, impressed.

202
00:06:35,960 --> 00:06:38,240
Tomorrow, that same person asks a follow-up, something

203
00:06:38,240 --> 00:06:41,240
like, what's blocking us from finishing that onboarding?

204
00:06:41,240 --> 00:06:43,320
Notice what just happened in that sentence.

205
00:06:43,320 --> 00:06:44,960
That onboarding.

206
00:06:44,960 --> 00:06:46,360
A human would hear this and instantly

207
00:06:46,360 --> 00:06:48,040
connected to yesterday's conversation,

208
00:06:48,040 --> 00:06:50,320
no clarification needed, co-pilot doesn't.

209
00:06:50,320 --> 00:06:52,800
It has no idea that onboarding refers to Meridian

210
00:06:52,800 --> 00:06:55,440
because it never stored the fact that Meridian came up yesterday.

211
00:06:55,440 --> 00:06:58,800
So it searches again from scratch, hoping the word onboarding alone

212
00:06:58,800 --> 00:07:00,640
points it back to the right documents.

213
00:07:00,640 --> 00:07:02,480
Sometimes it gets lucky, sometimes it grabs

214
00:07:02,480 --> 00:07:04,680
the different clients onboarding thread entirely

215
00:07:04,680 --> 00:07:06,800
and now you've got a confidently wrong answer

216
00:07:06,800 --> 00:07:09,360
standing right next to yesterday's confidently right one.

217
00:07:09,360 --> 00:07:10,960
Here's the part that actually costs you.

218
00:07:10,960 --> 00:07:11,920
Nothing compounds.

219
00:07:11,920 --> 00:07:14,480
There's no accumulated understanding of the Meridian project

220
00:07:14,480 --> 00:07:16,600
building up over these two conversations,

221
00:07:16,600 --> 00:07:19,520
no growing sense of who's involved, what decisions got made,

222
00:07:19,520 --> 00:07:21,560
what the blockers have historically been.

223
00:07:21,560 --> 00:07:23,080
Each question is an island.

224
00:07:23,080 --> 00:07:24,960
You could ask co-pilot about the same project

225
00:07:24,960 --> 00:07:27,960
50 times over three months and it would know exactly

226
00:07:27,960 --> 00:07:31,080
as much on question 50 as it did on question one

227
00:07:31,080 --> 00:07:34,160
because it never carried anything forward between them.

228
00:07:34,160 --> 00:07:36,880
Now compare that to how an actual colleague would handle it,

229
00:07:36,880 --> 00:07:39,560
say the person who handled Meridian's onboarding last month,

230
00:07:39,560 --> 00:07:42,440
the answer to what's blocking us comes wrapped in memory.

231
00:07:42,440 --> 00:07:43,760
They remember the client pushed back

232
00:07:43,760 --> 00:07:45,920
on the integration timeline two weeks ago.

233
00:07:45,920 --> 00:07:47,640
They remember a decision got made in a meeting

234
00:07:47,640 --> 00:07:48,920
to reprioritize.

235
00:07:48,920 --> 00:07:50,520
They don't need to redrive any of that

236
00:07:50,520 --> 00:07:52,240
because they were there and they kept it.

237
00:07:52,240 --> 00:07:55,280
Every follow-up question they answer builds on the last one,

238
00:07:55,280 --> 00:07:57,760
gets sharper, gets more useful over time.

239
00:07:57,760 --> 00:07:59,800
That's what a second conversation with a person actually

240
00:07:59,800 --> 00:08:02,280
feels like and it's exactly what a second conversation

241
00:08:02,280 --> 00:08:03,440
with co-pilot doesn't.

242
00:08:03,440 --> 00:08:05,240
So name the real cost here plainly

243
00:08:05,240 --> 00:08:06,960
because it's not what people assume.

244
00:08:06,960 --> 00:08:09,240
The cost isn't wrong answers, not usually.

245
00:08:09,240 --> 00:08:10,720
Rags often technically accurate

246
00:08:10,720 --> 00:08:12,520
on any single question you throw at it.

247
00:08:12,520 --> 00:08:15,120
The cost is shallow answers, ones that never deep in no matter

248
00:08:15,120 --> 00:08:16,800
how long you've been using the tool,

249
00:08:16,800 --> 00:08:20,000
no matter how much history exists between you and the system.

250
00:08:20,000 --> 00:08:21,560
You get competence without growth,

251
00:08:21,560 --> 00:08:24,240
fluency without accumulation.

252
00:08:24,240 --> 00:08:26,880
And here's the thing worth sitting with before we go further.

253
00:08:26,880 --> 00:08:29,120
This isn't a training problem, it's not a prompt problem

254
00:08:29,120 --> 00:08:31,240
and it's not a rollout problem you can fix

255
00:08:31,240 --> 00:08:32,640
with better change management.

256
00:08:32,640 --> 00:08:34,960
It's what happens structurally every single time

257
00:08:34,960 --> 00:08:37,680
you take a search mechanism and ask it to do memory's job.

258
00:08:37,680 --> 00:08:40,480
Search was never built to remember, it was built to find.

259
00:08:40,480 --> 00:08:42,080
And no amount of clever phrasing changes

260
00:08:42,080 --> 00:08:44,960
what the system was designed to do underneath.

261
00:08:44,960 --> 00:08:46,920
Why this isn't a prompting problem?

262
00:08:46,920 --> 00:08:50,120
At this point, someone in IT is going to have a reflex reaction

263
00:08:50,120 --> 00:08:51,600
to everything we just said.

264
00:08:51,600 --> 00:08:53,040
They're going to blame the prompts.

265
00:08:53,040 --> 00:08:55,880
Maybe the rollout was rushed, maybe users just don't know how

266
00:08:55,880 --> 00:08:57,400
to ask questions the right way yet

267
00:08:57,400 --> 00:08:59,640
and a few training sessions will close the gap.

268
00:08:59,640 --> 00:09:00,720
That instinct makes sense.

269
00:09:00,720 --> 00:09:02,280
It's usually where the blame lands first

270
00:09:02,280 --> 00:09:03,640
because it's the easiest thing to fix

271
00:09:03,640 --> 00:09:05,520
without touching the underlying system.

272
00:09:05,520 --> 00:09:08,120
Rewrite the prompt library, run a lunch and learn,

273
00:09:08,120 --> 00:09:09,640
tell people to be more specific

274
00:09:09,640 --> 00:09:12,160
and show a better phrasing helps at the margins.

275
00:09:12,160 --> 00:09:13,640
But it doesn't touch the actual problem

276
00:09:13,640 --> 00:09:16,480
and there's a number that makes this hard to ignore.

277
00:09:16,480 --> 00:09:18,000
Research on structured knowledge systems

278
00:09:18,000 --> 00:09:20,840
shows organizations using layered knowledge with LLMs

279
00:09:20,840 --> 00:09:23,080
cut AI error rates by more than 60%

280
00:09:23,080 --> 00:09:25,280
compared to standard rag setups, 60%.

281
00:09:25,280 --> 00:09:26,600
That's not a rounding error

282
00:09:26,600 --> 00:09:28,200
and it's not the kind of gap you close

283
00:09:28,200 --> 00:09:30,320
by teaching people to phrase their questions better.

284
00:09:30,320 --> 00:09:33,000
Think about what that number actually implies.

285
00:09:33,000 --> 00:09:35,240
It's not measuring how well people ask questions.

286
00:09:35,240 --> 00:09:37,240
It's measuring what the system does before a question

287
00:09:37,240 --> 00:09:38,320
ever gets typed.

288
00:09:38,320 --> 00:09:41,360
60% fewer errors didn't come from smarter users.

289
00:09:41,360 --> 00:09:43,440
It came from a completely different relationship

290
00:09:43,440 --> 00:09:45,800
between the AI and the knowledge underneath it,

291
00:09:45,800 --> 00:09:47,600
one where the knowledge had already been organized,

292
00:09:47,600 --> 00:09:50,800
connected and reconciled before anyone opened a chat window.

293
00:09:50,800 --> 00:09:52,920
So here's the reframe and it's worth sitting with

294
00:09:52,920 --> 00:09:54,920
because it changes where you point your effort.

295
00:09:54,920 --> 00:09:56,920
You can't prompt your way out of an architecture

296
00:09:56,920 --> 00:09:59,280
that has no place to store what it learns.

297
00:09:59,280 --> 00:10:01,760
It doesn't matter how carefully worded the question is.

298
00:10:01,760 --> 00:10:03,800
If the system searches fresh every time,

299
00:10:03,800 --> 00:10:06,000
then even a perfect prompt just retrieves a fresh set

300
00:10:06,000 --> 00:10:07,920
of chunks and generates a fresh answer,

301
00:10:07,920 --> 00:10:10,240
disconnected from anything that came before.

302
00:10:10,240 --> 00:10:12,360
A better prompt gets you a better single search.

303
00:10:12,360 --> 00:10:13,880
It doesn't get you memory because memory

304
00:10:13,880 --> 00:10:15,400
was never something a prompt controls.

305
00:10:15,400 --> 00:10:17,200
That's the part worth naming plainly.

306
00:10:17,200 --> 00:10:18,680
The fix doesn't live in the interface.

307
00:10:18,680 --> 00:10:20,400
It doesn't live in how questions get asked.

308
00:10:20,400 --> 00:10:23,080
It lives below all of that in how knowledge gets prepared

309
00:10:23,080 --> 00:10:25,440
and represented before Copilot ever touches it.

310
00:10:25,440 --> 00:10:27,240
If the knowledge sitting behind the assistant

311
00:10:27,240 --> 00:10:29,960
is a pile of raw documents waiting to be chunked on demand,

312
00:10:29,960 --> 00:10:31,640
no amount of prompt engineering changes

313
00:10:31,640 --> 00:10:32,640
what happens next.

314
00:10:32,640 --> 00:10:35,720
The system is built to search and search is what it will keep doing

315
00:10:35,720 --> 00:10:37,600
no matter how the question gets phrased,

316
00:10:37,600 --> 00:10:38,960
which raises the actual question worth

317
00:10:38,960 --> 00:10:40,360
spending the rest of this episode on.

318
00:10:40,360 --> 00:10:42,720
Not how do we ask Copilot better questions.

319
00:10:42,720 --> 00:10:44,840
But what does the system actually look like

320
00:10:44,840 --> 00:10:46,960
when it's built for memory instead of search

321
00:10:46,960 --> 00:10:48,960
because that system exists and it works

322
00:10:48,960 --> 00:10:50,240
on a completely different principle

323
00:10:50,240 --> 00:10:52,560
than the one copilot ships with today?

324
00:10:52,560 --> 00:10:53,880
Enter the LLM Wiki.

325
00:10:53,880 --> 00:10:55,040
Here's that system.

326
00:10:55,040 --> 00:10:56,400
It didn't come out of Redmond

327
00:10:56,400 --> 00:10:58,120
and it didn't come out of a product road map.

328
00:10:58,120 --> 00:11:00,280
It surfaced from AI research circles

329
00:11:00,280 --> 00:11:02,840
and the person most associated with articulating it clearly

330
00:11:02,840 --> 00:11:05,680
is Andre Kapati, one of the more influential voices

331
00:11:05,680 --> 00:11:07,120
in how people think about working

332
00:11:07,120 --> 00:11:08,600
with these models day to day.

333
00:11:08,600 --> 00:11:10,480
He wasn't building an enterprise product.

334
00:11:10,480 --> 00:11:12,160
He was solving his own problem,

335
00:11:12,160 --> 00:11:14,800
building a personal knowledge base for research he cared about

336
00:11:14,800 --> 00:11:17,520
and the pattern he landed on is the one worth stealing.

337
00:11:17,520 --> 00:11:18,680
The core move is simple to say

338
00:11:18,680 --> 00:11:20,160
and genuinely different in practice.

339
00:11:20,160 --> 00:11:21,560
Instead of searching raw documents

340
00:11:21,560 --> 00:11:23,440
every single time a question comes in

341
00:11:23,440 --> 00:11:25,320
and AI reads the sources once

342
00:11:25,320 --> 00:11:28,360
and builds a structured interlinked knowledge base out of them.

343
00:11:28,360 --> 00:11:29,960
Once, not once per question.

344
00:11:29,960 --> 00:11:32,320
Once, period, then it keeps that structure around

345
00:11:32,320 --> 00:11:33,320
and refers back to it.

346
00:11:33,320 --> 00:11:35,360
Picture three layers because this is where it stops

347
00:11:35,360 --> 00:11:37,160
being an abstract idea and starts

348
00:11:37,160 --> 00:11:38,760
being something you could actually build.

349
00:11:38,760 --> 00:11:40,240
Layer one is raw sources.

350
00:11:40,240 --> 00:11:42,400
These are your original documents whatever they are

351
00:11:42,400 --> 00:11:43,640
and they stay read only.

352
00:11:43,640 --> 00:11:45,400
Nobody touches them, nothing re-rides them.

353
00:11:45,400 --> 00:11:47,360
They're ground truth sitting there untouched

354
00:11:47,360 --> 00:11:49,160
the same way a source document should.

355
00:11:49,160 --> 00:11:50,560
Layer two is the Wiki itself.

356
00:11:50,560 --> 00:11:51,960
This is where the actual work happens.

357
00:11:51,960 --> 00:11:53,680
It's a set of pages, the AI writes.

358
00:11:53,680 --> 00:11:55,480
One per concept, one per entity,

359
00:11:55,480 --> 00:11:57,560
one per recurring topic, all cross linked

360
00:11:57,560 --> 00:11:59,200
to each other the way a real Wiki works.

361
00:11:59,200 --> 00:12:00,440
Not a database of chunks.

362
00:12:00,440 --> 00:12:02,800
Pages written in something like plain language

363
00:12:02,800 --> 00:12:04,280
meant to be read as a whole.

364
00:12:04,280 --> 00:12:05,680
Layer three is the schema.

365
00:12:05,680 --> 00:12:07,080
Think of it as the rule book.

366
00:12:07,080 --> 00:12:09,440
It tells the AI how to structure new pages,

367
00:12:09,440 --> 00:12:11,080
how to decide when something updates

368
00:12:11,080 --> 00:12:13,520
an existing page versus when it deserves a new one,

369
00:12:13,520 --> 00:12:15,480
how to format things consistently

370
00:12:15,480 --> 00:12:17,840
so the Wiki doesn't turn into chaos as it grows.

371
00:12:17,840 --> 00:12:21,680
Now here's the part people get wrong on first hearing this.

372
00:12:21,680 --> 00:12:23,960
They assume the Wiki is just static storage,

373
00:12:23,960 --> 00:12:26,640
a folder where documents get copied and organized once.

374
00:12:26,640 --> 00:12:28,160
It's not, it's maintained.

375
00:12:28,160 --> 00:12:30,680
Every time a new source comes in, the AI reads it

376
00:12:30,680 --> 00:12:32,000
and does real work with it.

377
00:12:32,000 --> 00:12:34,600
It updates existing pages if the new source adds detail

378
00:12:34,600 --> 00:12:36,040
to something already compiled.

379
00:12:36,040 --> 00:12:38,280
It creates new pages if the source introduces

380
00:12:38,280 --> 00:12:41,280
a genuinely new concept that didn't exist in the Wiki yet.

381
00:12:41,280 --> 00:12:43,320
And if the new source says something that contradicts

382
00:12:43,320 --> 00:12:46,240
what's already there, it doesn't just override quietly.

383
00:12:46,240 --> 00:12:48,120
It flags it, that flag is the whole point

384
00:12:48,120 --> 00:12:49,920
and we'll come back to why in a bit.

385
00:12:49,920 --> 00:12:52,440
So say the line plainly, because it's the whole argument

386
00:12:52,440 --> 00:12:53,440
in one sentence.

387
00:12:53,440 --> 00:12:55,480
Knowledge gets compiled once then kept current.

388
00:12:55,480 --> 00:12:57,440
That's the exact opposite of what Ragn does.

389
00:12:57,440 --> 00:13:00,560
Ragn rebuilds understanding from scratch on every query.

390
00:13:00,560 --> 00:13:02,840
A Wiki builds it once and then just maintains it,

391
00:13:02,840 --> 00:13:04,800
the way you'd maintain anything that's actually alive

392
00:13:04,800 --> 00:13:07,680
and growing instead of something you keep rebuilding from nothing.

393
00:13:07,680 --> 00:13:10,280
To see why this distinction actually matters for co-pilot

394
00:13:10,280 --> 00:13:13,120
and not just as a nice idea for personal research projects,

395
00:13:13,120 --> 00:13:15,160
you need to put the two systems side by side

396
00:13:15,160 --> 00:13:17,800
and watch where the actual work happens in each one.

397
00:13:17,800 --> 00:13:20,040
Ragn vs LLM Wiki side by side.

398
00:13:20,040 --> 00:13:22,720
Put the two systems next to each other and ask one question,

399
00:13:22,720 --> 00:13:24,480
where does the thinking actually happen?

400
00:13:24,480 --> 00:13:26,280
Not the answering, the thinking.

401
00:13:26,280 --> 00:13:28,640
The part where raw information turns into something

402
00:13:28,640 --> 00:13:30,240
organized enough to be useful.

403
00:13:30,240 --> 00:13:32,160
That question tells you almost everything

404
00:13:32,160 --> 00:13:34,240
you need to know about why one system compounds

405
00:13:34,240 --> 00:13:35,320
and the other doesn't.

406
00:13:35,320 --> 00:13:37,440
In Ragn, the thinking happens at query time.

407
00:13:37,440 --> 00:13:39,440
Every single question triggers the whole process

408
00:13:39,440 --> 00:13:40,360
from the beginning.

409
00:13:40,360 --> 00:13:42,120
Fresh search across the embeddings,

410
00:13:42,120 --> 00:13:44,880
fresh chunking logic deciding which fragments matter.

411
00:13:44,880 --> 00:13:47,840
Fresh synthesis, where the model reads whatever got pulled

412
00:13:47,840 --> 00:13:49,920
and stitches it into an answer on the spot.

413
00:13:49,920 --> 00:13:52,600
None of that work exists until the question shows up.

414
00:13:52,600 --> 00:13:55,240
The system has zero understanding sitting in reserve.

415
00:13:55,240 --> 00:13:58,080
It manufactures understanding, on demand every time

416
00:13:58,080 --> 00:14:00,880
and then throws it away the moment the answer gets delivered.

417
00:14:00,880 --> 00:14:04,000
In the LLM Wiki, the thinking happens at ingestion time.

418
00:14:04,000 --> 00:14:05,520
That's the entire structural difference

419
00:14:05,520 --> 00:14:07,560
and it's easy to miss because it sounds small.

420
00:14:07,560 --> 00:14:10,840
The synthesis is already done before anyone asks the question.

421
00:14:10,840 --> 00:14:14,520
When a new source comes in, the AI does the hard part right then.

422
00:14:14,520 --> 00:14:16,560
Reading it, deciding what it means,

423
00:14:16,560 --> 00:14:18,880
connecting it to what's already compiled,

424
00:14:18,880 --> 00:14:21,240
updating pages, flagging conflicts.

425
00:14:21,240 --> 00:14:24,320
By the time someone actually types a question into the system,

426
00:14:24,320 --> 00:14:25,600
there's nothing left to figure out.

427
00:14:25,600 --> 00:14:27,600
The system isn't generating understanding.

428
00:14:27,600 --> 00:14:30,000
It's retrieving understanding that already exists,

429
00:14:30,000 --> 00:14:31,920
fully formed, sitting on a page.

430
00:14:31,920 --> 00:14:34,680
That difference shows up in a number worth paying attention to.

431
00:14:34,680 --> 00:14:37,400
For focused knowledge basis, LLM Wiki can cut token usage

432
00:14:37,400 --> 00:14:40,480
by roughly 95% compared to naive rag document loading.

433
00:14:40,480 --> 00:14:42,960
95% that's not a marginal efficiency gain

434
00:14:42,960 --> 00:14:44,760
that's a different order of magnitude.

435
00:14:44,760 --> 00:14:46,960
Here's why that number exists and it's not magic.

436
00:14:46,960 --> 00:14:49,920
Rag spends tokens researching and restitching every single time.

437
00:14:49,920 --> 00:14:52,400
It has to pull multiple chunks, feed all of them

438
00:14:52,400 --> 00:14:54,520
into the prompt alongside the question,

439
00:14:54,520 --> 00:14:57,480
and let the model do the work of reconciling fragments

440
00:14:57,480 --> 00:15:00,040
that were never written to sit next to each other.

441
00:15:00,040 --> 00:15:02,800
All of that costs tokens every query, forever.

442
00:15:02,800 --> 00:15:04,120
A Wiki skips that entirely.

443
00:15:04,120 --> 00:15:05,960
You're not researching and restitching.

444
00:15:05,960 --> 00:15:08,080
You're reading a page that was already organized,

445
00:15:08,080 --> 00:15:10,880
already coherent, already written, to be read as a whole.

446
00:15:10,880 --> 00:15:13,640
Less raw material has to get shoved into the model's context window

447
00:15:13,640 --> 00:15:15,880
because someone or something already

448
00:15:15,880 --> 00:15:17,720
did the organizing work in advance.

449
00:15:17,720 --> 00:15:19,960
Now, let's be honest about where this doesn't apply

450
00:15:19,960 --> 00:15:22,720
because overselling it here would undercut the whole argument.

451
00:15:22,720 --> 00:15:26,080
Rag stays the default for large, relatively static repositories

452
00:15:26,080 --> 00:15:27,240
and that's completely fine.

453
00:15:27,240 --> 00:15:29,440
If you're sitting on hundreds of thousands of documents

454
00:15:29,440 --> 00:15:31,600
that don't change much and get searched broadly

455
00:15:31,600 --> 00:15:33,240
across a huge range of topics,

456
00:15:33,240 --> 00:15:37,040
Rag's search first design is doing exactly what it's built for.

457
00:15:37,040 --> 00:15:38,920
The Wiki pattern isn't trying to replace that.

458
00:15:38,920 --> 00:15:40,560
Where the Wiki pattern actually shines

459
00:15:40,560 --> 00:15:44,280
is below roughly 50 to 100,000 tokens of compiled knowledge.

460
00:15:44,280 --> 00:15:45,760
That's a bounded focused domain.

461
00:15:45,760 --> 00:15:47,840
Not everything your organization has ever produced,

462
00:15:47,840 --> 00:15:50,760
a specific slice of it, the stuff people ask about again and again

463
00:15:50,760 --> 00:15:52,320
compiled once in Keptrap.

464
00:15:52,320 --> 00:15:54,400
None of this stays theoretical for long though

465
00:15:54,400 --> 00:15:57,000
because there's a version of exactly this tension

466
00:15:57,000 --> 00:15:59,720
already sitting inside Microsoft 365 right now,

467
00:15:59,720 --> 00:16:01,680
quietly shaping how co-pilot behaves

468
00:16:01,680 --> 00:16:03,760
every time someone opens a chat window.

469
00:16:03,760 --> 00:16:05,960
Where co-pilot's retrieval actually lives,

470
00:16:05,960 --> 00:16:08,960
let's get concrete about where all of this actually plays out

471
00:16:08,960 --> 00:16:11,800
because up to now we've been talking architecture in the abstract.

472
00:16:11,800 --> 00:16:15,120
Inside Microsoft 365, co-pilot indexes SharePoint OneDrive,

473
00:16:15,120 --> 00:16:17,440
Teams Exchange, and whatever other connected sources

474
00:16:17,440 --> 00:16:18,920
your organization has plugged in,

475
00:16:18,920 --> 00:16:21,320
it respects your permission structure while it does this.

476
00:16:21,320 --> 00:16:23,800
So someone only sees what they're already allowed to see.

477
00:16:23,800 --> 00:16:25,440
That parts real and it matters.

478
00:16:25,440 --> 00:16:27,160
But look at what that list actually is.

479
00:16:27,160 --> 00:16:29,320
It's a description of Rag at Enterprise Scale.

480
00:16:29,320 --> 00:16:31,520
Microsoft calls it Graph Grounding.

481
00:16:31,520 --> 00:16:33,240
And the name makes it sound like something new,

482
00:16:33,240 --> 00:16:35,120
something built specifically for this moment.

483
00:16:35,120 --> 00:16:35,920
It isn't.

484
00:16:35,920 --> 00:16:37,520
Strip the branding off and its retrieval

485
00:16:37,520 --> 00:16:39,480
same as we've been describing this whole episode,

486
00:16:39,480 --> 00:16:41,240
just dressed in Enterprise clothing.

487
00:16:41,240 --> 00:16:44,280
Co-pilot reaches into Graph, pulls whatever content looks relevant

488
00:16:44,280 --> 00:16:47,120
to the question and generates an answer from those fragments.

489
00:16:47,120 --> 00:16:49,920
That's the identical mechanism running against a much bigger,

490
00:16:49,920 --> 00:16:52,400
much messier pile of source material.

491
00:16:52,400 --> 00:16:54,440
Now, the permission model deserves real credit here

492
00:16:54,440 --> 00:16:56,400
because it solves a genuinely hard problem.

493
00:16:56,400 --> 00:16:58,560
Making sure someone in finance doesn't accidentally

494
00:16:58,560 --> 00:17:01,240
see HR's compensation planning docs through a chat window

495
00:17:01,240 --> 00:17:03,640
that's not trivial and Microsoft built real infrastructure

496
00:17:03,640 --> 00:17:04,560
to handle it.

497
00:17:04,560 --> 00:17:06,920
But notice exactly what that infrastructure solves.

498
00:17:06,920 --> 00:17:09,520
It solves the access problem, who's allowed to see what.

499
00:17:09,520 --> 00:17:11,840
It does absolutely nothing about the memory problem

500
00:17:11,840 --> 00:17:13,680
we spent the last several sections on.

501
00:17:13,680 --> 00:17:16,920
Permissions decide whether a document is eligible to be retrieved.

502
00:17:16,920 --> 00:17:19,680
They say nothing about whether the system remembers retrieving it

503
00:17:19,680 --> 00:17:22,240
or builds on it or connects it to the conversation

504
00:17:22,240 --> 00:17:23,240
from yesterday.

505
00:17:23,240 --> 00:17:25,760
Access and memory are two entirely separate questions

506
00:17:25,760 --> 00:17:27,520
and solving one doesn't touch the other.

507
00:17:27,520 --> 00:17:29,240
Here's where it gets interesting though.

508
00:17:29,240 --> 00:17:32,240
Microsoft's own guidance quietly admits something worth pausing

509
00:17:32,240 --> 00:17:32,720
on.

510
00:17:32,720 --> 00:17:35,000
Admins can now register specific SharePoint sites

511
00:17:35,000 --> 00:17:36,800
as knowledge sources for Co-pilot.

512
00:17:36,800 --> 00:17:38,000
Read that carefully.

513
00:17:38,000 --> 00:17:40,240
That feature only makes sense if not all content

514
00:17:40,240 --> 00:17:42,360
is equally trustworthy or equally structured

515
00:17:42,360 --> 00:17:44,920
or equally worth trusting Co-pilot to pull from.

516
00:17:44,920 --> 00:17:47,040
If every SharePoint site were equally reliable,

517
00:17:47,040 --> 00:17:49,800
there'd be no reason to let admins flag some of them as special.

518
00:17:49,800 --> 00:17:52,400
The existence of that toggle is Microsoft tacitly admitting

519
00:17:52,400 --> 00:17:53,440
the obvious.

520
00:17:53,440 --> 00:17:55,720
Some of what's sitting in your tenant is good, organized,

521
00:17:55,720 --> 00:17:59,080
current, and some of it is stale, duplicated, or just noise.

522
00:17:59,080 --> 00:18:01,000
And Co-pilot, left to its own devices,

523
00:18:01,000 --> 00:18:02,360
treats all of it the same.

524
00:18:02,360 --> 00:18:04,960
That single feature is Microsoft nudging your organization

525
00:18:04,960 --> 00:18:08,200
toward curation without ever fully naming the underlying issue

526
00:18:08,200 --> 00:18:09,160
out loud.

527
00:18:09,160 --> 00:18:11,600
It's a quiet acknowledgement wrapped in an admin setting.

528
00:18:11,600 --> 00:18:13,120
Nobody in the documentation says,

529
00:18:13,120 --> 00:18:16,400
our retrieval system can't tell good content from bad content,

530
00:18:16,400 --> 00:18:18,000
so we built your workaround.

531
00:18:18,000 --> 00:18:20,440
But that's functionally what registering a knowledge source

532
00:18:20,440 --> 00:18:21,040
does.

533
00:18:21,040 --> 00:18:23,160
It's damage control for a structural limitation

534
00:18:23,160 --> 00:18:24,600
presented as a feature.

535
00:18:24,600 --> 00:18:28,040
And curation, as a starting point, is genuinely useful.

536
00:18:28,040 --> 00:18:30,040
But it's not the same thing as compilation,

537
00:18:30,040 --> 00:18:31,560
and the difference between those two words

538
00:18:31,560 --> 00:18:33,880
is exactly where this conversation needs to go next.

539
00:18:33,880 --> 00:18:35,680
Curation isn't compilation.

540
00:18:35,680 --> 00:18:37,600
So let's draw the line clearly, because it's

541
00:18:37,600 --> 00:18:39,360
easy to conflate these two ideas.

542
00:18:39,360 --> 00:18:42,080
And conflating them is exactly how organizations end up

543
00:18:42,080 --> 00:18:43,920
thinking they've solved the memory problem

544
00:18:43,920 --> 00:18:46,520
when they've only solved the access problem twice.

545
00:18:46,520 --> 00:18:48,760
Marking a SharePoint site as a trusted knowledge source

546
00:18:48,760 --> 00:18:50,120
tells Co-pilot where to look.

547
00:18:50,120 --> 00:18:50,800
That's all it does.

548
00:18:50,800 --> 00:18:51,480
It's a pointer.

549
00:18:51,480 --> 00:18:53,840
It says, when you're searching, wait this location more heavily

550
00:18:53,840 --> 00:18:54,880
than that one.

551
00:18:54,880 --> 00:18:56,760
What it doesn't do, what it can't do,

552
00:18:56,760 --> 00:18:59,040
is tell Co-pilot what it learned from that site.

553
00:18:59,040 --> 00:19:01,040
There's no learning happening in the act

554
00:19:01,040 --> 00:19:02,560
of registering a knowledge source.

555
00:19:02,560 --> 00:19:04,680
There's just a narrower, more confident search.

556
00:19:04,680 --> 00:19:06,160
And that's the part we're sitting with.

557
00:19:06,160 --> 00:19:08,640
Curated sources still get retrieved fresh, chunked fresh,

558
00:19:08,640 --> 00:19:10,440
synthesized fresh every single time someone

559
00:19:10,440 --> 00:19:11,400
asks a question.

560
00:19:11,400 --> 00:19:13,520
Registring a site as trusted doesn't change

561
00:19:13,520 --> 00:19:15,520
the mechanism underneath it at all.

562
00:19:15,520 --> 00:19:17,560
It changes which pile of documents get searched.

563
00:19:17,560 --> 00:19:19,800
It doesn't change the fact that it's still a search starting

564
00:19:19,800 --> 00:19:21,280
from nothing every time.

565
00:19:21,280 --> 00:19:23,520
You've made the haste-axe more and more reliable.

566
00:19:23,520 --> 00:19:26,400
You haven't given the system a memory of what's in it.

567
00:19:26,400 --> 00:19:27,920
The wiki pattern goes one step further,

568
00:19:27,920 --> 00:19:30,000
and this is the actual structural difference.

569
00:19:30,000 --> 00:19:31,280
It doesn't just point at good sources

570
00:19:31,280 --> 00:19:33,200
and say, look here first.

571
00:19:33,200 --> 00:19:35,960
It processes those sources into standing knowledge.

572
00:19:35,960 --> 00:19:38,160
It reads them once and turns them into something

573
00:19:38,160 --> 00:19:40,800
that already exists before the next question arrives.

574
00:19:40,800 --> 00:19:41,920
That's not a better pointer.

575
00:19:41,920 --> 00:19:43,720
That's a different kind of object entirely.

576
00:19:43,720 --> 00:19:45,080
Here's the planist way to say it.

577
00:19:45,080 --> 00:19:46,840
Curation is picking better ingredients.

578
00:19:46,840 --> 00:19:48,440
You've gone through the pantry, thrown out,

579
00:19:48,440 --> 00:19:50,240
what's expired, kept what's fresh,

580
00:19:50,240 --> 00:19:53,440
and told the kitchen, cook from the shelf, not that one.

581
00:19:53,440 --> 00:19:54,520
That's real and it matters.

582
00:19:54,520 --> 00:19:56,760
But compilation is actually cooking the dish once

583
00:19:56,760 --> 00:19:59,040
and keeping it in the fridge ready to serve.

584
00:19:59,040 --> 00:20:00,560
Nobody's back in the kitchen chopping

585
00:20:00,560 --> 00:20:03,240
and searing from scratch every time someone's hungry.

586
00:20:03,240 --> 00:20:04,360
The work happened once.

587
00:20:04,360 --> 00:20:05,800
What's left is just serving it.

588
00:20:05,800 --> 00:20:08,840
Without that second step, even perfectly curated content

589
00:20:08,840 --> 00:20:11,520
still forces co-pilot to start over on every question.

590
00:20:11,520 --> 00:20:13,840
You can register every sharepoint site in your tenant

591
00:20:13,840 --> 00:20:15,760
as a trusted knowledge source, get the pantry

592
00:20:15,760 --> 00:20:17,000
as clean as it's ever going to be

593
00:20:17,000 --> 00:20:19,400
and co-pilot will still chunk those documents fresh,

594
00:20:19,400 --> 00:20:21,400
still search them fresh, still synthesize

595
00:20:21,400 --> 00:20:23,560
an answer fresh every single time.

596
00:20:23,560 --> 00:20:25,560
Curation narrows what gets searched.

597
00:20:25,560 --> 00:20:27,400
It doesn't touch how the searching works

598
00:20:27,400 --> 00:20:28,800
or what happens to the understanding

599
00:20:28,800 --> 00:20:30,280
once the answer's been delivered.

600
00:20:30,280 --> 00:20:33,560
It evaporates, same as always, waiting for the next question

601
00:20:33,560 --> 00:20:34,880
to rebuild it from nothing.

602
00:20:34,880 --> 00:20:36,480
So curation is a real improvement

603
00:20:36,480 --> 00:20:38,400
and worth doing regardless of anything else

604
00:20:38,400 --> 00:20:40,200
in this episode, but it's not the fix.

605
00:20:40,200 --> 00:20:42,600
It's step one of a process that most organizations

606
00:20:42,600 --> 00:20:45,200
stop halfway through mistaking a cleaner haystack

607
00:20:45,200 --> 00:20:46,520
for an actual memory.

608
00:20:46,520 --> 00:20:48,720
This is exactly where the week is three-layer structure

609
00:20:48,720 --> 00:20:51,840
stops being an abstract idea from AI research circles

610
00:20:51,840 --> 00:20:53,680
and starts being something you could genuinely build

611
00:20:53,680 --> 00:20:55,760
inside your own tenant.

612
00:20:55,760 --> 00:20:57,920
The three layers applied to co-pilot.

613
00:20:57,920 --> 00:21:00,720
So let's put actual Microsoft 365 nouns

614
00:21:00,720 --> 00:21:02,200
onto those three layers because that's

615
00:21:02,200 --> 00:21:04,120
where this stops sounding like a research idea

616
00:21:04,120 --> 00:21:05,800
and starts sounding like something you could sketch

617
00:21:05,800 --> 00:21:08,080
on a whiteboard this week.

618
00:21:08,080 --> 00:21:09,800
Layer one raw sources.

619
00:21:09,800 --> 00:21:11,920
In your tenant, that's your SharePoint libraries,

620
00:21:11,920 --> 00:21:14,080
your Teams meeting transcripts, exchange threads

621
00:21:14,080 --> 00:21:16,920
that carry real decisions buried in reply chains.

622
00:21:16,920 --> 00:21:19,520
Viva engage posts where someone announced a policy change

623
00:21:19,520 --> 00:21:21,520
nobody bothered to document properly.

624
00:21:21,520 --> 00:21:23,160
All of it stays exactly where it is.

625
00:21:23,160 --> 00:21:24,520
Nobody touches it, nobody edits it,

626
00:21:24,520 --> 00:21:27,120
nobody deletes it to clean things up.

627
00:21:27,120 --> 00:21:28,880
It's ground truth, sitting untouched,

628
00:21:28,880 --> 00:21:30,720
the same way it should sit, whether or not

629
00:21:30,720 --> 00:21:32,520
you ever build anything on top of it.

630
00:21:32,520 --> 00:21:34,200
Layer two, the wiki itself.

631
00:21:34,200 --> 00:21:36,360
This is the part worth picturing concretely,

632
00:21:36,360 --> 00:21:37,800
not a folder of documents.

633
00:21:37,800 --> 00:21:40,000
A maintained set of pages, one per project,

634
00:21:40,000 --> 00:21:41,920
one per client, one per recurring decision,

635
00:21:41,920 --> 00:21:45,400
your organization keeps having to re-explain to somebody new.

636
00:21:45,400 --> 00:21:47,080
Cross-linked, so the page about a client

637
00:21:47,080 --> 00:21:48,520
connects to the page about the project,

638
00:21:48,520 --> 00:21:50,760
connects to the page about the compliance interpretation

639
00:21:50,760 --> 00:21:52,560
that shaped how that project got scripted.

640
00:21:52,560 --> 00:21:54,800
This is the layer that didn't exist an hour ago

641
00:21:54,800 --> 00:21:57,560
in your tenant and does exist now once someone builds it.

642
00:21:57,560 --> 00:21:59,160
Layer three, the schema.

643
00:21:59,160 --> 00:22:01,480
Think of this as the rule book, the thing that

644
00:22:01,480 --> 00:22:04,040
decides how any of this actually happens

645
00:22:04,040 --> 00:22:05,680
when a new Teams transcript comes in,

646
00:22:05,680 --> 00:22:07,480
does it update an existing client page

647
00:22:07,480 --> 00:22:08,800
or does it spin up a new one?

648
00:22:08,800 --> 00:22:10,480
How does a page get formatted?

649
00:22:10,480 --> 00:22:12,920
So it's consistent with every other page in the wiki.

650
00:22:12,920 --> 00:22:16,000
And critically, what happens when something in that transcript

651
00:22:16,000 --> 00:22:17,680
contradicts what's already written down?

652
00:22:17,680 --> 00:22:19,160
The schema is what answers those questions

653
00:22:19,160 --> 00:22:21,760
before anyone has to answer them manually every single time.

654
00:22:21,760 --> 00:22:24,040
Now here's the thing, none of this should feel foreign

655
00:22:24,040 --> 00:22:27,160
to anyone who spent time in Microsoft 365 governance.

656
00:22:27,160 --> 00:22:28,880
Map it onto what you already understand.

657
00:22:28,880 --> 00:22:31,800
Retention labels already decide what happens to content

658
00:22:31,800 --> 00:22:34,200
based on rules you defined in advance.

659
00:22:34,200 --> 00:22:35,840
Content types already give structure

660
00:22:35,840 --> 00:22:38,760
to what would otherwise be undifferentiated files.

661
00:22:38,760 --> 00:22:40,960
Metadata schemas already tell you what fields

662
00:22:40,960 --> 00:22:43,240
a document needs before it counts as complete.

663
00:22:43,240 --> 00:22:46,000
This isn't a new discipline landing on your desk out of nowhere.

664
00:22:46,000 --> 00:22:48,360
It's an extension of information architecture work.

665
00:22:48,360 --> 00:22:50,720
Your organization has probably already been doing

666
00:22:50,720 --> 00:22:52,160
just aimed at a different output.

667
00:22:52,160 --> 00:22:54,760
Instead of the schema governing how a document gets filed,

668
00:22:54,760 --> 00:22:57,640
it's governing how a wiki page gets written and maintained.

669
00:22:57,640 --> 00:22:59,440
But here's the part that doesn't happen automatically

670
00:22:59,440 --> 00:23:01,120
and it's worth being blunt about it.

671
00:23:01,120 --> 00:23:04,680
Someone or something has to actually do the compiling.

672
00:23:04,680 --> 00:23:06,240
The three layers we just walked through

673
00:23:06,240 --> 00:23:09,640
don't assemble themselves the moment you understand the concept.

674
00:23:09,640 --> 00:23:12,520
Current copilot deployments don't ship with a background process

675
00:23:12,520 --> 00:23:13,960
that reads your SharePoint libraries

676
00:23:13,960 --> 00:23:15,840
and writes structured pages out of them.

677
00:23:15,840 --> 00:23:16,840
That work is real work.

678
00:23:16,840 --> 00:23:18,960
It has to be built or run or triggered

679
00:23:18,960 --> 00:23:21,120
by something that treats it as its job.

680
00:23:21,120 --> 00:23:24,080
Right now, in a standard Microsoft 365 environment,

681
00:23:24,080 --> 00:23:25,440
nothing is doing that job.

682
00:23:25,440 --> 00:23:28,120
Copilot searches it doesn't compile the wiki layer

683
00:23:28,120 --> 00:23:31,160
if it exists at all exists because somebody decided to build it,

684
00:23:31,160 --> 00:23:33,160
not because it came in the box.

685
00:23:33,160 --> 00:23:35,160
Which puts a fairly obvious question on the table

686
00:23:35,160 --> 00:23:36,640
and it's the one worth answering next.

687
00:23:36,640 --> 00:23:38,560
If this layer doesn't ship by default

688
00:23:38,560 --> 00:23:42,000
and it doesn't build itself, who or what is actually supposed

689
00:23:42,000 --> 00:23:44,000
to build it inside a Microsoft ecosystem?

690
00:23:44,000 --> 00:23:45,760
Because the three layer structure only matters

691
00:23:45,760 --> 00:23:47,320
if there's a real answer to that question,

692
00:23:47,320 --> 00:23:49,920
not just a diagram that looks clean on a slide.

693
00:23:49,920 --> 00:23:51,400
Who builds the wiki layer?

694
00:23:51,400 --> 00:23:53,480
Let's be honest about something before we go any further

695
00:23:53,480 --> 00:23:56,160
because glossing over it would make the rest of this episode

696
00:23:56,160 --> 00:23:57,840
sound like a feature announcement.

697
00:23:57,840 --> 00:24:00,280
This isn't a Microsoft ship capability today.

698
00:24:00,280 --> 00:24:02,320
There's no toggle in the admin center labeled

699
00:24:02,320 --> 00:24:04,760
turn on compiled knowledge.

700
00:24:04,760 --> 00:24:06,440
What we've been describing is a pattern,

701
00:24:06,440 --> 00:24:08,920
something you'd implement on top of your existing knowledge

702
00:24:08,920 --> 00:24:10,320
state, not something that arrives

703
00:24:10,320 --> 00:24:11,960
with your next licensing renewal.

704
00:24:11,960 --> 00:24:13,960
So what does implementing it actually look like?

705
00:24:13,960 --> 00:24:16,160
Picture an AI agent and if you've spent any time

706
00:24:16,160 --> 00:24:19,040
in Copilot Studio, the shape of this will feel familiar.

707
00:24:19,040 --> 00:24:20,880
Not unlike how a Copilot Studio agent

708
00:24:20,880 --> 00:24:23,120
gets built to watch a trigger and take an action

709
00:24:23,120 --> 00:24:24,800
or how a custom graph connected agent

710
00:24:24,800 --> 00:24:28,160
gets set up to reach into specific data sources on a schedule.

711
00:24:28,160 --> 00:24:31,240
Same underlying capability pointed at a different job.

712
00:24:31,240 --> 00:24:33,760
Instead of answering a user's question in the moment,

713
00:24:33,760 --> 00:24:36,200
this agent's entire purpose is reading sources

714
00:24:36,200 --> 00:24:38,200
and writing structured pages out of them.

715
00:24:38,200 --> 00:24:40,720
It goes into a SharePoint library, reads what's there,

716
00:24:40,720 --> 00:24:42,480
decides what belongs on an existing page

717
00:24:42,480 --> 00:24:44,640
and what deserves a new one and writes accordingly.

718
00:24:44,640 --> 00:24:45,680
Nobody's chatting with it.

719
00:24:45,680 --> 00:24:47,040
It's not waiting for a prompt.

720
00:24:47,040 --> 00:24:49,360
It's just working quietly in the background

721
00:24:49,360 --> 00:24:51,080
whether or not anyone's logged in that day.

722
00:24:51,080 --> 00:24:52,520
That distinction is worth sitting with

723
00:24:52,520 --> 00:24:55,040
because it's where agente orchestration actually earns

724
00:24:55,040 --> 00:24:56,880
its name instead of just being a buzzword.

725
00:24:56,880 --> 00:24:59,120
This isn't a chatbot answering questions on demand.

726
00:24:59,120 --> 00:25:01,680
It's a background process running on its own schedule,

727
00:25:01,680 --> 00:25:03,160
compiling and maintaining knowledge,

728
00:25:03,160 --> 00:25:05,880
whether or not a single person asks it anything that day.

729
00:25:05,880 --> 00:25:07,360
The value isn't in the conversation.

730
00:25:07,360 --> 00:25:09,160
It's in the fact that the conversation never has

731
00:25:09,160 --> 00:25:11,240
to start from zero because something already did

732
00:25:11,240 --> 00:25:12,960
the work before the question showed up.

733
00:25:12,960 --> 00:25:14,960
And here's the part that should feel encouraging

734
00:25:14,960 --> 00:25:15,800
rather than daunting.

735
00:25:15,800 --> 00:25:18,240
None of the infrastructure to build this is missing.

736
00:25:18,240 --> 00:25:20,840
Copilot Studio already supports custom connectors,

737
00:25:20,840 --> 00:25:23,000
reaching into whatever data source you pointed at.

738
00:25:23,000 --> 00:25:24,720
It already supports scheduled flows,

739
00:25:24,720 --> 00:25:26,480
running processes on a timer without a human

740
00:25:26,480 --> 00:25:27,680
kicking them off manually.

741
00:25:27,680 --> 00:25:29,960
Those two capabilities connectors and schedules

742
00:25:29,960 --> 00:25:31,360
are most of what you need to build

743
00:25:31,360 --> 00:25:33,160
the compiling half of the system.

744
00:25:33,160 --> 00:25:35,000
The pieces exist in your tenant right now.

745
00:25:35,000 --> 00:25:36,840
They're just not assembled this way by default

746
00:25:36,840 --> 00:25:39,040
because nobody's shipped a template that says,

747
00:25:39,040 --> 00:25:40,680
"Here's how you wire a connector

748
00:25:40,680 --> 00:25:42,680
and a schedule together to build a wiki

749
00:25:42,680 --> 00:25:45,160
instead of just automating a workflow."

750
00:25:45,160 --> 00:25:47,960
Which means the real barrier here isn't technical capability.

751
00:25:47,960 --> 00:25:51,200
It's a shift in how IT thinks about what Copilot even is.

752
00:25:51,200 --> 00:25:54,400
Right now, most organizations think of Copilot as one thing,

753
00:25:54,400 --> 00:25:55,680
a chat interface.

754
00:25:55,680 --> 00:25:56,720
You type, it answers.

755
00:25:56,720 --> 00:25:58,080
That's the whole mental model.

756
00:25:58,080 --> 00:26:00,240
But once you see the wiki pattern clearly,

757
00:26:00,240 --> 00:26:02,840
Copilot stops being one system and starts being two.

758
00:26:02,840 --> 00:26:04,720
There's the compiler, the background agent quietly

759
00:26:04,720 --> 00:26:06,680
reading, writing and maintaining structured pages.

760
00:26:06,680 --> 00:26:09,080
And there's the responder, the familiar chat interface

761
00:26:09,080 --> 00:26:11,320
that answers questions, except now it's answering

762
00:26:11,320 --> 00:26:13,800
from compiled pages instead of raw fragments.

763
00:26:13,800 --> 00:26:16,440
Two systems doing two different jobs working together.

764
00:26:16,440 --> 00:26:19,800
Most IT teams today have only built or bought the second one.

765
00:26:19,800 --> 00:26:22,760
Once you see it that way, as two systems instead of one,

766
00:26:22,760 --> 00:26:24,200
a different question starts to matter

767
00:26:24,200 --> 00:26:25,760
more than efficiency ever did.

768
00:26:25,760 --> 00:26:28,040
Because splitting the compiler from the responder

769
00:26:28,040 --> 00:26:30,440
doesn't just make answers cheaper or faster.

770
00:26:30,440 --> 00:26:31,880
It changes something more fundamental

771
00:26:31,880 --> 00:26:34,040
about whether people trust what Copilot tells them

772
00:26:34,040 --> 00:26:35,280
in the first place.

773
00:26:35,280 --> 00:26:37,760
The trust problem, RAAG, can't solve.

774
00:26:37,760 --> 00:26:39,560
There's data behind this, and it's worth putting

775
00:26:39,560 --> 00:26:41,560
a real number on it before we go further.

776
00:26:41,560 --> 00:26:43,640
Over 60% of enterprise IT leaders

777
00:26:43,640 --> 00:26:46,000
said they intended to expand generative AI

778
00:26:46,000 --> 00:26:47,200
for knowledge discovery.

779
00:26:47,200 --> 00:26:48,320
That's not a small commitment.

780
00:26:48,320 --> 00:26:50,520
That's leadership teams planning to lean harder

781
00:26:50,520 --> 00:26:52,960
into this budgeting forward, telling their boards

782
00:26:52,960 --> 00:26:54,520
it's the next phase of the rollout.

783
00:26:54,520 --> 00:26:56,680
And then in a lot of those same organizations adoption

784
00:26:56,680 --> 00:26:59,160
stalls, not because people forgot to use the tool,

785
00:26:59,160 --> 00:27:00,520
because somewhere along the way,

786
00:27:00,520 --> 00:27:03,000
the answer started feeling inconsistent.

787
00:27:03,000 --> 00:27:05,160
And inconsistency is quietly fatal

788
00:27:05,160 --> 00:27:06,920
to whether people keep coming back.

789
00:27:06,920 --> 00:27:09,480
Here's what's worth understanding about that inconsistency,

790
00:27:09,480 --> 00:27:11,960
because it isn't random noise you can chalk up to bad luck

791
00:27:11,960 --> 00:27:13,480
or a model having an off day.

792
00:27:13,480 --> 00:27:14,600
It's structural.

793
00:27:14,600 --> 00:27:17,240
Ask the same underlying question two different ways,

794
00:27:17,240 --> 00:27:19,400
and RAAG doesn't just phrase the answer differently.

795
00:27:19,400 --> 00:27:21,640
It retrieves an entirely different set of chunks,

796
00:27:21,640 --> 00:27:23,320
different search, different fragments pulled

797
00:27:23,320 --> 00:27:25,080
from the index different material handed

798
00:27:25,080 --> 00:27:26,800
to the model to synthesize from.

799
00:27:26,800 --> 00:27:28,280
Two people asking about the same policy,

800
00:27:28,280 --> 00:27:31,240
one saying, what's our remote work policy and the other saying,

801
00:27:31,240 --> 00:27:33,080
can I work from home three days a week?

802
00:27:33,080 --> 00:27:35,480
Might get two answers that don't even agree with each other.

803
00:27:35,480 --> 00:27:38,280
Because under the hood, they triggered two separate searches

804
00:27:38,280 --> 00:27:40,640
that happened to land on different documents.

805
00:27:40,640 --> 00:27:42,240
Users notice this fast.

806
00:27:42,240 --> 00:27:44,040
And here's the part that should worry you more

807
00:27:44,040 --> 00:27:45,360
than a formal complaint would.

808
00:27:45,360 --> 00:27:46,320
They don't complain loudly.

809
00:27:46,320 --> 00:27:49,080
Nobody files a ticket that says co-pilot gave me inconsistent

810
00:27:49,080 --> 00:27:49,880
answers.

811
00:27:49,880 --> 00:27:52,400
What actually happens is quieter and harder to track.

812
00:27:52,400 --> 00:27:55,600
Someone gets burned once, maybe twice, on a question that mattered.

813
00:27:55,600 --> 00:27:57,960
And they just stop asking co-pilot the harder questions.

814
00:27:57,960 --> 00:27:59,520
They still use it for quick summaries,

815
00:27:59,520 --> 00:28:02,480
for drafting an email, for things with no real stakes attached.

816
00:28:02,480 --> 00:28:04,600
But the moment a real decision rides on the answer,

817
00:28:04,600 --> 00:28:06,160
they go find a person instead.

818
00:28:06,160 --> 00:28:08,720
That shift happens silently one person at a time.

819
00:28:08,720 --> 00:28:10,360
And it never shows up as a complaint.

820
00:28:10,360 --> 00:28:12,480
It shows up months later as usage numbers

821
00:28:12,480 --> 00:28:15,200
that look fine on paper, but never touch the questions

822
00:28:15,200 --> 00:28:16,280
that actually mattered.

823
00:28:16,280 --> 00:28:19,120
Now compare that to what a compiled wiki page does.

824
00:28:19,120 --> 00:28:21,680
Ask it the same underlying question five different ways,

825
00:28:21,680 --> 00:28:24,240
phrased five different ways, worded by five different people.

826
00:28:24,240 --> 00:28:25,800
You get the same answer every time.

827
00:28:25,800 --> 00:28:28,160
Because the synthesis already happened once carefully

828
00:28:28,160 --> 00:28:30,680
before any of those five people ever typed anything.

829
00:28:30,680 --> 00:28:32,440
There's no fresh search rolling the dice

830
00:28:32,440 --> 00:28:33,920
on which fragments get pulled.

831
00:28:33,920 --> 00:28:36,320
There's one page already reconciled, already settled,

832
00:28:36,320 --> 00:28:38,840
and every version of the question just gets pointed at it.

833
00:28:38,840 --> 00:28:40,840
This is the part worth naming directly,

834
00:28:40,840 --> 00:28:44,480
because it's easy to file under nice to have and move past.

835
00:28:44,480 --> 00:28:46,200
Consistency isn't a polish feature.

836
00:28:46,200 --> 00:28:47,280
It's a trust mechanism.

837
00:28:47,280 --> 00:28:49,400
It's the actual thing that decides whether someone

838
00:28:49,400 --> 00:28:52,880
is willing to let a tools answer carry weight in a real decision,

839
00:28:52,880 --> 00:28:55,080
or whether they keep it around for the easy stuff

840
00:28:55,080 --> 00:28:58,520
and quietly root anything important through a person instead.

841
00:28:58,520 --> 00:29:00,400
Once trust breaks on the hard questions,

842
00:29:00,400 --> 00:29:02,240
it doesn't come back just because the interface

843
00:29:02,240 --> 00:29:03,640
gets a new code of paint.

844
00:29:03,640 --> 00:29:05,280
And trust is only half the payoff here.

845
00:29:05,280 --> 00:29:07,960
The other half shows up somewhere you might not expect.

846
00:29:07,960 --> 00:29:10,520
And how much redundant work is happening across your organization

847
00:29:10,520 --> 00:29:11,560
right now.

848
00:29:11,560 --> 00:29:14,040
Work that exists purely because nobody trusted search

849
00:29:14,040 --> 00:29:17,120
to surface the answer that already existed.

850
00:29:17,120 --> 00:29:19,080
Killing redundant knowledge creation?

851
00:29:19,080 --> 00:29:20,880
There's another number worth putting on the table

852
00:29:20,880 --> 00:29:22,600
and it points at something most organizations

853
00:29:22,600 --> 00:29:24,280
never think to measure.

854
00:29:24,280 --> 00:29:26,880
Early adopters of AI-based knowledge discovery report

855
00:29:26,880 --> 00:29:30,360
a 25% to 40% reduction in redundant knowledge creation.

856
00:29:30,360 --> 00:29:32,440
That's a strange phrase, redundant knowledge creation.

857
00:29:32,440 --> 00:29:34,000
So let's unpack what it actually means.

858
00:29:34,000 --> 00:29:36,480
Because once you see it, you start noticing it everywhere.

859
00:29:36,480 --> 00:29:38,960
Here's the mechanism, and it's simpler than it sounds.

860
00:29:38,960 --> 00:29:41,840
When nobody trusts search to surface the right answer,

861
00:29:41,840 --> 00:29:43,000
people don't just give up.

862
00:29:43,000 --> 00:29:43,600
They rebuild.

863
00:29:43,600 --> 00:29:46,000
Someone needs a document explaining how a process works.

864
00:29:46,000 --> 00:29:47,680
They run a search, get nothing solid,

865
00:29:47,680 --> 00:29:49,920
and instead of digging further, they just write it again

866
00:29:49,920 --> 00:29:51,080
from scratch.

867
00:29:51,080 --> 00:29:53,120
Except that document already exists.

868
00:29:53,120 --> 00:29:55,840
It's sitting in a team's thread from four months ago

869
00:29:55,840 --> 00:29:58,280
or buried on page three of an old SharePoint site.

870
00:29:58,280 --> 00:29:59,840
Nobody thinks to check anymore.

871
00:29:59,840 --> 00:30:01,080
The information isn't missing.

872
00:30:01,080 --> 00:30:04,040
It's unfindable, which functionally amounts to the same thing.

873
00:30:04,040 --> 00:30:07,360
And unfindable information gets recreated over and over

874
00:30:07,360 --> 00:30:09,240
by different people who have no idea.

875
00:30:09,240 --> 00:30:11,400
They're duplicating work that already happened.

876
00:30:11,400 --> 00:30:13,640
This is where a compiled weeky changes the math.

877
00:30:13,640 --> 00:30:15,240
Instead of a search turning up nothing

878
00:30:15,240 --> 00:30:18,200
useful or five slightly different fragments that half answer

879
00:30:18,200 --> 00:30:20,520
the question, it surfaces the existing answer

880
00:30:20,520 --> 00:30:24,120
directly with a citation pointing back to where it came from.

881
00:30:24,120 --> 00:30:26,520
Someone doesn't spend an afternoon reconstructing a document

882
00:30:26,520 --> 00:30:28,680
that already lives in the company's collective memory.

883
00:30:28,680 --> 00:30:30,560
They find it in minutes because the weeky

884
00:30:30,560 --> 00:30:33,120
already did the work of deciding that this concept exists

885
00:30:33,120 --> 00:30:34,320
and here's where it lives.

886
00:30:34,320 --> 00:30:36,240
And this connects back to something we touched on earlier,

887
00:30:36,240 --> 00:30:38,600
but it's worth stating in its sharpest form here.

888
00:30:38,600 --> 00:30:40,680
Ragtreeze every piece of content is equally hard

889
00:30:40,680 --> 00:30:41,880
to find every single time.

890
00:30:41,880 --> 00:30:44,040
It doesn't matter if someone asked the exact same question

891
00:30:44,040 --> 00:30:45,960
last week and got a great answer.

892
00:30:45,960 --> 00:30:48,280
The system starts over searching the same haystack

893
00:30:48,280 --> 00:30:50,280
with the same odds of missing what it's looking for,

894
00:30:50,280 --> 00:30:51,840
a wiki flips that completely.

895
00:30:51,840 --> 00:30:54,480
It treats concepts as already found permanently.

896
00:30:54,480 --> 00:30:57,440
Once something's compiled onto a page, it stays found.

897
00:30:57,440 --> 00:31:00,280
Nobody has to get lucky with their search terms ever again.

898
00:31:00,280 --> 00:31:02,920
Frame the business case here in the planist terms possible

899
00:31:02,920 --> 00:31:05,320
because this is where a lot of the argument in this episode

900
00:31:05,320 --> 00:31:08,080
stops being architectural and starts being financial.

901
00:31:08,080 --> 00:31:10,160
Less duplicated effort isn't a technical

902
00:31:10,160 --> 00:31:12,120
when you put on a slide about system design.

903
00:31:12,120 --> 00:31:16,000
It's hours, actual hours given back to people doing actual work

904
00:31:16,000 --> 00:31:17,560
instead of quietly rebuilding something

905
00:31:17,560 --> 00:31:19,600
that already existed somewhere in the tenant,

906
00:31:19,600 --> 00:31:22,800
multiply that across a few hundred employees over a year

907
00:31:22,800 --> 00:31:24,680
and you're not talking about a nice efficiency gain.

908
00:31:24,680 --> 00:31:26,480
You're talking about a meaningful chunk

909
00:31:26,480 --> 00:31:28,400
of your organization's collective time

910
00:31:28,400 --> 00:31:30,800
spent solving a problem that had already been solved.

911
00:31:30,800 --> 00:31:32,480
But none of this holds up if the material

912
00:31:32,480 --> 00:31:34,600
feeding the wiki in the first place is a mess

913
00:31:34,600 --> 00:31:36,040
and that's the next real constraint

914
00:31:36,040 --> 00:31:37,600
worth being honest about.

915
00:31:37,600 --> 00:31:39,440
Garbage in, still garbage out.

916
00:31:39,440 --> 00:31:40,640
Let's be direct about something

917
00:31:40,640 --> 00:31:42,800
the whole wiki pattern depends on because it's the part

918
00:31:42,800 --> 00:31:45,640
that's easiest to skip past when you're excited about the architecture.

919
00:31:45,640 --> 00:31:47,680
The wiki doesn't fix bad source material.

920
00:31:47,680 --> 00:31:49,040
It organizes what you give it.

921
00:31:49,040 --> 00:31:51,480
It doesn't audit whether what you gave it is true, current

922
00:31:51,480 --> 00:31:53,040
or even still relevant.

923
00:31:53,040 --> 00:31:54,840
That job never belonged to the compiler

924
00:31:54,840 --> 00:31:55,720
and it never will.

925
00:31:55,720 --> 00:31:57,440
Here's what that means in practice.

926
00:31:57,440 --> 00:32:00,280
If your share point is full of outdated policy documents,

927
00:32:00,280 --> 00:32:02,200
documents nobody archived properly,

928
00:32:02,200 --> 00:32:04,520
documents that describe a process your organization

929
00:32:04,520 --> 00:32:06,480
stopped using two reogs ago,

930
00:32:06,480 --> 00:32:08,480
the wiki will read every one of them faithfully

931
00:32:08,480 --> 00:32:12,000
and compile them into clean, confident looking pages.

932
00:32:12,000 --> 00:32:13,360
That's the uncomfortable part.

933
00:32:13,360 --> 00:32:15,080
The output won't look messy or uncertain.

934
00:32:15,080 --> 00:32:17,280
It'll look organized, it'll look authoritative

935
00:32:17,280 --> 00:32:19,440
and it'll be wrong presented with the same polish

936
00:32:19,440 --> 00:32:21,680
as everything else in the wiki that happens to be right.

937
00:32:21,680 --> 00:32:24,120
This is exactly where content lifecycle management

938
00:32:24,120 --> 00:32:26,120
stops being optional and it's worth pointing out

939
00:32:26,120 --> 00:32:27,440
that Microsoft has been saying this

940
00:32:27,440 --> 00:32:30,720
since its own January 2024 Copilot readiness guidance

941
00:32:30,720 --> 00:32:32,240
that guidance wasn't subtle about it.

942
00:32:32,240 --> 00:32:34,640
Content needs owners, it needs version control,

943
00:32:34,640 --> 00:32:36,440
it needs a process for retiring things

944
00:32:36,440 --> 00:32:38,840
that are no longer accurate instead of letting them sit

945
00:32:38,840 --> 00:32:40,880
indefinitely next to whatever replaced them.

946
00:32:40,880 --> 00:32:43,400
None of that is new advice invented for this episode.

947
00:32:43,400 --> 00:32:45,880
It's been sitting in Microsoft's own documentation

948
00:32:45,880 --> 00:32:47,440
for a while mostly ignored

949
00:32:47,440 --> 00:32:49,120
because it sounded like housekeeping

950
00:32:49,120 --> 00:32:50,680
rather than something that determines

951
00:32:50,680 --> 00:32:53,920
whether your AI investment actually works.

952
00:32:53,920 --> 00:32:56,200
So say it plainly because this is the part organizations

953
00:32:56,200 --> 00:32:57,160
keep skipping.

954
00:32:57,160 --> 00:32:59,800
Structured knowledge work starts before any AI

955
00:32:59,800 --> 00:33:02,080
touches your content retention labels still matter,

956
00:33:02,080 --> 00:33:04,000
content type still matter, clear ownership.

957
00:33:04,000 --> 00:33:05,280
So someone's actually responsible

958
00:33:05,280 --> 00:33:08,400
for knowing whether a document is still true, still matters.

959
00:33:08,400 --> 00:33:09,840
None of that becomes less important

960
00:33:09,840 --> 00:33:11,360
once you add a wiki layer on top.

961
00:33:11,360 --> 00:33:13,400
If anything, it becomes the entire foundation

962
00:33:13,400 --> 00:33:14,520
the wiki is standing on.

963
00:33:14,520 --> 00:33:16,120
And here's the sharper version of the risk

964
00:33:16,120 --> 00:33:17,280
the one worth sitting with.

965
00:33:17,280 --> 00:33:19,520
The wiki pattern raises the stakes on curation

966
00:33:19,520 --> 00:33:21,840
because it makes bad content look more authoritative,

967
00:33:21,840 --> 00:33:22,680
not less.

968
00:33:22,680 --> 00:33:25,240
A messy pile of raw documents, at least looks messy.

969
00:33:25,240 --> 00:33:26,960
People approach it with some skepticism.

970
00:33:26,960 --> 00:33:28,520
A beautifully compiled wiki page

971
00:33:28,520 --> 00:33:30,120
doesn't invite that same skepticism.

972
00:33:30,120 --> 00:33:32,040
It reads like something's already been verified,

973
00:33:32,040 --> 00:33:34,440
organized and settled even when the underlying source

974
00:33:34,440 --> 00:33:35,760
was garbage the whole time.

975
00:33:35,760 --> 00:33:38,120
So assuming your source material actually is solid,

976
00:33:38,120 --> 00:33:39,120
here's what happens next.

977
00:33:39,120 --> 00:33:40,760
Once new content starts arriving

978
00:33:40,760 --> 00:33:43,640
and the system has to decide what to do with it.

979
00:33:43,640 --> 00:33:45,560
When bad knowledge looks authoritative,

980
00:33:45,560 --> 00:33:47,280
here's what that means in practice.

981
00:33:47,280 --> 00:33:50,240
If your share point is full of outdated policy documents,

982
00:33:50,240 --> 00:33:51,880
documents nobody archived properly,

983
00:33:51,880 --> 00:33:54,160
documents that describe a process your organization

984
00:33:54,160 --> 00:33:56,000
stopped using two reoggs ago,

985
00:33:56,000 --> 00:33:58,240
the wiki will read every one of them faithfully

986
00:33:58,240 --> 00:34:01,160
and compile them into clean, confident looking pages.

987
00:34:01,160 --> 00:34:02,520
That's the uncomfortable part.

988
00:34:02,520 --> 00:34:04,640
The output won't look messy or uncertain.

989
00:34:04,640 --> 00:34:06,720
It'll look organized, it'll look authoritative

990
00:34:06,720 --> 00:34:09,040
and it'll be wrong, presented with the same polish

991
00:34:09,040 --> 00:34:11,440
as everything else in the wiki that happens to be right.

992
00:34:11,440 --> 00:34:13,880
This is exactly where content lifecycle management

993
00:34:13,880 --> 00:34:15,120
stops being optional.

994
00:34:15,120 --> 00:34:16,840
And it's worth pointing out that Microsoft

995
00:34:16,840 --> 00:34:20,080
has been saying this since its own January 2024

996
00:34:20,080 --> 00:34:21,760
co-pilot readiness guidance.

997
00:34:21,760 --> 00:34:23,480
That guidance wasn't subtle about it.

998
00:34:23,480 --> 00:34:25,000
Content needs owners.

999
00:34:25,000 --> 00:34:26,200
It needs version controlled.

1000
00:34:26,200 --> 00:34:28,000
It needs a process for retiring things

1001
00:34:28,000 --> 00:34:30,160
that are no longer accurate instead of letting them sit

1002
00:34:30,160 --> 00:34:32,440
indefinitely next to whatever replaced them.

1003
00:34:32,440 --> 00:34:34,800
None of that is new advice invented for this episode.

1004
00:34:34,800 --> 00:34:36,840
It's been sitting in Microsoft's own documentation

1005
00:34:36,840 --> 00:34:40,440
for a while, mostly ignored because it sounded like housekeeping

1006
00:34:40,440 --> 00:34:42,680
rather than something that determines whether your AI

1007
00:34:42,680 --> 00:34:44,560
investment actually works.

1008
00:34:44,560 --> 00:34:46,480
So say it plainly, because this is the part

1009
00:34:46,480 --> 00:34:47,960
organizations keep skipping.

1010
00:34:47,960 --> 00:34:51,680
Structured knowledge work starts before any AI touches your content.

1011
00:34:51,680 --> 00:34:54,480
Retention labels still matter, content types still matter.

1012
00:34:54,480 --> 00:34:56,480
Clear ownership, so someone's actually responsible

1013
00:34:56,480 --> 00:34:58,760
for knowing whether a document is still true still matters.

1014
00:34:58,760 --> 00:35:00,200
None of that becomes less important

1015
00:35:00,200 --> 00:35:01,960
once you add a wiki layer on top.

1016
00:35:01,960 --> 00:35:03,760
If anything, it becomes the entire foundation

1017
00:35:03,760 --> 00:35:05,160
the wiki is standing on.

1018
00:35:05,160 --> 00:35:06,840
And here's the sharper version of the risk

1019
00:35:06,840 --> 00:35:08,080
the one worth sitting with.

1020
00:35:08,080 --> 00:35:09,960
The wiki pattern raises the stakes on curation

1021
00:35:09,960 --> 00:35:12,320
because it makes bad content look more authoritative,

1022
00:35:12,320 --> 00:35:13,000
not less.

1023
00:35:13,000 --> 00:35:15,320
A messy pile of raw documents at least looks messy.

1024
00:35:15,320 --> 00:35:17,320
People approach it with some skepticism.

1025
00:35:17,320 --> 00:35:18,520
A beautifully compiled wiki page

1026
00:35:18,520 --> 00:35:20,160
doesn't invite that same skepticism.

1027
00:35:20,160 --> 00:35:22,560
It reads like something's already been verified, organized

1028
00:35:22,560 --> 00:35:24,640
and settled, even when the underlying source

1029
00:35:24,640 --> 00:35:25,920
was garbage the whole time.

1030
00:35:25,920 --> 00:35:28,320
So assuming your source material actually is solid,

1031
00:35:28,320 --> 00:35:29,960
here's what happens next.

1032
00:35:29,960 --> 00:35:33,680
Once new content starts arriving and the system has to decide

1033
00:35:33,680 --> 00:35:34,560
what to do with it.

1034
00:35:34,560 --> 00:35:37,280
So picture what actually happens when a new source lands

1035
00:35:37,280 --> 00:35:40,080
because this is where the wiki pattern stops being a diagram

1036
00:35:40,080 --> 00:35:42,240
and starts being a repeatable process.

1037
00:35:42,240 --> 00:35:44,240
A document shows up, maybe a policy update,

1038
00:35:44,240 --> 00:35:46,400
maybe a new team's transcript from a client call.

1039
00:35:46,400 --> 00:35:48,520
The system reads it once, not skims it,

1040
00:35:48,520 --> 00:35:50,960
not chunks it into fragments for later retrieval.

1041
00:35:50,960 --> 00:35:52,560
Reads it, the way you'd read something

1042
00:35:52,560 --> 00:35:54,440
if your job was to understand it well enough

1043
00:35:54,440 --> 00:35:56,040
to explain it to someone else.

1044
00:35:56,040 --> 00:35:58,280
Key concepts get pulled out during that read.

1045
00:35:58,280 --> 00:35:59,680
What's this document actually about?

1046
00:35:59,680 --> 00:36:00,520
What does it claim?

1047
00:36:00,520 --> 00:36:01,600
What does it connect to?

1048
00:36:01,600 --> 00:36:03,560
From there, one of two things happens.

1049
00:36:03,560 --> 00:36:05,360
And this is worth being precise about because it's

1050
00:36:05,360 --> 00:36:06,400
the whole mechanism.

1051
00:36:06,400 --> 00:36:08,520
If the new source adds detail to something already

1052
00:36:08,520 --> 00:36:11,040
compiled, an existing page gets updated.

1053
00:36:11,040 --> 00:36:13,200
Say there's already a page about a client's onboarding

1054
00:36:13,200 --> 00:36:15,800
process, and this new transcript adds a wrinkle nobody

1055
00:36:15,800 --> 00:36:17,040
had documented yet.

1056
00:36:17,040 --> 00:36:18,000
That page gets richer.

1057
00:36:18,000 --> 00:36:19,120
It doesn't get duplicated.

1058
00:36:19,120 --> 00:36:21,840
It doesn't sit next to an old version creating confusion.

1059
00:36:21,840 --> 00:36:23,520
It gets updated in place.

1060
00:36:23,520 --> 00:36:25,680
But if the source introduces something genuinely new,

1061
00:36:25,680 --> 00:36:27,640
something that doesn't map onto any page that already

1062
00:36:27,640 --> 00:36:29,400
exists, a new page gets created.

1063
00:36:29,400 --> 00:36:31,520
And whichever of those two things happens,

1064
00:36:31,520 --> 00:36:33,120
related ideas get linked.

1065
00:36:33,120 --> 00:36:35,080
The new page connects to the client it belongs to,

1066
00:36:35,080 --> 00:36:37,720
the project it touches, the policy that shaped it, nothing

1067
00:36:37,720 --> 00:36:40,600
sits alone.

1068
00:36:40,600 --> 00:36:42,360
Here's the part that matters most, though.

1069
00:36:42,360 --> 00:36:45,080
And it's the part, rag has no equivalent for it all.

1070
00:36:45,080 --> 00:36:47,480
What happens when the new information contradicts something

1071
00:36:47,480 --> 00:36:48,920
already sitting in the wiki?

1072
00:36:48,920 --> 00:36:51,160
Say the old page claims a process works one way,

1073
00:36:51,160 --> 00:36:52,960
and the new source says it changed.

1074
00:36:52,960 --> 00:36:54,960
The system doesn't just override the old claim

1075
00:36:54,960 --> 00:36:56,240
and move on like nothing happened.

1076
00:36:56,240 --> 00:36:57,120
It flags it.

1077
00:36:57,120 --> 00:36:58,760
Both versions get preserved with a note

1078
00:36:58,760 --> 00:37:00,440
that something shifted, and a signal

1079
00:37:00,440 --> 00:37:02,480
that someone should look at this before treating either

1080
00:37:02,480 --> 00:37:03,800
version as settled.

1081
00:37:03,800 --> 00:37:05,640
This flagging behavior isn't a nice to have

1082
00:37:05,640 --> 00:37:07,240
detailed buried in the mechanics.

1083
00:37:07,240 --> 00:37:09,720
It matters enormously for actual enterprise use,

1084
00:37:09,720 --> 00:37:11,840
because businesses change constantly.

1085
00:37:11,840 --> 00:37:14,000
And most of that change happens quietly.

1086
00:37:14,000 --> 00:37:17,040
In a policy update here, a strategy pivot there,

1087
00:37:17,040 --> 00:37:19,840
a decision that got reversed without anyone sending out

1088
00:37:19,840 --> 00:37:21,240
a formal memo about it.

1089
00:37:21,240 --> 00:37:23,400
When a policy update contradicts an old assumption

1090
00:37:23,400 --> 00:37:25,000
that's been floating around for months,

1091
00:37:25,000 --> 00:37:27,080
that contradiction doesn't get buried under a pile

1092
00:37:27,080 --> 00:37:29,280
of documents that all look equally current.

1093
00:37:29,280 --> 00:37:32,760
It gets surfaced directly to whoever's supposed to notice it.

1094
00:37:32,760 --> 00:37:34,200
Now hold that against what rag does

1095
00:37:34,200 --> 00:37:36,400
with the exact same situation, because the contrast

1096
00:37:36,400 --> 00:37:37,680
is where this really lands.

1097
00:37:37,680 --> 00:37:39,400
A stale document sits in your share point

1098
00:37:39,400 --> 00:37:41,480
next to the document that replaced it.

1099
00:37:41,480 --> 00:37:44,080
Someone asks a question, and rag retrieves both.

1100
00:37:44,080 --> 00:37:46,680
Not one, both, as two equally weighted chunks,

1101
00:37:46,680 --> 00:37:49,200
stitched together into an answer with absolutely no signal

1102
00:37:49,200 --> 00:37:50,960
about which one is actually current.

1103
00:37:50,960 --> 00:37:52,880
The model doesn't know one superseded the other.

1104
00:37:52,880 --> 00:37:54,880
It just sees two fragments that seem relevant

1105
00:37:54,880 --> 00:37:57,640
and does its best to synthesize something coherent

1106
00:37:57,640 --> 00:37:59,840
out of material that's quietly contradicting itself.

1107
00:37:59,840 --> 00:38:01,040
Nobody flagged anything.

1108
00:38:01,040 --> 00:38:03,320
Nobody got a signal that something needed a second look.

1109
00:38:03,320 --> 00:38:05,800
The contradiction just sits there, silently corrupting

1110
00:38:05,800 --> 00:38:07,640
whatever answer comes out the other end.

1111
00:38:07,640 --> 00:38:10,560
That difference, flagging versus silent averaging,

1112
00:38:10,560 --> 00:38:13,000
points at something bigger than a technical detail

1113
00:38:13,000 --> 00:38:14,440
about how updates get handled.

1114
00:38:14,440 --> 00:38:15,800
It's a concept worth naming directly,

1115
00:38:15,800 --> 00:38:17,480
because once you see it, you can't unsee it

1116
00:38:17,480 --> 00:38:19,600
in every rag deployment you've ever looked at.

1117
00:38:19,600 --> 00:38:22,040
Contradiction as a feature, not a bug.

1118
00:38:22,040 --> 00:38:25,400
Most systems, when they run into contradictory information,

1119
00:38:25,400 --> 00:38:26,720
treat it as noise.

1120
00:38:26,720 --> 00:38:28,680
Something to smooth over, average out,

1121
00:38:28,680 --> 00:38:30,760
resolve into a single, clean answer,

1122
00:38:30,760 --> 00:38:32,560
so nobody has to look at the mess underneath.

1123
00:38:32,560 --> 00:38:35,240
The wiki pattern does something almost nobody expects

1124
00:38:35,240 --> 00:38:36,440
the first time they see it.

1125
00:38:36,440 --> 00:38:38,480
It treats contradiction as a signal,

1126
00:38:38,480 --> 00:38:40,680
worth preserving, not a problem to hide.

1127
00:38:40,680 --> 00:38:42,600
Think about what a contradiction actually represents

1128
00:38:42,600 --> 00:38:43,760
inside a business.

1129
00:38:43,760 --> 00:38:45,120
It's rarely a mistake.

1130
00:38:45,120 --> 00:38:46,800
Most of the time, it's a policy shift somebody

1131
00:38:46,800 --> 00:38:49,040
made three months ago and never fully announced,

1132
00:38:49,040 --> 00:38:52,240
or a strategy pivot that quietly replaced last year's plan,

1133
00:38:52,240 --> 00:38:54,880
or a decision that got reversed after a client pushed back,

1134
00:38:54,880 --> 00:38:57,160
and only half the team ever heard about the reversal.

1135
00:38:57,160 --> 00:38:59,360
Contradictions in an organizational context

1136
00:38:59,360 --> 00:39:02,480
are usually just changed that hasn't finished propagating yet.

1137
00:39:02,480 --> 00:39:05,160
Here's where rag runs into a wall it can't get passed.

1138
00:39:05,160 --> 00:39:08,320
A stateless system has no way to represent this used to be true,

1139
00:39:08,320 --> 00:39:09,720
now it isn't.

1140
00:39:09,720 --> 00:39:11,200
There's no mechanism for that.

1141
00:39:11,200 --> 00:39:13,880
Rag just retrieves, whichever chunk seems more relevant

1142
00:39:13,880 --> 00:39:15,600
at the moment the question gets asked,

1143
00:39:15,600 --> 00:39:18,400
and depending on wording, embedding similarity,

1144
00:39:18,400 --> 00:39:20,400
whatever's closest in the vector space,

1145
00:39:20,400 --> 00:39:23,200
it might hand you the old policy or the new one.

1146
00:39:23,200 --> 00:39:24,840
Neither answer is technically wrong.

1147
00:39:24,840 --> 00:39:26,000
The document exists.

1148
00:39:26,000 --> 00:39:28,520
It's just that rag has no concept of existed,

1149
00:39:28,520 --> 00:39:30,120
but got superseded.

1150
00:39:30,120 --> 00:39:33,120
Every chunk is equally present, equally retrievable,

1151
00:39:33,120 --> 00:39:36,560
equally weighted, forever, regardless of whether it's still true.

1152
00:39:36,560 --> 00:39:39,200
A compiled wiki does something structurally different.

1153
00:39:39,200 --> 00:39:42,520
It can hold both the old state and the new state on the same page,

1154
00:39:42,520 --> 00:39:45,720
with a note attached explaining when the change happened and why,

1155
00:39:45,720 --> 00:39:49,600
if that reason is known, not too competing documents floating in a search index

1156
00:39:49,600 --> 00:39:51,360
with no relationship to each other.

1157
00:39:51,360 --> 00:39:53,600
One page, with history built into it.

1158
00:39:53,600 --> 00:39:58,160
The old assumption is still there, but it's marked as superseded, dated, explained.

1159
00:39:58,160 --> 00:40:00,960
Nothing gets erased, nothing gets hidden, it just gets sequenced.

1160
00:40:00,960 --> 00:40:03,560
This is the actual difference between an assistant that answers

1161
00:40:03,560 --> 00:40:06,360
and an assistant that understands your organization's history.

1162
00:40:06,360 --> 00:40:09,160
Answering means producing something plausible when asked.

1163
00:40:09,160 --> 00:40:10,960
Understanding history means knowing that

1164
00:40:10,960 --> 00:40:12,880
what's true today used to be something else

1165
00:40:12,880 --> 00:40:15,800
and knowing why the shift happened and being able to tell you that story

1166
00:40:15,800 --> 00:40:18,680
instead of just picking aside and hoping it's the right one.

1167
00:40:18,680 --> 00:40:20,600
Understanding history like that is one thing,

1168
00:40:20,600 --> 00:40:23,080
making it usable across an entire enterprise,

1169
00:40:23,080 --> 00:40:25,680
not just one project page or one client file,

1170
00:40:25,680 --> 00:40:27,920
but thousands of them, spanning years,

1171
00:40:27,920 --> 00:40:30,640
across every team, is an entirely different problem.

1172
00:40:30,640 --> 00:40:32,840
And it's the one worth turning to next.

1173
00:40:32,840 --> 00:40:34,080
The scale question.

1174
00:40:34,080 --> 00:40:35,520
So let's be honest about the limit here,

1175
00:40:35,520 --> 00:40:37,840
because this pattern doesn't scale the way you'd wanted to,

1176
00:40:37,840 --> 00:40:39,520
just by wishing hard enough.

1177
00:40:39,520 --> 00:40:42,360
Personal scale wikis, the kind cup he's been describing,

1178
00:40:42,360 --> 00:40:44,920
work well around 100 or so source documents.

1179
00:40:44,920 --> 00:40:48,040
That's the range where one person or one small compiling agent

1180
00:40:48,040 --> 00:40:50,480
can keep everything coherent, cross-linked and current

1181
00:40:50,480 --> 00:40:53,400
without the whole thing turning into its own management problem.

1182
00:40:53,400 --> 00:40:55,040
Your enterprise content estate is not that.

1183
00:40:55,040 --> 00:40:57,320
It's not 100 documents, it's not 1,000.

1184
00:40:57,320 --> 00:41:00,280
Most organizations are sitting on content spread across years,

1185
00:41:00,280 --> 00:41:03,640
teams and systems that nobody's fully audited, let alone compiled.

1186
00:41:03,640 --> 00:41:04,920
So where does that leave the pattern

1187
00:41:04,920 --> 00:41:06,760
once your past personal scale?

1188
00:41:06,760 --> 00:41:09,640
The honest answer is pointing towards 2026 outlooks

1189
00:41:09,640 --> 00:41:11,440
that keep landing on the same conclusion.

1190
00:41:11,440 --> 00:41:13,920
Roughly 75% of enterprise applications

1191
00:41:13,920 --> 00:41:15,960
are moving toward hybrid architectures,

1192
00:41:15,960 --> 00:41:18,200
ones that combine wiki style compilation

1193
00:41:18,200 --> 00:41:21,160
with rag and agentic search, rather than picking one

1194
00:41:21,160 --> 00:41:22,400
and walking away from the other.

1195
00:41:22,400 --> 00:41:24,520
That's not a compromise bone out of indecision.

1196
00:41:24,520 --> 00:41:26,080
It's a recognition that these two systems

1197
00:41:26,080 --> 00:41:29,040
are solving different problems, and neither one covers the other's job.

1198
00:41:29,040 --> 00:41:30,800
Here's why hybrid actually makes sense,

1199
00:41:30,800 --> 00:41:33,320
once you stop thinking of this as a competition.

1200
00:41:33,320 --> 00:41:36,680
Rag remains the default for large, relatively static repositories,

1201
00:41:36,680 --> 00:41:39,200
the sprawling stuff that doesn't get asked about constantly,

1202
00:41:39,200 --> 00:41:42,120
but still needs to be searchable when someone does need it.

1203
00:41:42,120 --> 00:41:44,080
Wiki's excel somewhere else entirely,

1204
00:41:44,080 --> 00:41:47,160
the stable, high value knowledge that gets asked about again and again,

1205
00:41:47,160 --> 00:41:48,680
the stuff worth the effort of compiling

1206
00:41:48,680 --> 00:41:50,440
because people keep coming back to it.

1207
00:41:50,440 --> 00:41:52,920
Frame it as a division of labor because that's really what it is.

1208
00:41:52,920 --> 00:41:54,880
Rag handles breadth, it covers the long tail,

1209
00:41:54,880 --> 00:41:56,480
everything that exists, but doesn't get touched

1210
00:41:56,480 --> 00:41:58,840
often enough to justify compiling it by hand.

1211
00:41:58,840 --> 00:42:01,080
The wiki handles depths on the narrow set of things

1212
00:42:01,080 --> 00:42:04,840
that actually matter enough to deserve careful, maintained synthesis.

1213
00:42:04,840 --> 00:42:06,600
Neither one is trying to do the other's job.

1214
00:42:06,600 --> 00:42:07,640
That's the whole point.

1215
00:42:07,640 --> 00:42:09,440
Worth a cost reality check here, too,

1216
00:42:09,440 --> 00:42:11,600
because none of this is free either direction.

1217
00:42:11,600 --> 00:42:15,440
MVP Rag systems typically run 15 to $40,000 to stand up,

1218
00:42:15,440 --> 00:42:17,080
covering the vector infrastructure,

1219
00:42:17,080 --> 00:42:19,360
the embedding pipeline, the retrieval tooling,

1220
00:42:19,360 --> 00:42:20,720
the wiki layer isn't free,

1221
00:42:20,720 --> 00:42:22,320
but its cost doesn't show up the same way.

1222
00:42:22,320 --> 00:42:24,680
It shows up as ongoing maintenance, not infrastructure.

1223
00:42:24,680 --> 00:42:25,960
You're not buying a bigger system.

1224
00:42:25,960 --> 00:42:27,760
You're paying for the compiling and upkeep

1225
00:42:27,760 --> 00:42:31,360
that keeps the wiki accurate as your organization keeps changing.

1226
00:42:31,360 --> 00:42:32,640
Once you see it that way,

1227
00:42:32,640 --> 00:42:34,960
breadth handled one way, depth handled another,

1228
00:42:34,960 --> 00:42:37,640
the question stops being which architecture do we pick,

1229
00:42:37,640 --> 00:42:39,640
and starts being something much more practical.

1230
00:42:39,640 --> 00:42:41,600
How does this hybrid framing actually change

1231
00:42:41,600 --> 00:42:45,960
what a rollout looks like inside a real Microsoft 365 environment?

1232
00:42:45,960 --> 00:42:47,760
What hybrid looks like for co-pilot?

1233
00:42:47,760 --> 00:42:49,880
So picture what this actually looks like once it's running,

1234
00:42:49,880 --> 00:42:52,280
not as a diagram, but as something sitting inside your tenant

1235
00:42:52,280 --> 00:42:54,040
doing two different jobs at once.

1236
00:42:54,040 --> 00:42:55,360
Co-pilot with two tiers,

1237
00:42:55,360 --> 00:42:57,680
one tier is the broad, graph grounded retrieval

1238
00:42:57,680 --> 00:42:58,800
you already have today,

1239
00:42:58,800 --> 00:43:00,560
reaching across the long tail of everything

1240
00:43:00,560 --> 00:43:02,840
in SharePoint teams exchange all of it,

1241
00:43:02,840 --> 00:43:04,880
searchable the way it's always been searchable.

1242
00:43:04,880 --> 00:43:06,720
The other tier is a compiled wiki layer,

1243
00:43:06,720 --> 00:43:08,200
sitting quietly underneath,

1244
00:43:08,200 --> 00:43:10,920
built specifically for the topics your organization asks about

1245
00:43:10,920 --> 00:43:11,960
again and again,

1246
00:43:11,960 --> 00:43:14,040
and that second tier only earns its keep

1247
00:43:14,040 --> 00:43:15,440
on specific kinds of content.

1248
00:43:15,440 --> 00:43:17,840
So it's worth naming what actually belongs there,

1249
00:43:17,840 --> 00:43:18,840
onboarding knowledge,

1250
00:43:18,840 --> 00:43:21,000
the stuff every new hire needs and every team ends up

1251
00:43:21,000 --> 00:43:24,360
re-explaning verbally because nobody trusts the search results.

1252
00:43:24,360 --> 00:43:26,520
Product positioning, the language your organization

1253
00:43:26,520 --> 00:43:27,720
has already agreed on,

1254
00:43:27,720 --> 00:43:29,640
that keeps drifting because three different people

1255
00:43:29,640 --> 00:43:31,240
describe it three different ways,

1256
00:43:31,240 --> 00:43:32,600
depending on who's asked.

1257
00:43:32,600 --> 00:43:34,040
Recurring client questions,

1258
00:43:34,040 --> 00:43:36,280
the ones that show up every quarter with new phrasing,

1259
00:43:36,280 --> 00:43:38,080
but the same underlying ask.

1260
00:43:38,080 --> 00:43:39,480
Compliance interpretations,

1261
00:43:39,480 --> 00:43:42,400
where getting it slightly wrong isn't a minor inconvenience.

1262
00:43:42,400 --> 00:43:44,240
Project histories, the decisions and reversals

1263
00:43:44,240 --> 00:43:45,960
that shaped how something got built,

1264
00:43:45,960 --> 00:43:48,280
the kind of context a new team member has no way of finding

1265
00:43:48,280 --> 00:43:50,560
without asking someone who's been there for years.

1266
00:43:50,560 --> 00:43:52,000
Notice what all of those have in common.

1267
00:43:52,000 --> 00:43:53,240
They're exactly the categories

1268
00:43:53,240 --> 00:43:55,440
where consistency and trust matter most,

1269
00:43:55,440 --> 00:43:58,280
the ones we spent real time on earlier in this episode,

1270
00:43:58,280 --> 00:43:59,520
and they're exactly the categories

1271
00:43:59,520 --> 00:44:01,800
where redundant recreation waste the most time,

1272
00:44:01,800 --> 00:44:03,600
people rebuilding something that already exists

1273
00:44:03,600 --> 00:44:04,920
because search let them down once

1274
00:44:04,920 --> 00:44:06,000
and they stop trusting it.

1275
00:44:06,000 --> 00:44:07,040
That's not a coincidence.

1276
00:44:07,040 --> 00:44:09,240
Those two lists overlap almost completely,

1277
00:44:09,240 --> 00:44:12,000
which is precisely why this is where compiling earns its cost

1278
00:44:12,000 --> 00:44:14,600
instead of just adding complexity for its own sake.

1279
00:44:14,600 --> 00:44:16,880
Here's the mechanic and it's worth describing plainly,

1280
00:44:16,880 --> 00:44:19,480
instead of dressing it up as something more dramatic than it is.

1281
00:44:19,480 --> 00:44:21,560
An agent process runs periodically,

1282
00:44:21,560 --> 00:44:23,240
not constantly, not on every query,

1283
00:44:23,240 --> 00:44:24,320
just on a schedule,

1284
00:44:24,320 --> 00:44:26,440
reading through relevant, share point libraries

1285
00:44:26,440 --> 00:44:28,560
and teams content tied to whichever domain

1286
00:44:28,560 --> 00:44:29,880
it's responsible for.

1287
00:44:29,880 --> 00:44:32,240
It compiles what it finds into structured pages,

1288
00:44:32,240 --> 00:44:33,640
updating what already exists,

1289
00:44:33,640 --> 00:44:36,760
creating new pages where something genuinely new shows up.

1290
00:44:36,760 --> 00:44:39,320
And then when someone asks Copilot a question

1291
00:44:39,320 --> 00:44:41,440
that touches one of those compiled topics,

1292
00:44:41,440 --> 00:44:44,280
Copilot's studio surfaces those pages preferentially,

1293
00:44:44,280 --> 00:44:46,720
not exclusively, not instead of graph grounding entirely,

1294
00:44:46,720 --> 00:44:48,320
just weighted higher, trusted more,

1295
00:44:48,320 --> 00:44:49,920
pulled first when the topic matches.

1296
00:44:49,920 --> 00:44:51,440
Worth being clear about what this isn't

1297
00:44:51,440 --> 00:44:53,560
because it's easy to oversell a pattern like this

1298
00:44:53,560 --> 00:44:56,080
once you've spent an episode building the case for it.

1299
00:44:56,080 --> 00:44:57,800
This isn't a rip and replace of Copilot.

1300
00:44:57,800 --> 00:44:59,680
You're not tearing out graph grounding

1301
00:44:59,680 --> 00:45:01,000
and swapping in something else.

1302
00:45:01,000 --> 00:45:02,000
It's an added layer.

1303
00:45:02,000 --> 00:45:04,160
Sitting alongside what already exists,

1304
00:45:04,160 --> 00:45:06,920
changing what grounding actually means specifically

1305
00:45:06,920 --> 00:45:09,640
for the topics your organization keeps asking about.

1306
00:45:09,640 --> 00:45:11,840
Everything else still works the way it always has.

1307
00:45:11,840 --> 00:45:13,640
The long tail is still there, still searchable,

1308
00:45:13,640 --> 00:45:15,280
still handle the old way,

1309
00:45:15,280 --> 00:45:17,960
because most of your content doesn't need this kind of treatment

1310
00:45:17,960 --> 00:45:19,040
and never will.

1311
00:45:19,040 --> 00:45:20,120
But none of this runs itself.

1312
00:45:20,120 --> 00:45:21,560
Somebody has to decide what counts

1313
00:45:21,560 --> 00:45:24,400
as a recurring topic worth compiling in the first place.

1314
00:45:24,400 --> 00:45:26,840
Somebody has to define how those pages get formatted,

1315
00:45:26,840 --> 00:45:28,040
what triggers an update,

1316
00:45:28,040 --> 00:45:29,560
when a contradiction gets flagged

1317
00:45:29,560 --> 00:45:31,360
instead of quietly resolved.

1318
00:45:31,360 --> 00:45:32,480
And that's a governance question.

1319
00:45:32,480 --> 00:45:34,480
One IT leader's can't really skip

1320
00:45:34,480 --> 00:45:37,360
once they've decided this layer is worth building at all.

1321
00:45:37,360 --> 00:45:39,360
The governance layer nobody's talking about.

1322
00:45:39,360 --> 00:45:40,600
Somebody has to own this.

1323
00:45:40,600 --> 00:45:42,920
Not in a vague, it will figure it out sense,

1324
00:45:42,920 --> 00:45:45,840
but a real role with a real name attached to it.

1325
00:45:45,840 --> 00:45:47,480
Someone has to own the schema,

1326
00:45:47,480 --> 00:45:49,400
the actual rules for what gets compiled

1327
00:45:49,400 --> 00:45:51,600
and what doesn't, how a page gets formatted,

1328
00:45:51,600 --> 00:45:54,120
so it matches every other page in the wiki.

1329
00:45:54,120 --> 00:45:56,680
And critically, when a contradiction gets escalated

1330
00:45:56,680 --> 00:45:59,440
to an actual human, instead of just sitting flagged

1331
00:45:59,440 --> 00:46:01,080
in a system nobody's watching.

1332
00:46:01,080 --> 00:46:02,800
Without that owner, the schema drifts,

1333
00:46:02,800 --> 00:46:05,480
the pages get inconsistent and the whole thing slowly turns

1334
00:46:05,480 --> 00:46:08,120
into exactly the mess it was supposed to replace.

1335
00:46:08,120 --> 00:46:10,320
Here's the good news, and it's worth saying plainly

1336
00:46:10,320 --> 00:46:12,680
because this part tends to get overcomplicated.

1337
00:46:12,680 --> 00:46:14,440
This role isn't new, it maps directly

1338
00:46:14,440 --> 00:46:16,040
onto jobs that already exist inside

1339
00:46:16,040 --> 00:46:18,320
most Microsoft 365 environments.

1340
00:46:18,320 --> 00:46:19,800
Information architects already think

1341
00:46:19,800 --> 00:46:21,560
about how content gets structured.

1342
00:46:21,560 --> 00:46:24,000
Records managers already think about lifecycle,

1343
00:46:24,000 --> 00:46:25,760
about what stays, what gets retired,

1344
00:46:25,760 --> 00:46:27,160
what needs a review cycle.

1345
00:46:27,160 --> 00:46:29,320
Content owners already carry responsibility

1346
00:46:29,320 --> 00:46:30,520
for whether something's accurate.

1347
00:46:30,520 --> 00:46:32,280
None of these people need a new job title.

1348
00:46:32,280 --> 00:46:34,160
They need this added to a job they're already doing

1349
00:46:34,160 --> 00:46:36,400
because the skill underneath it is identical.

1350
00:46:36,400 --> 00:46:39,640
Deciding what belongs where and who's accountable when it's wrong.

1351
00:46:39,640 --> 00:46:41,400
There's already a preview of where this is heading

1352
00:46:41,400 --> 00:46:43,120
and it's sitting inside Pervue right now.

1353
00:46:43,120 --> 00:46:44,840
The data loss prevention controls,

1354
00:46:44,840 --> 00:46:46,920
Microsoft built for co-pilot's web searches,

1355
00:46:46,920 --> 00:46:49,880
the ones that stop sensitive data from leaking into a prompt,

1356
00:46:49,880 --> 00:46:52,240
while still letting co-pilot ground its answer

1357
00:46:52,240 --> 00:46:54,640
in internal sources, that's the same instinct

1358
00:46:54,640 --> 00:46:56,760
applied to a narrower problem.

1359
00:46:56,760 --> 00:46:59,640
Guard rails around what an AI-driven system is allowed to see

1360
00:46:59,640 --> 00:47:01,800
and by extension, what it's allowed to compile.

1361
00:47:01,800 --> 00:47:03,400
Once you're running an agent that reads

1362
00:47:03,400 --> 00:47:06,240
across your SharePoint libraries and writes structured pages,

1363
00:47:06,240 --> 00:47:08,000
you need the same kind of thinking

1364
00:47:08,000 --> 00:47:10,480
just aimed at compilation instead of search.

1365
00:47:10,480 --> 00:47:13,280
So frame this the right way because how you frame it changes

1366
00:47:13,280 --> 00:47:14,840
whether it actually gets funded.

1367
00:47:14,840 --> 00:47:18,320
This isn't a brand new discipline landing on IT's desk out of nowhere.

1368
00:47:18,320 --> 00:47:19,840
It's an extension of governance work

1369
00:47:19,840 --> 00:47:22,120
that's probably already underway in some form.

1370
00:47:22,120 --> 00:47:24,760
If your organization already has content lifecycle policies,

1371
00:47:24,760 --> 00:47:26,600
retention rules, ownership models,

1372
00:47:26,600 --> 00:47:28,880
you already have the starting point for a Wiki schema.

1373
00:47:28,880 --> 00:47:30,440
You're not building from zero.

1374
00:47:30,440 --> 00:47:32,320
You're extending something that exists.

1375
00:47:32,320 --> 00:47:34,640
And this is really the dividing line between organizations

1376
00:47:34,640 --> 00:47:36,320
that get real value out of this pattern

1377
00:47:36,320 --> 00:47:38,600
and organizations that build something impressive looking

1378
00:47:38,600 --> 00:47:40,880
and then watch it decay within a year.

1379
00:47:40,880 --> 00:47:42,920
The ones that benefit most treat this

1380
00:47:42,920 --> 00:47:44,960
as a knowledge management initiative

1381
00:47:44,960 --> 00:47:46,880
that happens to have an AI component,

1382
00:47:46,880 --> 00:47:49,680
not an AI initiative that happens to touch knowledge,

1383
00:47:49,680 --> 00:47:52,200
that ordering matters more than it sounds like it should.

1384
00:47:52,200 --> 00:47:55,200
One puts governance first and lets the technology serve it.

1385
00:47:55,200 --> 00:47:56,720
The other bolts governance on afterward

1386
00:47:56,720 --> 00:47:58,360
once something's already gone wrong.

1387
00:47:58,360 --> 00:48:00,720
With the mechanics covered and the governance question

1388
00:48:00,720 --> 00:48:02,240
at least named honestly,

1389
00:48:02,240 --> 00:48:04,040
it's worth stepping back from the how

1390
00:48:04,040 --> 00:48:07,760
and naming what's actually different here underneath all of it.

1391
00:48:07,760 --> 00:48:09,080
What actually changes?

1392
00:48:09,080 --> 00:48:12,520
So here's the shift stated as plainly as it deserves to be stated.

1393
00:48:12,520 --> 00:48:15,280
Copilot stops being a search box with better manners,

1394
00:48:15,280 --> 00:48:18,200
one that phrases things politely and summarizes nicely,

1395
00:48:18,200 --> 00:48:19,920
but still starts from zero every time

1396
00:48:19,920 --> 00:48:22,280
and starts becoming a system that actually has

1397
00:48:22,280 --> 00:48:23,720
somewhere to put what it learns.

1398
00:48:23,720 --> 00:48:25,680
That's the whole difference, not a smarter model,

1399
00:48:25,680 --> 00:48:27,880
not a better prompt, a place for knowledge to live

1400
00:48:27,880 --> 00:48:29,760
once it's been figured out instead of a mechanism

1401
00:48:29,760 --> 00:48:31,760
that refigures it out on command forever.

1402
00:48:31,760 --> 00:48:33,560
And that difference doesn't show up on day one.

1403
00:48:33,560 --> 00:48:34,960
It shows up over time,

1404
00:48:34,960 --> 00:48:38,560
which is exactly why so many organizations miss it during a pilot.

1405
00:48:38,560 --> 00:48:40,960
A stateless system is the same tool on day 500

1406
00:48:40,960 --> 00:48:42,400
as it was on day one.

1407
00:48:42,400 --> 00:48:45,480
Ask at the same category of question a year from now

1408
00:48:45,480 --> 00:48:47,320
and it goes through the identical motions

1409
00:48:47,320 --> 00:48:48,800
it went through in week one,

1410
00:48:48,800 --> 00:48:53,160
search, chunk, retrieve, synthesize, forget.

1411
00:48:53,160 --> 00:48:55,160
Nothing about that process improves with age,

1412
00:48:55,160 --> 00:48:57,600
it doesn't get faster, it doesn't get more accurate,

1413
00:48:57,600 --> 00:48:58,600
it just repeats.

1414
00:48:58,600 --> 00:49:00,960
A compiled system by contrast is measurably different

1415
00:49:00,960 --> 00:49:04,120
a year in because every project page, every client history,

1416
00:49:04,120 --> 00:49:06,200
every resolved contradiction from month three

1417
00:49:06,200 --> 00:49:08,320
is still sitting there in month 14,

1418
00:49:08,320 --> 00:49:10,840
ready to be built on instead of rediscovered.

1419
00:49:10,840 --> 00:49:12,120
That's not a theoretical claim.

1420
00:49:12,120 --> 00:49:14,600
It shows up in the numbers we've already touched on earlier

1421
00:49:14,600 --> 00:49:16,640
in this episode around structured knowledge

1422
00:49:16,640 --> 00:49:19,520
and it's worth tying back to the sharpest version of them.

1423
00:49:19,520 --> 00:49:22,120
Enterprise knowledge graphs and structured content approaches

1424
00:49:22,120 --> 00:49:25,720
have delivered results like 287% ROI over three years

1425
00:49:25,720 --> 00:49:28,040
with multi million dollar annual benefits reported

1426
00:49:28,040 --> 00:49:29,440
in some organizations.

1427
00:49:29,440 --> 00:49:30,960
Those aren't numbers from a system

1428
00:49:30,960 --> 00:49:32,280
that got asked more questions,

1429
00:49:32,280 --> 00:49:34,320
they're numbers from a system that got smarter

1430
00:49:34,320 --> 00:49:35,680
at answering the same questions

1431
00:49:35,680 --> 00:49:37,840
because the underlying knowledge kept compounding

1432
00:49:37,840 --> 00:49:39,080
instead of resetting.

1433
00:49:39,080 --> 00:49:41,440
And that's really the mechanism worth naming directly

1434
00:49:41,440 --> 00:49:43,720
because it's easy to miss if you're only looking

1435
00:49:43,720 --> 00:49:45,160
at usage dashboards.

1436
00:49:45,160 --> 00:49:47,120
Compounding knowledge, not repeated search,

1437
00:49:47,120 --> 00:49:49,160
is what actually saves time at scale.

1438
00:49:49,160 --> 00:49:52,080
Repeated search just means people are using the tool more often.

1439
00:49:52,080 --> 00:49:54,600
It says nothing about whether the tool is getting better.

1440
00:49:54,600 --> 00:49:56,600
Compounding knowledge means the thousandth question

1441
00:49:56,600 --> 00:49:58,720
benefits from everything the system learned

1442
00:49:58,720 --> 00:50:00,680
answering the first 999.

1443
00:50:00,680 --> 00:50:02,160
That's a completely different kind of value

1444
00:50:02,160 --> 00:50:04,120
and it's the kind drag structurally can't produce

1445
00:50:04,120 --> 00:50:05,440
no matter how much you use it.

1446
00:50:05,440 --> 00:50:07,000
So say the quiet part plainly

1447
00:50:07,000 --> 00:50:08,800
because it's the sentence this whole episode

1448
00:50:08,800 --> 00:50:10,160
has been building toward.

1449
00:50:10,160 --> 00:50:12,760
An assistant that gets smarter is a fundamentally different

1450
00:50:12,760 --> 00:50:15,080
asset than one that just gets asked more questions.

1451
00:50:15,080 --> 00:50:17,720
One compounds the other just accumulates traffic.

1452
00:50:17,720 --> 00:50:20,240
Confusing the two is exactly how organizations end up

1453
00:50:20,240 --> 00:50:21,920
with impressive adoption metrics

1454
00:50:21,920 --> 00:50:24,320
and a tool that still can't hold a coherent memory

1455
00:50:24,320 --> 00:50:26,640
of its own history, which naturally raises

1456
00:50:26,640 --> 00:50:28,440
the only question left worth answering.

1457
00:50:28,440 --> 00:50:29,760
Not whether this is worth doing

1458
00:50:29,760 --> 00:50:31,680
but where an organization actually starts,

1459
00:50:31,680 --> 00:50:33,680
small and deliberate instead of trying to compile

1460
00:50:33,680 --> 00:50:34,960
everything at once.

1461
00:50:34,960 --> 00:50:36,120
What if you start today?

1462
00:50:36,120 --> 00:50:38,120
So pick one thing, not the whole knowledge estate,

1463
00:50:38,120 --> 00:50:41,120
not every department, one recurring high value domain.

1464
00:50:41,120 --> 00:50:42,480
Maybe it's a single product line

1465
00:50:42,480 --> 00:50:44,120
that support keeps asking about.

1466
00:50:44,120 --> 00:50:46,000
Maybe it's one recurring client type,

1467
00:50:46,000 --> 00:50:47,640
the kind of account whose questions repeat

1468
00:50:47,640 --> 00:50:49,480
every quarter with new phrasing.

1469
00:50:49,480 --> 00:50:51,240
Maybe it's a single compliance area

1470
00:50:51,240 --> 00:50:53,440
where getting the answer wrong actually costs something.

1471
00:50:53,440 --> 00:50:55,320
The domain doesn't matter as much as the discipline

1472
00:50:55,320 --> 00:50:56,360
of picking just one.

1473
00:50:56,360 --> 00:50:58,800
Here's what a first 90 days actually looks like

1474
00:50:58,800 --> 00:51:01,120
without dressing it up as bigger than it is.

1475
00:51:01,120 --> 00:51:03,080
Start by identifying the raw sources,

1476
00:51:03,080 --> 00:51:04,720
whatever SharePoint libraries,

1477
00:51:04,720 --> 00:51:06,600
Teams, Threads and old policy docs

1478
00:51:06,600 --> 00:51:09,320
already touch this domain, then define a schema

1479
00:51:09,320 --> 00:51:11,080
and keep it lightweight on purpose.

1480
00:51:11,080 --> 00:51:13,400
Not a governance framework spanning 50 pages,

1481
00:51:13,400 --> 00:51:15,480
just enough rules to say what a page looks like

1482
00:51:15,480 --> 00:51:17,200
and how contradictions get flagged,

1483
00:51:17,200 --> 00:51:19,480
run an initial compilation pass against that schema.

1484
00:51:19,480 --> 00:51:21,320
Then test it against real questions,

1485
00:51:21,320 --> 00:51:23,000
phrase the way actual people phrase them,

1486
00:51:23,000 --> 00:51:25,240
not the clean version you'd write in a demo script.

1487
00:51:25,240 --> 00:51:27,720
Picture the end of that pilot next to where you started.

1488
00:51:27,720 --> 00:51:30,600
Before, co-pilot retrieves five slightly different chunks

1489
00:51:30,600 --> 00:51:32,640
about the same policy and depending on

1490
00:51:32,640 --> 00:51:34,240
how someone phrased the question,

1491
00:51:34,240 --> 00:51:36,240
they get a different one of those five.

1492
00:51:36,240 --> 00:51:39,000
After, it answers from one compiled current page

1493
00:51:39,000 --> 00:51:41,120
every time, no matter how the question gets asked.

1494
00:51:41,120 --> 00:51:43,920
That's the whole test, not whether the answer sounds smart,

1495
00:51:43,920 --> 00:51:45,560
whether it's the same answer twice,

1496
00:51:45,560 --> 00:51:47,200
worth connecting this back to the research

1497
00:51:47,200 --> 00:51:48,600
we've been leaning on all episode

1498
00:51:48,600 --> 00:51:50,720
because it matters that this isn't theoretical.

1499
00:51:50,720 --> 00:51:53,560
The organization's reporting a 25 to 40% reduction

1500
00:51:53,560 --> 00:51:55,000
in redundant knowledge creation,

1501
00:51:55,000 --> 00:51:57,280
the ones seeing meaningfully lower error rates.

1502
00:51:57,280 --> 00:51:59,440
They didn't get there by rebuilding everything at once.

1503
00:51:59,440 --> 00:52:02,040
They got their one domain at a time, exactly like this.

1504
00:52:02,040 --> 00:52:05,160
Small-bounded, tested against real use before expanding,

1505
00:52:05,160 --> 00:52:06,840
and there's a thread here worth planting

1506
00:52:06,840 --> 00:52:08,680
without pulling on a too hard right now.

1507
00:52:08,680 --> 00:52:09,840
The moment you run this pilot,

1508
00:52:09,840 --> 00:52:12,680
you're going to hit real questions about who owns the schema,

1509
00:52:12,680 --> 00:52:14,840
who decides when a contradiction gets escalated,

1510
00:52:14,840 --> 00:52:17,160
who's accountable when a page goes stale.

1511
00:52:17,160 --> 00:52:18,120
That's governance.

1512
00:52:18,120 --> 00:52:21,240
And it's worth its own conversation rather than a rushed afterthought here.

1513
00:52:21,240 --> 00:52:23,760
So that's the technical case and the strategic case,

1514
00:52:23,760 --> 00:52:26,880
sitting next to each other, time to bring them together.

1515
00:52:26,880 --> 00:52:28,080
Key transformation.

1516
00:52:28,080 --> 00:52:31,320
Here's the transformation, stated as plainly as it deserves.

1517
00:52:31,320 --> 00:52:34,320
Copilot's value ceiling isn't set by the model underneath it.

1518
00:52:34,320 --> 00:52:37,440
It's set by whether your organization's knowledge gets compiled

1519
00:52:37,440 --> 00:52:40,120
or just searched over and over forever.

1520
00:52:40,120 --> 00:52:43,320
Rags will keep giving you fast, plausible, forgettable answers.

1521
00:52:43,320 --> 00:52:44,560
So that's what it's built to do,

1522
00:52:44,560 --> 00:52:45,720
and it'll keep doing it well.

1523
00:52:45,720 --> 00:52:48,480
A wiki layer gives you something structurally different,

1524
00:52:48,480 --> 00:52:51,880
an assistant that remembers what your organization has already figured out

1525
00:52:51,880 --> 00:52:54,800
instead of refiguring it out every time someone asks.

1526
00:52:54,800 --> 00:52:57,400
And here's the part worth sitting with on your way out of this episode.

1527
00:52:57,400 --> 00:53:00,920
The tools inside Microsoft 365 to build this exist today.

1528
00:53:00,920 --> 00:53:01,720
Connectors.

1529
00:53:01,720 --> 00:53:02,960
Scheduled flows.

1530
00:53:02,960 --> 00:53:04,320
Copilot studio agents.

1531
00:53:04,320 --> 00:53:05,800
All of it's already sitting in your tenant.

1532
00:53:05,800 --> 00:53:08,040
They're just not assembled this way by default.

1533
00:53:08,040 --> 00:53:09,720
So here's your challenge for this month.

1534
00:53:09,720 --> 00:53:11,760
Pick one recurring knowledge domain,

1535
00:53:11,760 --> 00:53:14,880
and go find out honestly whether Copilot is searching it

1536
00:53:14,880 --> 00:53:17,600
or actually compiling it, not what the vendor deck says,

1537
00:53:17,600 --> 00:53:19,280
what actually happens when someone asks.

1538
00:53:19,280 --> 00:53:21,920
Here's the test, and it's simple enough to run this week.

1539
00:53:21,920 --> 00:53:24,320
Ask the same question two or three different ways.

1540
00:53:24,320 --> 00:53:26,680
If the answer changes depending on how it's phrased,

1541
00:53:26,680 --> 00:53:29,480
you've found a retrieval problem, not a copilot problem.

1542
00:53:29,480 --> 00:53:31,280
That distinction is the whole episode,

1543
00:53:31,280 --> 00:53:33,240
boiled down to one homework assignment.

1544
00:53:33,240 --> 00:53:35,720
If this changed how you think about what's actually running

1545
00:53:35,720 --> 00:53:37,280
under your copilot deployment,

1546
00:53:37,280 --> 00:53:39,640
subscribe to M365FM podcast.

1547
00:53:39,640 --> 00:53:42,000
Next time we're going deeper into the governance side of this,

1548
00:53:42,000 --> 00:53:44,120
who owns the schema, who owns the wiki,

1549
00:53:44,120 --> 00:53:46,240
and what happens when that ownership isn't clear.

1550
00:53:46,240 --> 00:53:49,360
And if you want more people, more IT pros, more decision makers,

1551
00:53:49,360 --> 00:53:52,520
to run into this conversation about fixing Copilot's memory problem,

1552
00:53:52,520 --> 00:53:53,520
leave a review.

1553
00:53:53,520 --> 00:53:56,200
It's genuinely how this finds the people who need it.

1554
00:53:56,200 --> 00:53:59,000
One last thing, connect with me, Mirko Peters on LinkedIn.

1555
00:53:59,000 --> 00:54:01,800
I want to know what recurring knowledge domain you'd compile first

1556
00:54:01,800 --> 00:54:04,480
and what you think the next episode should cover.