IoT Hub Message Routing vs Event Grid — Why Telemetry and Events Are Not the Same Problem
Key Takeaways
- Confusing continuous telemetry with discrete events creates noisy workflows, incomplete production histories, and systems that react without sufficient context.
- IoT Hub Message Routing acts as the data plane for high-volume telemetry, preserving the chronological sequence of measurements needed for operational history and analytics.
- Azure Event Grid is designed for publish-and-subscribe notifications about state changes—such as device connections or threshold alerts—that require an immediate reaction.
- Treating every telemetry reading as an event triggers unnecessary workflows, whereas burying critical connectivity changes in a raw data pipeline leaves teams uninformed.
- Neither IoT Hub nor Event Grid inherently understands business context; combining message buses with Manufacturing Execution Systems (MES) and asset models is essential for accurate production insights.
A temperature reading, a vibration sample, a machine cycle, and a device disconnect may all be described as events, but they do not represent the same architectural problem. In industrial IoT and manufacturing environments, confusing continuous telemetry with discrete events can create noisy workflows, incomplete production histories, unnecessary processing, and systems that react without enough context.
In this episode, we break down the architectural difference between Azure IoT Hub Message Routing and Azure Event Grid by following a realistic manufacturing scenario. A press line continuously sends temperature, vibration, cycle count, energy consumption, and machine-state information through an industrial gateway. Those measurements create an operational history that engineers, data teams, maintenance teams, and production systems may need to analyze later. A device disconnect is different because it represents a change that may require another system or person to react.
TELEMETRY IS A RECORD, NOT AN ALERT
Telemetry represents repeated measurements over time. A single temperature value or vibration measurement usually tells you very little on its own. The real information exists in the sequence: how quickly values changed, what the machine was doing at that moment, what happened before a stop, whether measurements disappeared during a network interruption, and whether the same behavior appeared in earlier production runs.
Typical industrial telemetry includes:
• Temperature, vibration, pressure, energy consumption, and current draw
• Machine states such as running, idle, stopped, or faulted
• Cycle counts, production counters, and process measurements
• Source timestamps, device identifiers, sequence numbers, and correlation information
Those records may later support condition monitoring, quality investigations, energy analysis, OEE calculations, Microsoft Fabric analytics, Power BI reporting, and production optimization. That is why telemetry needs retention, replay, duplicate handling, independent consumers, and a reliable way to reconstruct the production timeline.
IOT HUB MESSAGE ROUTING AS THE TELEMETRY DATA PLANE
Azure IoT Hub provides the controlled device-to-cloud boundary. Devices and gateways authenticate with their own identities, send device-to-cloud messages, maintain device-management state, and can participate in controlled cloud-to-device communication.
Once a telemetry message reaches IoT Hub, Message Routing determines where that data should go. Routing can inspect message properties, system properties, parts of the message body, and device twin information. This makes it possible to separate production telemetry, energy measurements, diagnostics, or other message classes before they reach downstream consumers.
A common architecture might look like this: Industrial gateway → Azure IoT Hub → Message Routing → Event Hubs or Storage → Processing → Microsoft Fabric.
One consumer may perform near-real-time analysis while another keeps a raw archive. A third consumer may prepare curated operational data for Microsoft Fabric. Each consumer can work independently without turning every telemetry reading into a workflow invocation.
WHY ORDERING MATTERS
Industrial telemetry is particularly sensitive to sequence. Imagine a machine reporting that it entered a running state, then transmitting several cycle counts, followed by a process deviation and finally a stopped state. If those records are reconstructed incorrectly, a downstream system could conclude that the machine produced parts while stopped or that a process deviation happened after production had already ended.
The same problem affects downtime calculations, OEE, production counts, and condition monitoring. You therefore need to distinguish between when the source observed something, when IoT Hub received the message, and when a downstream system processed it.
A robust telemetry architecture should therefore consider:
• Source timestamps and cloud receipt timestamps
• Stable partitioning appropriate to the asset or workload
• Message IDs or sequence numbers for duplicate detection
• Idempotent consumers capable of handling at-least-once delivery
Network interruptions make this especially important. A gateway may buffer telemetry and send it when connectivity returns, which means Azure arrival time may be much later than the actual machine timestamp.
EVENT GRID SOLVES A DIFFERENT PROBLEM
Azure Event Grid is designed around publish-and-subscribe notifications. Instead of continuously reconstructing the state of a machine from thousands of readings, Event Grid tells interested systems that something changed and gives subscribers an opportunity to react.
Examples in an IoT environment include device creation, device deletion, connection, disconnection, or carefully selected telemetry-derived conditions. A device-created event might trigger an asset onboarding process. A device-disconnected event might start a technical investigation. A detected engineering condition might trigger a maintenance workflow.
Typical subscribers include:
• Azure Functions
• Logic Apps
• Webhooks
• Security workflows
• Asset-management systems
• Operational applications
The fundamental architectural difference is simple: telemetry asks what an asset has been doing, while an event asks what changed and whether something should react.
EVENTS SHOULD START INVESTIGATIONS, NOT DEFINE REALITY
Event Grid should generally be treated as a notification mechanism, not as the final authoritative state of the physical world. Event delivery can be repeated, and events are not something a subscriber should blindly interpret as the final current state without checking.
A device-disconnected event should therefore usually trigger a state check. An Azure Function might verify the current device state, inspect recent telemetry, check relevant registry information, and then decide whether the investigation should remain open.
Handlers should be designed around:
• Idempotent processing
• Current-state verification
• Event identity and timestamps
• Clearly defined ownership of the resulting action
This matters because duplicate events should not create duplicate tickets, duplicate alerts, or conflicting operational records.
A DEVICE DISCONNECT DOES NOT MEAN THE MACHINE STOPPED
One of the most important distinctions in industrial IoT is the difference between cloud connectivity, data-collection health, and production state. If an industrial gateway disconnects from IoT Hub, the cloud has lost visibility into that gateway. That does not automatically mean that the physical machine stopped.
The PLC may continue controlling the equipment locally while the gateway temporarily loses connectivity. Production may continue normally, and the gateway may even buffer telemetry and upload it later. The opposite can also happen: a gateway may remain connected to Azure while its connection to the PLC or local equipment has failed.
A useful architecture therefore separates cloud connection state, gateway health, machine production state, and MES work-order state. Only by combining those sources can the business determine whether a technical connectivity issue actually affected production.
FILTERING IN IOT HUB VS FILTERING IN EVENT GRID
Both technologies support filtering, but the purpose is different. IoT Hub Message Routing filtering answers which operational messages should enter a specific data path. Event Grid subscription filtering answers which notifications a particular subscriber should receive.
For example, production telemetry may go to one Event Hubs stream while energy data goes to another destination. At the same time, a support workflow may subscribe only to disconnected events from a particular group of gateways.
Confusing those two types of filtering often results in an architecture where every subscriber creates its own interpretation of the telemetry stream. A cleaner design keeps broad telemetry classification governed centrally and keeps Event Grid subscriptions focused on clearly defined reactions.
THE MANUFACTURING CONTEXT LIVES OUTSIDE THE MESSAGE BUS
Neither IoT Hub nor Event Grid knows what a machine means to the business. A device ID may identify a gateway, but it does not inherently know which production line the gateway belongs to, which machine it observes, which work order is running, whether a stop is planned, or whether production is already at risk.
That context usually comes from multiple systems. MES provides execution context such as work orders, operations, recipes, production state, and shift information. ERP provides demand, commitments, inventory, and planning information. An asset model connects devices, gateways, PLCs, machines, cells, lines, and sites.
In more advanced architectures, a digital twin or knowledge graph can help maintain those relationships. The result is a more reliable operational picture in which telemetry explains what the equipment reported, MES explains what production was doing, ERP explains why it matters, and the asset model explains how the technical components relate to the physical plant.
Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.
🚀 Want to be part of m365.fm?
Then stop just listening… and start showing up.
👉 Connect with me on LinkedIn and let’s make something happen:
- 🎙️ Be a podcast guest and share your story
- 🎧 Host your own episode (yes, seriously)
- 💡 Pitch topics the community actually wants to hear
- 🌍 Build your personal brand in the Microsoft 365 space
This isn’t just a podcast — it’s a platform for people who take action.
🔥 Most people wait. The best ones don’t.
👉 Connect with me on LinkedIn and send me a message:
"I want in"
Let’s build something awesome 👊
Frequently Asked Questions
What is the difference between IoT Hub Message Routing and Event Grid in Azure?
IoT Hub Message Routing is built for handling continuous, high-volume device telemetry that requires ordering, retention, and replay, whereas Event Grid is designed for publish-and-subscribe notifications that signal a state change requiring a reaction.
Why is message ordering important for industrial IoT telemetry?
Industrial telemetry relies on chronological sequence to accurately reconstruct machine states, calculate overall equipment effectiveness (OEE), and perform condition monitoring. Out-of-order messages can lead downstream systems to misinterpret normal shutdowns as equipment faults.
Does a device disconnect event mean that a manufacturing machine has stopped?
No, a gateway losing cloud connectivity only means visibility has changed, not that the physical machine stopped. The PLC may continue local operations while telemetry is buffered and uploaded later.
How should filtering be handled between IoT Hub and Event Grid?
IoT Hub Message Routing filtering should govern broad telemetry classification and data paths, while Event Grid subscription filtering should determine which notifications specific subscribers receive for operational reactions.
00:00:00,000 --> 00:00:02,440
A machine sends a temperature value every few seconds.
2
00:00:02,440 --> 00:00:03,480
It sends a cycle count.
3
00:00:03,480 --> 00:00:06,360
It sends vibration data, current draw, and a run state.
4
00:00:06,360 --> 00:00:08,280
Then at some point, it's connection drops.
5
00:00:08,280 --> 00:00:10,800
People often call all of that events.
6
00:00:10,800 --> 00:00:14,120
Azure uses the word too, so the confusion is understandable.
7
00:00:14,120 --> 00:00:16,800
But a stream of measurements and a notice that something changed
8
00:00:16,800 --> 00:00:19,360
create two very different architecture jobs.
9
00:00:19,360 --> 00:00:21,680
One path needs to preserve the production record.
10
00:00:21,680 --> 00:00:24,160
The other needs to tell the right system, or person,
11
00:00:24,160 --> 00:00:25,560
that it may need to react.
12
00:00:25,560 --> 00:00:27,040
If you treat both paths the same,
13
00:00:27,040 --> 00:00:29,160
you either build a very expensive alarm system
14
00:00:29,160 --> 00:00:31,960
for routine sensor data, or you bury urgent signals
15
00:00:31,960 --> 00:00:34,000
in a data pipeline in nobody watches.
16
00:00:34,000 --> 00:00:36,720
So let's follow one machine through a normal production shift,
17
00:00:36,720 --> 00:00:39,160
because that makes the difference much easier to hear.
18
00:00:39,160 --> 00:00:42,200
A press line, a sensor stream, and one missing signal.
19
00:00:42,200 --> 00:00:45,800
Picture a press line producing formed metal parts
20
00:00:45,800 --> 00:00:47,280
during a live work order.
21
00:00:47,280 --> 00:00:49,560
The line has a PLC, a few sensors,
22
00:00:49,560 --> 00:00:52,680
and an industrial gateway that connects the OT network to Azure.
23
00:00:52,680 --> 00:00:54,080
None of that is unusual.
24
00:00:54,080 --> 00:00:56,040
Most plants already have some version of this,
25
00:00:56,040 --> 00:00:59,000
even if the gateway talks to a mix of old and new equipment.
26
00:00:59,000 --> 00:01:01,280
Every few seconds, the gateway sends a device
27
00:01:01,280 --> 00:01:04,120
to cloud message into Azure IoT Hub.
28
00:01:04,120 --> 00:01:06,760
The message may include press temperature, vibration level,
29
00:01:06,760 --> 00:01:09,600
cycle count, energy draw, and whether the machine reports
30
00:01:09,600 --> 00:01:12,280
a running, idle, faulted, or stopped state.
31
00:01:12,280 --> 00:01:14,920
It may also carry a device identifier and a timestamp
32
00:01:14,920 --> 00:01:17,560
from the gateway that flow creates a record over time.
33
00:01:17,560 --> 00:01:20,040
The MES, the manufacturing execution system,
34
00:01:20,040 --> 00:01:21,840
tracks a different part of the story.
35
00:01:21,840 --> 00:01:23,760
It knows which work order is active,
36
00:01:23,760 --> 00:01:26,640
which operation should run, which recipe applies,
37
00:01:26,640 --> 00:01:28,920
perhaps which operators signed into the cell,
38
00:01:28,920 --> 00:01:31,440
and whether the order has started or completed.
39
00:01:31,440 --> 00:01:33,960
The machine can tell you that a cycle happened.
40
00:01:33,960 --> 00:01:36,120
The MES can tell you which production operation
41
00:01:36,120 --> 00:01:38,120
that cycle belongs to, both matter,
42
00:01:38,120 --> 00:01:39,640
but they don't come from the same place.
43
00:01:39,640 --> 00:01:41,920
Now imagine the press starts its shift normally.
44
00:01:41,920 --> 00:01:43,920
Temperature rises as the equipment warms up.
45
00:01:43,920 --> 00:01:46,320
Cycle counts increase, vibration stays inside
46
00:01:46,320 --> 00:01:47,760
its normal operating band.
47
00:01:47,760 --> 00:01:50,720
A production engineer may want that data later to study where.
48
00:01:50,720 --> 00:01:53,760
A maintenance engineer may want it to watch for a gradual change.
49
00:01:53,760 --> 00:01:56,120
A data team may use it to build a reliable history
50
00:01:56,120 --> 00:01:57,960
for analysis in Microsoft Fabric.
51
00:01:57,960 --> 00:02:00,360
Those consumers don't need a new workflow for every reading.
52
00:02:00,360 --> 00:02:03,280
They need the facts to keep arriving in a usable sequence
53
00:02:03,280 --> 00:02:06,120
so they can reconstruct what happened across the shift.
54
00:02:06,120 --> 00:02:07,640
A planner has another concern.
55
00:02:07,640 --> 00:02:09,320
The planner doesn't need a notification
56
00:02:09,320 --> 00:02:11,040
every time the press completes a cycle
57
00:02:11,040 --> 00:02:12,720
because that would turn production planning
58
00:02:12,720 --> 00:02:14,400
into a fairly noisy hobby.
59
00:02:14,400 --> 00:02:15,960
The planner needs trustworthy evidence
60
00:02:15,960 --> 00:02:18,000
when production has actually slowed, stopped,
61
00:02:18,000 --> 00:02:19,320
or moved off plan.
62
00:02:19,320 --> 00:02:21,200
That evidence may start with machine data,
63
00:02:21,200 --> 00:02:24,520
but it needs MES context before anyone changes a schedule.
64
00:02:24,520 --> 00:02:26,200
Then the gateway loses its connection.
65
00:02:26,200 --> 00:02:27,560
That is a different kind of message.
66
00:02:27,560 --> 00:02:29,640
The loss of connection may need a prompt response
67
00:02:29,640 --> 00:02:31,680
from an IT or OT support team.
68
00:02:31,680 --> 00:02:33,560
A workflow might open an investigation.
69
00:02:33,560 --> 00:02:35,240
It might notify the line support group.
70
00:02:35,240 --> 00:02:37,200
It might check whether the device has reconnected
71
00:02:37,200 --> 00:02:40,120
whether other devices on the same network segment also
72
00:02:40,120 --> 00:02:42,920
dropped or whether the gateway itself reports a fault.
73
00:02:42,920 --> 00:02:45,560
But notice what the connection loss does not prove.
74
00:02:45,560 --> 00:02:47,360
It doesn't prove that the press stopped.
75
00:02:47,360 --> 00:02:50,120
The press may continue running under local PLC control.
76
00:02:50,120 --> 00:02:51,920
It doesn't prove the work order is late.
77
00:02:51,920 --> 00:02:53,680
It doesn't prove a safety issue.
78
00:02:53,680 --> 00:02:56,240
And it doesn't tell a planner which alternative machine
79
00:02:56,240 --> 00:02:57,400
can take the work.
80
00:02:57,400 --> 00:03:00,560
A disconnected gateway tells you that cloud visibility has changed.
81
00:03:00,560 --> 00:03:02,440
That's enough to trigger an investigation.
82
00:03:02,440 --> 00:03:04,560
It isn't enough to write the production story.
83
00:03:04,560 --> 00:03:07,440
This is where teams often create trouble without meaning to.
84
00:03:07,440 --> 00:03:10,040
They see the word event, then send every telemetry
85
00:03:10,040 --> 00:03:11,920
reading through an event triggered workflow,
86
00:03:11,920 --> 00:03:13,960
a sensor value arrives, a function runs,
87
00:03:13,960 --> 00:03:16,000
another value arrives, another function runs.
88
00:03:16,000 --> 00:03:18,320
Soon, ordinary operating data creates a flood
89
00:03:18,320 --> 00:03:21,160
of small reactions, each one carrying less context
90
00:03:21,160 --> 00:03:22,880
than the person who needs to act.
91
00:03:22,880 --> 00:03:25,320
The opposite mistake can happen to a connection change
92
00:03:25,320 --> 00:03:28,280
enters a large telemetry pipeline, lands in storage,
93
00:03:28,280 --> 00:03:30,960
waits for a consumer, and eventually appears in a report.
94
00:03:30,960 --> 00:03:32,760
Technically, the signal arrived.
95
00:03:32,760 --> 00:03:34,560
Operationally, the team who needed to know
96
00:03:34,560 --> 00:03:36,160
may hear about it far too late.
97
00:03:36,160 --> 00:03:38,440
Think about the two messages leaving that same press line.
98
00:03:38,440 --> 00:03:41,560
One message says, "The press temperature is 82 degrees
99
00:03:41,560 --> 00:03:43,600
and the cycle count is still increasing."
100
00:03:43,600 --> 00:03:45,880
It belongs in a continuous operational record.
101
00:03:45,880 --> 00:03:48,360
The other says, "This device connection changed state."
102
00:03:48,360 --> 00:03:50,120
It belongs in a response path,
103
00:03:50,120 --> 00:03:52,800
where subscribed systems can decide whether they need to act.
104
00:03:52,800 --> 00:03:54,120
They may leave the same gateway,
105
00:03:54,120 --> 00:03:56,240
they may pass through the same IoT hub,
106
00:03:56,240 --> 00:03:58,240
but downstream they have different jobs.
107
00:03:58,240 --> 00:04:00,200
Telemetry means a time-ordered record.
108
00:04:00,200 --> 00:04:04,000
Telemetry is a record of repeated measurements over time.
109
00:04:04,000 --> 00:04:06,160
One reading has limited meaning on its own,
110
00:04:06,160 --> 00:04:08,240
but the sequence tells you how the machine behaved
111
00:04:08,240 --> 00:04:09,520
during a real production run.
112
00:04:09,520 --> 00:04:10,720
Take the press line again.
113
00:04:10,720 --> 00:04:13,120
A temperature reading of 82 degrees might be normal.
114
00:04:13,120 --> 00:04:14,280
It might also be a problem.
115
00:04:14,280 --> 00:04:15,960
You can't tell from that one number,
116
00:04:15,960 --> 00:04:17,680
unless you know what came before it,
117
00:04:17,680 --> 00:04:20,400
how fast it rose, what the press was doing at the time,
118
00:04:20,400 --> 00:04:22,760
and where the similar runs followed the same pattern.
119
00:04:22,760 --> 00:04:24,040
The order creates the story.
120
00:04:24,040 --> 00:04:25,520
A machine may start from cold,
121
00:04:25,520 --> 00:04:26,640
then enter a warmer period
122
00:04:26,640 --> 00:04:28,520
where temperature and current draw climb.
123
00:04:28,520 --> 00:04:30,680
After that, it reaches a stable run state.
124
00:04:30,680 --> 00:04:33,120
Much later, vibration might begin to creep upward
125
00:04:33,120 --> 00:04:35,040
while cycle time stretches slightly.
126
00:04:35,040 --> 00:04:37,920
The line then stops, perhaps for a planned material change,
127
00:04:37,920 --> 00:04:39,360
perhaps because a fault occurred,
128
00:04:39,360 --> 00:04:41,200
perhaps because the operator paused it.
129
00:04:41,200 --> 00:04:43,440
Those facts only form a usable production record
130
00:04:43,440 --> 00:04:45,280
when you keep their sequence intact.
131
00:04:45,280 --> 00:04:47,560
If a vibration reading from after the stop appears
132
00:04:47,560 --> 00:04:49,440
before the last running state message,
133
00:04:49,440 --> 00:04:51,680
a downstream system can tell the wrong story.
134
00:04:51,680 --> 00:04:54,120
It may classify a normal shutdown as a fault.
135
00:04:54,120 --> 00:04:56,920
It may calculate a downtime period that never existed.
136
00:04:56,920 --> 00:04:59,560
It may attach the wrong sensor pattern to the wrong batch.
137
00:04:59,560 --> 00:05:01,600
That kind of error rarely looks dramatic
138
00:05:01,600 --> 00:05:03,000
in a cloud architecture meeting.
139
00:05:03,000 --> 00:05:04,880
It looks like a small timestamp issue
140
00:05:04,880 --> 00:05:07,720
or a consumer problem someone plans to clean up later.
141
00:05:07,720 --> 00:05:08,840
On a factory floor though,
142
00:05:08,840 --> 00:05:11,360
those small errors can turn into hours of people arguing
143
00:05:11,360 --> 00:05:13,920
over which system tells the right version of the shift.
144
00:05:13,920 --> 00:05:16,560
And this is where OEE, overall equipment effectiveness,
145
00:05:16,560 --> 00:05:17,760
often gets misunderstood.
146
00:05:17,760 --> 00:05:19,600
OEE needs more than a machine counter.
147
00:05:19,600 --> 00:05:22,360
It needs a defensible timeline of planned time,
148
00:05:22,360 --> 00:05:25,440
runtime stops, speed loss, production count,
149
00:05:25,440 --> 00:05:26,640
and quality results.
150
00:05:26,640 --> 00:05:28,400
Some of that comes from machine telemetry.
151
00:05:28,400 --> 00:05:29,760
Some comes from the MS.
152
00:05:29,760 --> 00:05:31,960
But if the telemetry record arrives late,
153
00:05:31,960 --> 00:05:34,800
disappears during a network gap or lands out of sequence,
154
00:05:34,800 --> 00:05:37,360
your OEE calculation can become a polished answer
155
00:05:37,360 --> 00:05:38,520
to the wrong question.
156
00:05:38,520 --> 00:05:40,480
Condition monitoring works the same way.
157
00:05:40,480 --> 00:05:42,480
A maintenance engineer usually doesn't care only
158
00:05:42,480 --> 00:05:44,400
that vibration passed a number once.
159
00:05:44,400 --> 00:05:46,280
They care whether the value rose over days,
160
00:05:46,280 --> 00:05:48,240
whether it changes under a certain product load,
161
00:05:48,240 --> 00:05:50,920
whether the pattern repeats at a particular point in the cycle,
162
00:05:50,920 --> 00:05:54,040
and whether the reading returns to normal after maintenance.
163
00:05:54,040 --> 00:05:56,880
That analysis depends on a record you can inspect again.
164
00:05:56,880 --> 00:05:58,680
Think of telemetry as a production logbook
165
00:05:58,680 --> 00:06:00,120
that writes itself.
166
00:06:00,120 --> 00:06:03,120
It records what the equipment reported, when it reported it,
167
00:06:03,120 --> 00:06:05,000
and in what order the signals appeared.
168
00:06:05,000 --> 00:06:07,000
A logbook isn't useful because somebody
169
00:06:07,000 --> 00:06:09,040
reads every line the instant it arrives.
170
00:06:09,040 --> 00:06:11,560
It's useful because later, when a question comes up,
171
00:06:11,560 --> 00:06:13,720
you can trace the sequence and test an explanation
172
00:06:13,720 --> 00:06:14,680
against the record.
173
00:06:14,680 --> 00:06:17,360
That also explains why one consumer is rarely enough.
174
00:06:17,360 --> 00:06:20,080
Your operational team may need near real time processing
175
00:06:20,080 --> 00:06:21,200
for a current condition.
176
00:06:21,200 --> 00:06:23,160
Your data engineers may need the same telemetry
177
00:06:23,160 --> 00:06:24,560
for a curated history.
178
00:06:24,560 --> 00:06:27,040
A quality team may need to retrieve a slice of the record
179
00:06:27,040 --> 00:06:28,720
tied to a serial number or batch.
180
00:06:28,720 --> 00:06:30,440
And months later, an engineer may need
181
00:06:30,440 --> 00:06:33,600
to compare a failure pattern with a prior production run.
182
00:06:33,600 --> 00:06:35,000
Each consumer has a different job.
183
00:06:35,000 --> 00:06:37,880
They shouldn't need to compete for the same one-time notification,
184
00:06:37,880 --> 00:06:40,040
and they shouldn't depend on one workflow
185
00:06:40,040 --> 00:06:42,120
keeping perfect notes for everybody else.
186
00:06:42,120 --> 00:06:44,480
A proper telemetry path lets consumers read the stream
187
00:06:44,480 --> 00:06:45,640
for their own purpose.
188
00:06:45,640 --> 00:06:48,240
It retains raw facts long enough for traceability
189
00:06:48,240 --> 00:06:49,520
and later analysis.
190
00:06:49,520 --> 00:06:52,080
It also accepts an uncomfortable but normal manufacturing
191
00:06:52,080 --> 00:06:52,840
condition.
192
00:06:52,840 --> 00:06:55,720
Messages can arrive late and devices can lose contact.
193
00:06:55,720 --> 00:06:57,640
When that happens, you need to know the difference
194
00:06:57,640 --> 00:06:59,760
between a missing measurement and a machine state.
195
00:06:59,760 --> 00:07:02,800
A gap in telemetry might mean the gateway lost network access.
196
00:07:02,800 --> 00:07:04,240
It might mean the device rebooted,
197
00:07:04,240 --> 00:07:05,960
it might mean the sensor stopped sending.
198
00:07:05,960 --> 00:07:07,560
None of those explanations should quietly
199
00:07:07,560 --> 00:07:10,520
turn into the machine was down just because the data stopped.
200
00:07:10,520 --> 00:07:12,480
So the telemetry pipeline needs time, sequence,
201
00:07:12,480 --> 00:07:15,120
device identity, and enough source detail to interpret
202
00:07:15,120 --> 00:07:16,040
gaps, honestly.
203
00:07:16,040 --> 00:07:18,560
It needs consumers that can handle duplicates too,
204
00:07:18,560 --> 00:07:20,600
because reliable industrial messaging often
205
00:07:20,600 --> 00:07:22,720
means a message can appear more than once.
206
00:07:22,720 --> 00:07:25,320
Receiving the same fact twice is safer than silently losing it,
207
00:07:25,320 --> 00:07:27,520
but your processing logic still needs to recognize it.
208
00:07:27,520 --> 00:07:28,640
This is a stream problem.
209
00:07:28,640 --> 00:07:31,240
It needs replay, retention, and independent readers.
210
00:07:31,240 --> 00:07:33,240
A discrete event has a different job.
211
00:07:33,240 --> 00:07:36,440
Events mean something requires attention.
212
00:07:36,440 --> 00:07:38,080
An event tells a different kind of story.
213
00:07:38,080 --> 00:07:40,440
It says that a condition changed or that something
214
00:07:40,440 --> 00:07:43,200
happened and one or more systems may need to react.
215
00:07:43,200 --> 00:07:46,040
Think about a device being registered in IoT Hub.
216
00:07:46,040 --> 00:07:48,080
Somebody may need to assign it to a plant and owner
217
00:07:48,080 --> 00:07:49,200
and an asset record.
218
00:07:49,200 --> 00:07:51,120
When a device connects, a support process
219
00:07:51,120 --> 00:07:53,280
may record that the cloud link is active.
220
00:07:53,280 --> 00:07:55,240
When it disconnects, a support team may need
221
00:07:55,240 --> 00:07:57,520
to investigate whether the problem sits in the gateway,
222
00:07:57,520 --> 00:07:59,760
the network, the device, or the cloud path.
223
00:07:59,760 --> 00:08:02,000
Those are events because they change the state of the world
224
00:08:02,000 --> 00:08:03,720
in a way that can start work.
225
00:08:03,720 --> 00:08:05,800
The recipient doesn't usually need a long sequence
226
00:08:05,800 --> 00:08:07,480
of readings before it can respond.
227
00:08:07,480 --> 00:08:09,920
It needs a clear notice, an identity, a time,
228
00:08:09,920 --> 00:08:11,920
and enough detail to decide the next step.
229
00:08:11,920 --> 00:08:13,640
A workflow may create a support ticket,
230
00:08:13,640 --> 00:08:15,920
and as your function may check a device record,
231
00:08:15,920 --> 00:08:17,720
a logic app may send a message to the team
232
00:08:17,720 --> 00:08:18,920
responsible for the site.
233
00:08:18,920 --> 00:08:21,840
That response path needs to be fast and loosely connected.
234
00:08:21,840 --> 00:08:23,640
The system that handles device registration
235
00:08:23,640 --> 00:08:26,000
shouldn't need to know anything about vibration analysis.
236
00:08:26,000 --> 00:08:27,680
The team responsible for access control
237
00:08:27,680 --> 00:08:29,960
shouldn't need to consume the same high volume stream
238
00:08:29,960 --> 00:08:31,200
as the quality team.
239
00:08:31,200 --> 00:08:33,520
Each subscriber can receive the event it cares about,
240
00:08:33,520 --> 00:08:35,320
then do its own work without becoming part
241
00:08:35,320 --> 00:08:36,920
of the production data pipeline.
242
00:08:36,920 --> 00:08:39,200
Consider a new gateway arriving at a plant.
243
00:08:39,200 --> 00:08:40,960
The gateway gets registered in the cloud.
244
00:08:40,960 --> 00:08:42,640
That registration can trigger a process
245
00:08:42,640 --> 00:08:45,880
that creates an asset record, checks the device naming rule,
246
00:08:45,880 --> 00:08:48,920
records the site, and assigns an operational owner.
247
00:08:48,920 --> 00:08:50,400
None of those actions need the gateway
248
00:08:50,400 --> 00:08:52,320
to send thousands of measurements first.
249
00:08:52,320 --> 00:08:55,240
They need one reliable notice that the device exists,
250
00:08:55,240 --> 00:08:57,200
or take a disconnect event during a shift.
251
00:08:57,200 --> 00:08:59,160
A support workflow receives the notification
252
00:08:59,160 --> 00:09:01,680
and checks whether the gateway reconnects shortly after.
253
00:09:01,680 --> 00:09:04,360
It may compare the event with network monitoring data.
254
00:09:04,360 --> 00:09:06,440
It may raise the issue to the right team
255
00:09:06,440 --> 00:09:09,400
if other devices in the same cell disappear at the same time.
256
00:09:09,400 --> 00:09:11,560
That workflow responds to a changed condition.
257
00:09:11,560 --> 00:09:13,680
It doesn't need to reconstruct every press cycle
258
00:09:13,680 --> 00:09:14,960
since the start of the day.
259
00:09:14,960 --> 00:09:16,960
This is why I separate a response signal
260
00:09:16,960 --> 00:09:18,120
from an operational record.
261
00:09:18,120 --> 00:09:19,440
A response signal says,
262
00:09:19,440 --> 00:09:21,120
"Something deserves attention now."
263
00:09:21,120 --> 00:09:23,640
It can prompt a person, start an automated check,
264
00:09:23,640 --> 00:09:24,960
or create a work item.
265
00:09:24,960 --> 00:09:26,400
The signal needs clear ownership.
266
00:09:26,400 --> 00:09:28,040
If nobody owns the next action,
267
00:09:28,040 --> 00:09:30,680
the event simply becomes another message in another system.
268
00:09:30,680 --> 00:09:32,600
Factories already have enough of those.
269
00:09:32,600 --> 00:09:34,280
A useful event also has a boundary.
270
00:09:34,280 --> 00:09:37,120
Suppose the temperature in our press moves above a limit.
271
00:09:37,120 --> 00:09:39,120
The raw reading still belongs in the production record
272
00:09:39,120 --> 00:09:41,400
because an engineer may later need the full pattern,
273
00:09:41,400 --> 00:09:43,120
but a rule can interpret those readings
274
00:09:43,120 --> 00:09:44,720
and emit a separate event.
275
00:09:44,720 --> 00:09:47,080
Temperature condition requires review.
276
00:09:47,080 --> 00:09:50,120
That second message isn't pretending to replace the measurements.
277
00:09:50,120 --> 00:09:51,640
It tells a maintenance process
278
00:09:51,640 --> 00:09:54,000
that the measurements crossed a defined business
279
00:09:54,000 --> 00:09:55,320
or engineering threshold,
280
00:09:55,320 --> 00:09:57,280
that split keeps the architecture honest.
281
00:09:57,280 --> 00:10:00,080
The source data answers where what did the machine report?
282
00:10:00,080 --> 00:10:01,200
The event answers,
283
00:10:01,200 --> 00:10:04,400
what should a person or system consider doing because of it?
284
00:10:04,400 --> 00:10:06,680
You may send both at nearly the same time,
285
00:10:06,680 --> 00:10:08,080
but they carry different meaning
286
00:10:08,080 --> 00:10:09,920
and they should have different contracts.
287
00:10:09,920 --> 00:10:12,000
A fire alarm and a charter recorder both relate to safety,
288
00:10:12,000 --> 00:10:13,240
but they don't do the same job.
289
00:10:13,240 --> 00:10:15,160
The chart recorder tracks conditions through time,
290
00:10:15,160 --> 00:10:17,040
so an engineer can inspect the pattern.
291
00:10:17,040 --> 00:10:18,920
The fire alarm tells people to react.
292
00:10:18,920 --> 00:10:20,840
Nobody would replace the chart recorder
293
00:10:20,840 --> 00:10:22,120
with a room full of alarms
294
00:10:22,120 --> 00:10:23,680
and nobody would ask people to inspect
295
00:10:23,680 --> 00:10:26,280
a long trend chart before leaving a dangerous area.
296
00:10:26,280 --> 00:10:28,400
Industrial systems need the same discipline.
297
00:10:28,400 --> 00:10:31,160
Not every notification demands a human response either.
298
00:10:31,160 --> 00:10:32,800
Some events start technical checks.
299
00:10:32,800 --> 00:10:35,320
A device created event can trigger a registry update.
300
00:10:35,320 --> 00:10:37,880
A disconnect event can trigger a status check.
301
00:10:37,880 --> 00:10:39,720
A completed inspection can prompt a system
302
00:10:39,720 --> 00:10:41,600
to release the next workflow step.
303
00:10:41,600 --> 00:10:44,240
In each case, the event announces that a boundary was crossed,
304
00:10:44,240 --> 00:10:46,240
then a subscriber decides what to do next.
305
00:10:46,240 --> 00:10:48,440
That decision may still need more context.
306
00:10:48,440 --> 00:10:50,880
A device disconnect can start an investigation,
307
00:10:50,880 --> 00:10:53,320
but it can't by itself establish production loss.
308
00:10:53,320 --> 00:10:55,520
A device created event can start setup work,
309
00:10:55,520 --> 00:10:57,560
but it can't prove the device configuration
310
00:10:57,560 --> 00:10:59,200
is correct for a machine.
311
00:10:59,200 --> 00:11:01,000
An event opens the door to a response.
312
00:11:01,000 --> 00:11:02,760
It doesn't contain every fact needed
313
00:11:02,760 --> 00:11:04,320
for a safe operational decision.
314
00:11:04,320 --> 00:11:05,520
Keep that distinction in mind
315
00:11:05,520 --> 00:11:07,680
because telemetry can absolutely create an event,
316
00:11:07,680 --> 00:11:09,400
a stream rule, an engineering limit,
317
00:11:09,400 --> 00:11:12,720
or a model result may decide that the data now requires attention,
318
00:11:12,720 --> 00:11:15,560
but the event and the telemetry remain separate concerns,
319
00:11:15,560 --> 00:11:17,800
even when one leads directly to the other.
320
00:11:17,800 --> 00:11:18,640
Why?
321
00:11:18,640 --> 00:11:20,960
Everything is an event breaks down on the shop floor.
322
00:11:20,960 --> 00:11:25,560
Software teams often use the word event very broadly.
323
00:11:25,560 --> 00:11:28,240
A button click is an event, a record update is an event,
324
00:11:28,240 --> 00:11:29,760
a sensor reading is an event.
325
00:11:29,760 --> 00:11:31,680
At one level, that language is fine
326
00:11:31,680 --> 00:11:34,160
because all of those things happened at a point in time,
327
00:11:34,160 --> 00:11:36,480
but the factory doesn't care about language purity.
328
00:11:36,480 --> 00:11:38,440
It cares whether the information arrives
329
00:11:38,440 --> 00:11:40,920
in a form that supports the next operational job
330
00:11:40,920 --> 00:11:43,240
without adding noise, losing history,
331
00:11:43,240 --> 00:11:46,120
or waking up a workflow for no good reason.
332
00:11:46,120 --> 00:11:47,840
Picture a gateway is sending a temperature,
333
00:11:47,840 --> 00:11:49,880
pressure, vibration value, and power reading
334
00:11:49,880 --> 00:11:52,040
every few seconds from several machines.
335
00:11:52,040 --> 00:11:54,200
If every message enters a notification path,
336
00:11:54,200 --> 00:11:56,520
each reading can trigger a downstream handler.
337
00:11:56,520 --> 00:11:59,320
A function wakes up, a logic app evaluates a rule,
338
00:11:59,320 --> 00:12:02,680
a web hook receives a call, another system writes a status record,
339
00:12:02,680 --> 00:12:04,680
nothing necessarily fails at first.
340
00:12:04,680 --> 00:12:07,400
That's why this design can survive a proof of concept.
341
00:12:07,400 --> 00:12:10,440
Then the number of devices grows, the sampling rate changes,
342
00:12:10,440 --> 00:12:13,680
and more teams subscribe because the messages seem useful.
343
00:12:13,680 --> 00:12:16,480
The maintenance team wants one view, quality wants another,
344
00:12:16,480 --> 00:12:18,240
energy management wants a third.
345
00:12:18,240 --> 00:12:21,120
Somebody adds an email rule because a sensor value looks unusual
346
00:12:21,120 --> 00:12:23,120
and soon the same normal operating pattern
347
00:12:23,120 --> 00:12:25,120
starts several unrelated processes.
348
00:12:25,120 --> 00:12:26,880
A temperature reading isn't an incident,
349
00:12:26,880 --> 00:12:28,960
it may contribute to an incident later.
350
00:12:28,960 --> 00:12:31,000
It may be part of a pattern that points to where
351
00:12:31,000 --> 00:12:33,960
poor cooling and incorrect recipe or a sensor fault,
352
00:12:33,960 --> 00:12:36,520
but treating every measurement as an urgent business signal
353
00:12:36,520 --> 00:12:39,720
forces each downstream consumer to rediscover the same distinction.
354
00:12:39,720 --> 00:12:42,320
Is this just data or does somebody need to act?
355
00:12:42,320 --> 00:12:43,960
That is wasted architecture effort.
356
00:12:43,960 --> 00:12:46,120
You move filtering aggregation and state tracking
357
00:12:46,120 --> 00:12:49,400
into dozens of small handlers, where it becomes hard to test
358
00:12:49,400 --> 00:12:51,840
and even harder to trace when a production engineer asks
359
00:12:51,840 --> 00:12:54,720
a simple question about last Tuesday's shift.
360
00:12:54,720 --> 00:12:56,880
The delivery semantics create another problem.
361
00:12:56,880 --> 00:12:59,600
A stream consumer often needs to reason about sequence.
362
00:12:59,600 --> 00:13:03,400
It needs to know whether a run-state change came before a count increase,
363
00:13:03,400 --> 00:13:06,240
whether a threshold occurred during a production cycle,
364
00:13:06,240 --> 00:13:10,640
and whether a gap reflects a real stop or a loss of connectivity.
365
00:13:10,640 --> 00:13:14,440
A notification system doesn't exist to rebuild that kind of timeline.
366
00:13:14,440 --> 00:13:16,240
EventGrid doesn't guarantee delivery order.
367
00:13:16,240 --> 00:13:18,720
If you use it as the main home for high volume telemetry,
368
00:13:18,720 --> 00:13:20,760
a subscriber can receive messages in an order
369
00:13:20,760 --> 00:13:23,480
that doesn't match the order in which the machine produced them.
370
00:13:23,480 --> 00:13:25,840
You can try to repair that in every handler.
371
00:13:25,840 --> 00:13:28,280
Add timestamps, hold messages briefly,
372
00:13:28,280 --> 00:13:31,480
sort them, deduplicate them, track the last known state,
373
00:13:31,480 --> 00:13:34,360
then you discover that your simple event-driven design
374
00:13:34,360 --> 00:13:37,000
now contains several home-built stream processes
375
00:13:37,000 --> 00:13:39,680
each with slightly different logic and no shared record of truth.
376
00:13:39,680 --> 00:13:41,640
That's not an event-driven architecture.
377
00:13:41,640 --> 00:13:44,640
That's a distributed troubleshooting exercise with better branding.
378
00:13:44,640 --> 00:13:46,560
The reverse mistake causes trouble too.
379
00:13:46,560 --> 00:13:48,840
Teams sometimes take sparse lifecycle signals,
380
00:13:48,840 --> 00:13:51,320
such as device registration or a connection change,
381
00:13:51,320 --> 00:13:53,360
and send them only into a broad data path.
382
00:13:53,360 --> 00:13:56,880
The notification reaches storage and becomes available for analysis,
383
00:13:56,880 --> 00:13:59,200
but no subscriber receives a prompt trigger
384
00:13:59,200 --> 00:14:01,640
to check the device, update a registry,
385
00:14:01,640 --> 00:14:03,920
or start an owner assignment process.
386
00:14:03,920 --> 00:14:06,440
The data exists, yet the work doesn't start.
387
00:14:06,440 --> 00:14:08,520
Those two failures come from treating all messages
388
00:14:08,520 --> 00:14:11,000
as interchangeable because they share a transport name.
389
00:14:11,000 --> 00:14:12,400
They aren't interchangeable.
390
00:14:12,400 --> 00:14:13,760
Their volume may differ.
391
00:14:13,760 --> 00:14:15,160
Their need for order may differ.
392
00:14:15,160 --> 00:14:16,720
Their consumers may differ.
393
00:14:16,720 --> 00:14:19,640
Most of all, the question each message supports may differ.
394
00:14:19,640 --> 00:14:21,400
A telemetry message often asks,
395
00:14:21,400 --> 00:14:23,440
"What has this asset been doing?"
396
00:14:23,440 --> 00:14:25,760
A lifecycle event asks, "What changed?"
397
00:14:25,760 --> 00:14:27,080
and who should check it?
398
00:14:27,080 --> 00:14:29,320
Both questions can refer to the same device
399
00:14:29,320 --> 00:14:31,600
that doesn't turn them into the same data contract.
400
00:14:31,600 --> 00:14:34,320
There is also a people problem here, not just a cloud problem.
401
00:14:34,320 --> 00:14:36,400
When ordinary sensor traffic starts workflows,
402
00:14:36,400 --> 00:14:38,600
plan teams lose trust in alerts.
403
00:14:38,600 --> 00:14:40,920
They learn that a notification might mean anything
404
00:14:40,920 --> 00:14:43,400
from a routine reading to a real production issue,
405
00:14:43,400 --> 00:14:44,880
so they begin to ignore it.
406
00:14:44,880 --> 00:14:46,720
That response makes perfect sense.
407
00:14:46,720 --> 00:14:49,680
The system trained them to treat its messages as background noise.
408
00:14:49,680 --> 00:14:51,840
Good industrial architecture protects attention.
409
00:14:51,840 --> 00:14:54,720
It keeps the continuous record available for analysis,
410
00:14:54,720 --> 00:14:57,520
and it raises a response signal when a defined condition
411
00:14:57,520 --> 00:14:58,440
needs a response.
412
00:14:58,440 --> 00:15:01,320
Those are connected paths, but they should stay distinct.
413
00:15:01,320 --> 00:15:04,440
With that boundary clear, we can define the first Azure role
414
00:15:04,440 --> 00:15:06,320
without turning this into a product tour.
415
00:15:06,320 --> 00:15:09,680
What Azure IoT Hub actually handles?
416
00:15:09,680 --> 00:15:11,480
Azure IoT Hub sits at the point
417
00:15:11,480 --> 00:15:13,520
where a physical device meets the cloud.
418
00:15:13,520 --> 00:15:15,880
In a plant, that device may be an industrial gateway,
419
00:15:15,880 --> 00:15:18,360
a sensor package, a PLC-connected edge computer,
420
00:15:18,360 --> 00:15:20,120
or a purpose-built machine controller
421
00:15:20,120 --> 00:15:22,520
that can securely communicate beyond the OT network.
422
00:15:22,520 --> 00:15:23,720
Its first job is identity.
423
00:15:23,720 --> 00:15:26,280
Each device connects with its own identity and access
424
00:15:26,280 --> 00:15:28,000
rights, rather than every machine
425
00:15:28,000 --> 00:15:30,320
sharing one broad connection secret.
426
00:15:30,320 --> 00:15:32,320
That matters in manufacturing because a device
427
00:15:32,320 --> 00:15:34,160
isn't just another application client.
428
00:15:34,160 --> 00:15:37,920
It belongs to an asset, a line, a site, and an owner.
429
00:15:37,920 --> 00:15:40,200
And it may need to lose access without disrupting
430
00:15:40,200 --> 00:15:41,840
every other device in the plant.
431
00:15:41,840 --> 00:15:44,880
You can revoke or change a device's access independently.
432
00:15:44,880 --> 00:15:47,240
You can track which device communicates with the hub.
433
00:15:47,240 --> 00:15:48,720
And you can build a managed boundary
434
00:15:48,720 --> 00:15:50,960
between equipment networks and cloud services,
435
00:15:50,960 --> 00:15:54,040
rather than allowing every gateway to send data directly
436
00:15:54,040 --> 00:15:55,920
into every downstream system.
437
00:15:55,920 --> 00:15:58,520
That boundary doesn't remove the need for OT security.
438
00:15:58,520 --> 00:16:01,240
It doesn't replace network segmentation, firewall rules,
439
00:16:01,240 --> 00:16:04,000
industrial protocol controls, or a sensible edge design.
440
00:16:04,000 --> 00:16:06,360
But it gives the cloud side a controlled entry point
441
00:16:06,360 --> 00:16:07,720
for device communication.
442
00:16:07,720 --> 00:16:10,000
The next job is device-to-cloud messaging.
443
00:16:10,000 --> 00:16:12,040
A gateway can send its telemetry to IoT Hub
444
00:16:12,040 --> 00:16:13,720
where the platform accepts the message
445
00:16:13,720 --> 00:16:15,680
under the device identity that sent it.
446
00:16:15,680 --> 00:16:17,560
The message may carry sensor values,
447
00:16:17,560 --> 00:16:20,000
a machine state, an edge-calculated result,
448
00:16:20,000 --> 00:16:22,960
or a production signal prepared by software near the machine.
449
00:16:22,960 --> 00:16:24,680
That doesn't mean IoT Hub understands
450
00:16:24,680 --> 00:16:27,040
the press, the recipe, or the work order.
451
00:16:27,040 --> 00:16:29,360
It knows that an approved device sent a message.
452
00:16:29,360 --> 00:16:30,920
The meaning of that message still depends
453
00:16:30,920 --> 00:16:32,480
on the data contract you define
454
00:16:32,480 --> 00:16:34,840
and the manufacturing context you connect later.
455
00:16:34,840 --> 00:16:37,120
That distinction saves a lot of bad assumptions.
456
00:16:37,120 --> 00:16:39,160
If a gateway sends a field called state
457
00:16:39,160 --> 00:16:41,400
with the value running, IoT Hub can receive it safely
458
00:16:41,400 --> 00:16:43,160
and forward it based on your rules.
459
00:16:43,160 --> 00:16:45,920
It can't decide whether running means producing good parts,
460
00:16:45,920 --> 00:16:48,520
dry cycling, setup, rework, or a machine moving
461
00:16:48,520 --> 00:16:49,360
without material.
462
00:16:49,360 --> 00:16:52,480
Your MES and asset models still carry that part of the truth.
463
00:16:52,480 --> 00:16:55,360
IoT Hub also supports communication in the other direction
464
00:16:55,360 --> 00:16:57,040
from cloud to device.
465
00:16:57,040 --> 00:16:59,200
That can support commands, configuration changes,
466
00:16:59,200 --> 00:17:01,480
acknowledgments, or a request for a device
467
00:17:01,480 --> 00:17:02,880
to perform a defined action.
468
00:17:02,880 --> 00:17:04,920
In a factory, though, cloud to device messaging leads
469
00:17:04,920 --> 00:17:05,440
restrained.
470
00:17:05,440 --> 00:17:06,960
A cloud command isn't a substitute
471
00:17:06,960 --> 00:17:08,880
for a safety-rated control system.
472
00:17:08,880 --> 00:17:11,800
If a command could affect machine motion, product quality,
473
00:17:11,800 --> 00:17:14,480
or operator safety, you need clear control rules,
474
00:17:14,480 --> 00:17:16,560
local safeguards, and an OT design that
475
00:17:16,560 --> 00:17:18,840
can behave safely when the cloud disappears.
476
00:17:18,840 --> 00:17:21,160
The cloud can coordinate or request actions.
477
00:17:21,160 --> 00:17:22,960
The machine control layer remains responsible
478
00:17:22,960 --> 00:17:24,160
for safe physical behavior.
479
00:17:24,160 --> 00:17:26,680
There is another part of IoT Hub that often gets mixed up
480
00:17:26,680 --> 00:17:28,800
with telemetry, device twins.
481
00:17:28,800 --> 00:17:30,440
A device twin is a cloud-side document
482
00:17:30,440 --> 00:17:32,880
that represents state associated with a device.
483
00:17:32,880 --> 00:17:34,920
It separates desired properties, which
484
00:17:34,920 --> 00:17:36,960
the cloud wants the device to apply,
485
00:17:36,960 --> 00:17:39,680
from reported properties, which the device reports about its
486
00:17:39,680 --> 00:17:41,520
current condition or configuration.
487
00:17:41,520 --> 00:17:43,800
Imagine a gateway assigned to line two.
488
00:17:43,800 --> 00:17:45,800
The cloud may set a desired sampling interval
489
00:17:45,800 --> 00:17:47,240
or a configuration version.
490
00:17:47,240 --> 00:17:49,760
The gateway reports which version it actually applied
491
00:17:49,760 --> 00:17:51,840
when it last checked in, and perhaps which firmware
492
00:17:51,840 --> 00:17:52,880
build it runs.
493
00:17:52,880 --> 00:17:55,400
That creates a manageable conversation about configuration
494
00:17:55,400 --> 00:17:57,480
without treating every change as a command
495
00:17:57,480 --> 00:17:59,120
sent at the exact same moment.
496
00:17:59,120 --> 00:18:01,400
For industrial systems, twins work well for device
497
00:18:01,400 --> 00:18:02,240
and gateway state.
498
00:18:02,240 --> 00:18:04,200
They don't automatically become a complete digital
499
00:18:04,200 --> 00:18:05,480
twin of the factory.
500
00:18:05,480 --> 00:18:07,800
A device twin can describe a gateway's reported software
501
00:18:07,800 --> 00:18:08,440
version.
502
00:18:08,440 --> 00:18:11,080
It doesn't buy itself describe which machine the gateway
503
00:18:11,080 --> 00:18:13,840
observes, which cell contains that machine, which products
504
00:18:13,840 --> 00:18:17,200
it can produce, or how a failed sensor affects a customer
505
00:18:17,200 --> 00:18:17,800
order.
506
00:18:17,800 --> 00:18:20,520
You need an asset model, MS links, and sometimes a knowledge
507
00:18:20,520 --> 00:18:22,200
graph for that wider context.
508
00:18:22,200 --> 00:18:25,000
So think of IoT Hub as the device-facing boundary.
509
00:18:25,000 --> 00:18:27,040
It handles secure device identity,
510
00:18:27,040 --> 00:18:30,000
device-to-cloud messages, cloud-to-device communication,
511
00:18:30,000 --> 00:18:31,720
and device management state in a way that fits
512
00:18:31,720 --> 00:18:32,600
connected equipment.
513
00:18:32,600 --> 00:18:33,840
It doesn't replace the PLC.
514
00:18:33,840 --> 00:18:35,040
It doesn't replace the MES.
515
00:18:35,040 --> 00:18:37,760
It doesn't turn raw sensor data into production truth simply
516
00:18:37,760 --> 00:18:38,920
by accepting it.
517
00:18:38,920 --> 00:18:42,480
Once IoT Hub accepts a message, another question takes over.
518
00:18:42,480 --> 00:18:44,080
Where should that message travel?
519
00:18:44,080 --> 00:18:46,440
And which downstream system should receive it?
520
00:18:46,440 --> 00:18:50,720
IoT Hub message, rooting as the telemetry data plane.
521
00:18:50,720 --> 00:18:53,200
Once a device message enters IoT Hub, message rooting
522
00:18:53,200 --> 00:18:55,520
decides which approved downstream path receives it.
523
00:18:55,520 --> 00:18:58,080
Think of rooting as traffic rules at the edge of the cloud.
524
00:18:58,080 --> 00:19:00,960
The device sends one message, and IoT Hub evaluates rules
525
00:19:00,960 --> 00:19:03,400
you define before it sends that message onward.
526
00:19:03,400 --> 00:19:04,680
That matters because a plant rarely
527
00:19:04,680 --> 00:19:06,520
has one consumer for telemetry.
528
00:19:06,520 --> 00:19:08,520
Production data may need to feed a stream path
529
00:19:08,520 --> 00:19:10,240
for current operational analysis,
530
00:19:10,240 --> 00:19:13,600
while another copy needs to land in storage for traceability.
531
00:19:13,600 --> 00:19:16,400
A separate consumer may need a filtered set of messages
532
00:19:16,400 --> 00:19:19,160
for an application that tracks energy or machine condition.
533
00:19:19,160 --> 00:19:21,880
Rooting lets you send those messages to selected endpoints
534
00:19:21,880 --> 00:19:24,640
without teaching every gateway about every downstream system,
535
00:19:24,640 --> 00:19:26,200
the gateway sends to IoT Hub.
536
00:19:26,200 --> 00:19:28,040
The cloud architecture takes responsibility
537
00:19:28,040 --> 00:19:29,600
for distribution from there.
538
00:19:29,600 --> 00:19:32,200
The root can inspect information attached to the message.
539
00:19:32,200 --> 00:19:35,080
That includes application properties, system properties,
540
00:19:35,080 --> 00:19:38,400
and where your design needs it, parts of the message body.
541
00:19:38,400 --> 00:19:40,960
It can also use device twin tags and properties
542
00:19:40,960 --> 00:19:43,200
as part of a rooting query, say a gateway
543
00:19:43,200 --> 00:19:46,160
sends telemetry from several assets through one connection.
544
00:19:46,160 --> 00:19:49,040
The message can carry an asset identifier, a signal type,
545
00:19:49,040 --> 00:19:50,520
and a source timestamp.
546
00:19:50,520 --> 00:19:52,240
You might root production state messages
547
00:19:52,240 --> 00:19:54,360
to one path, energy readings to another,
548
00:19:54,360 --> 00:19:56,400
and diagnostic messages to a third.
549
00:19:56,400 --> 00:19:58,240
Or you may use a twin tag that identifies
550
00:19:58,240 --> 00:20:01,240
a device as belonging to a particular site or production area,
551
00:20:01,240 --> 00:20:04,520
then apply a root based on that managed device context.
552
00:20:04,520 --> 00:20:06,320
The important part is where that filtering happens.
553
00:20:06,320 --> 00:20:09,040
You don't want every downstream consumer receiving every message,
554
00:20:09,040 --> 00:20:10,600
then each team writing its own logic
555
00:20:10,600 --> 00:20:12,120
to discard most of the traffic.
556
00:20:12,120 --> 00:20:14,640
That creates duplicate work and inconsistent rules.
557
00:20:14,640 --> 00:20:16,920
A root gives you a controlled first decision
558
00:20:16,920 --> 00:20:18,600
about where a message belongs.
559
00:20:18,600 --> 00:20:21,080
It doesn't decide what the message means operationally.
560
00:20:21,080 --> 00:20:22,920
It decides where the message should go.
561
00:20:22,920 --> 00:20:25,040
IoT Hub can wrote device-to-cloud messages
562
00:20:25,040 --> 00:20:27,360
to Azure EventHubs, Azure Storage,
563
00:20:27,360 --> 00:20:29,120
Azure Service Bus, Qs, or topics,
564
00:20:29,120 --> 00:20:33,000
and on paid IoT Hub tiers that supported Azure Cosmos DB.
565
00:20:33,000 --> 00:20:35,480
Each destination fits a different downstream job,
566
00:20:35,480 --> 00:20:38,160
so naming the endpoint alone doesn't finish the design.
567
00:20:38,160 --> 00:20:40,080
EventHubs usually fits when you need a stream
568
00:20:40,080 --> 00:20:42,400
that several independent consumers can read.
569
00:20:42,400 --> 00:20:44,280
One consumer may process current data.
570
00:20:44,280 --> 00:20:47,240
Another may prepare data for later analysis, the stream.
571
00:20:47,240 --> 00:20:49,640
Let's those consumers work at their own pace.
572
00:20:49,640 --> 00:20:52,680
Subject to the retention and consumer design you choose.
573
00:20:52,680 --> 00:20:55,200
Azure Storage fits when you need an economical raw record
574
00:20:55,200 --> 00:20:57,000
that can support later inspection,
575
00:20:57,000 --> 00:20:58,560
audit work or batch processing.
576
00:20:58,560 --> 00:21:02,280
Service Bus can fit when telemetry produces a defined business message
577
00:21:02,280 --> 00:21:05,800
that an application must process through a Q or topic pattern.
578
00:21:05,800 --> 00:21:08,800
Cosmos DB can fit selected operational application patterns
579
00:21:08,800 --> 00:21:11,040
where low latency document access makes sense.
580
00:21:11,040 --> 00:21:14,040
None of those choices turns telemetry into a business decision.
581
00:21:14,040 --> 00:21:16,280
They give the telemetry a destination
582
00:21:16,280 --> 00:21:18,600
where a later process can do useful work with it.
583
00:21:18,600 --> 00:21:21,280
There is a practical advantage here that people sometimes miss.
584
00:21:21,280 --> 00:21:24,840
IoT Hub message routing doesn't add a separate routing charge
585
00:21:24,840 --> 00:21:26,840
beyond the telemetry ingress charge.
586
00:21:26,840 --> 00:21:28,880
If you wrote the same incoming device message
587
00:21:28,880 --> 00:21:30,520
to several configured endpoints,
588
00:21:30,520 --> 00:21:32,880
the routing itself doesn't turn that one incoming message
589
00:21:32,880 --> 00:21:34,920
into several IoT Hub ingress charges.
590
00:21:34,920 --> 00:21:37,560
That doesn't mean the rest of the architecture costs nothing.
591
00:21:37,560 --> 00:21:41,080
Storage, stream processing, downstream compute and data movement
592
00:21:41,080 --> 00:21:43,400
all need their own cost and capacity planning.
593
00:21:43,400 --> 00:21:45,640
But it does mean you shouldn't force unrelated consumers
594
00:21:45,640 --> 00:21:47,880
to share one fragile path just because you assume
595
00:21:47,880 --> 00:21:51,160
each additional route creates another IoT Hub ingestion bill.
596
00:21:51,160 --> 00:21:52,960
Ordering also shapes the decision.
597
00:21:52,960 --> 00:21:55,240
Microsoft documents that IoT Hub message routing
598
00:21:55,240 --> 00:21:57,560
maintains the order of routed messages.
599
00:21:57,560 --> 00:22:00,400
For production telemetry that gives you a stronger starting point
600
00:22:00,400 --> 00:22:02,400
than a notification path where consumers
601
00:22:02,400 --> 00:22:05,200
may receive messages in a different order from the source.
602
00:22:05,200 --> 00:22:07,600
Still, don't turn that statement into a guarantee
603
00:22:07,600 --> 00:22:09,600
your full business process can't support.
604
00:22:09,600 --> 00:22:12,640
Your partition strategy, consumer design, time stamp rules,
605
00:22:12,640 --> 00:22:15,400
retries and device behavior still matter.
606
00:22:15,400 --> 00:22:18,000
A gateway may buffer data during a network interruption.
607
00:22:18,000 --> 00:22:19,360
A device clock may drift.
608
00:22:19,360 --> 00:22:21,800
A consumer may process the same message more than once
609
00:22:21,800 --> 00:22:25,280
because the delivery model aims for at least one's delivery.
610
00:22:25,280 --> 00:22:27,040
Routing preserves an ordered path,
611
00:22:27,040 --> 00:22:29,520
but your application still needs to identify duplicates
612
00:22:29,520 --> 00:22:33,400
and reason carefully about source time versus cloud arrival time.
613
00:22:33,400 --> 00:22:35,840
That's normal engineering, especially in OT and IT
614
00:22:35,840 --> 00:22:36,840
convergence.
615
00:22:36,840 --> 00:22:40,160
Physical systems don't pause politely while networks reconnect.
616
00:22:40,160 --> 00:22:42,160
So the right way to view IoT Hub message routing
617
00:22:42,160 --> 00:22:43,560
is as the telemetry data plane.
618
00:22:43,560 --> 00:22:46,120
It takes device messages that need a controlled, filtered,
619
00:22:46,120 --> 00:22:48,520
ordered path and directs them to the services
620
00:22:48,520 --> 00:22:52,120
built to retain, process or distribute operational facts.
621
00:22:52,120 --> 00:22:54,000
A route can move the facts reliably.
622
00:22:54,000 --> 00:22:55,840
It doesn't decide which person needs an alert
623
00:22:55,840 --> 00:22:58,040
when those facts point to a problem.
624
00:22:58,040 --> 00:23:01,400
Route one, preserving the production story.
625
00:23:01,400 --> 00:23:03,440
Let's put a real route around the press line.
626
00:23:03,440 --> 00:23:05,640
The gateway sends a message each time
627
00:23:05,640 --> 00:23:08,360
it has a new set of readings worth sending upstream.
628
00:23:08,360 --> 00:23:09,920
That message needs more than a temperature
629
00:23:09,920 --> 00:23:12,240
and a vibration value if you want to use it later.
630
00:23:12,240 --> 00:23:14,960
It should carry a source time stamp, a sequence number
631
00:23:14,960 --> 00:23:17,480
where the gateway can provide one the asset identity
632
00:23:17,480 --> 00:23:19,040
and a clear machine state.
633
00:23:19,040 --> 00:23:21,280
Context travels with the message where it can.
634
00:23:21,280 --> 00:23:23,320
For this press, the gateway might also
635
00:23:23,320 --> 00:23:25,080
include the active work order reference
636
00:23:25,080 --> 00:23:27,280
that it received from the local production system
637
00:23:27,280 --> 00:23:30,240
or a correlation ID that lets another service connect
638
00:23:30,240 --> 00:23:32,280
the reading to the right operation.
639
00:23:32,280 --> 00:23:34,680
I'd still treat the MES as the authority for work order
640
00:23:34,680 --> 00:23:37,320
status, but adding a reference to the telemetry
641
00:23:37,320 --> 00:23:39,560
creates a usable link between a physical signal
642
00:23:39,560 --> 00:23:40,760
and the production record.
643
00:23:40,760 --> 00:23:44,160
A message could say, in plain terms, this came from press 12.
644
00:23:44,160 --> 00:23:46,360
It was the next message in this device sequence.
645
00:23:46,360 --> 00:23:48,160
The press reported that it was running
646
00:23:48,160 --> 00:23:50,400
and the gateway observed it at this time.
647
00:23:50,400 --> 00:23:52,600
It may also carry a cycle count, motor current,
648
00:23:52,600 --> 00:23:54,720
die temperature, vibration values,
649
00:23:54,720 --> 00:23:57,320
and a small set of quality-related process signals.
650
00:23:57,320 --> 00:23:59,680
That's enough to reconstruct part of the shift later.
651
00:23:59,680 --> 00:24:02,600
At IoT Hub, a route can select this production telemetry
652
00:24:02,600 --> 00:24:05,040
based on properties that the gateway sets consistently.
653
00:24:05,040 --> 00:24:07,680
You might mark the message type as production telemetry.
654
00:24:07,680 --> 00:24:09,920
You might separate it from gateway diagnostics,
655
00:24:09,920 --> 00:24:11,440
configuration acknowledgments,
656
00:24:11,440 --> 00:24:13,840
or infrequent maintenance status messages.
657
00:24:13,840 --> 00:24:16,000
The exact naming matters less than discipline.
658
00:24:16,000 --> 00:24:18,160
Everyone who produces and consumes the data
659
00:24:18,160 --> 00:24:19,920
needs to mean the same thing by it.
660
00:24:19,920 --> 00:24:22,960
The selected telemetry can then flow into Azure Event Hub.
661
00:24:22,960 --> 00:24:25,720
Event Hubs gives downstream consumers a stream they can read
662
00:24:25,720 --> 00:24:27,000
independently.
663
00:24:27,000 --> 00:24:30,400
One consumer might calculate current operating conditions.
664
00:24:30,400 --> 00:24:32,520
Another might prepare data for a data platform.
665
00:24:32,520 --> 00:24:34,680
A third might watch for data quality issues,
666
00:24:34,680 --> 00:24:37,160
such as a gateway that suddenly stop sending one signal
667
00:24:37,160 --> 00:24:39,200
while continuing to send everything else.
668
00:24:39,200 --> 00:24:41,080
Those consumers should not block each other.
669
00:24:41,080 --> 00:24:42,920
That matters during a busy shift.
670
00:24:42,920 --> 00:24:46,120
A data engineering process may fall behind for a short period.
671
00:24:46,120 --> 00:24:48,600
A near real time process may still need to continue.
672
00:24:48,600 --> 00:24:50,600
If both depend on one custom service,
673
00:24:50,600 --> 00:24:52,560
reading a message once and forwarding it perfectly
674
00:24:52,560 --> 00:24:55,600
to every other system, you've created a narrow point of failure
675
00:24:55,600 --> 00:24:57,920
right in the middle of your production record.
676
00:24:57,920 --> 00:24:59,720
The stream creates a better boundary.
677
00:24:59,720 --> 00:25:00,920
Alongside that streaming path,
678
00:25:00,920 --> 00:25:02,920
I'd usually keep a raw copy of the incoming messages
679
00:25:02,920 --> 00:25:05,720
in storage, not because raw data is automatically useful,
680
00:25:05,720 --> 00:25:07,400
but because you'll eventually need to answer
681
00:25:07,400 --> 00:25:09,720
a question that the curated model didn't anticipate.
682
00:25:09,720 --> 00:25:12,160
An engineer may ask why a condition model flagged
683
00:25:12,160 --> 00:25:14,160
a machine on a specific day.
684
00:25:14,160 --> 00:25:16,520
A quality team may need to inspect the sensor history
685
00:25:16,520 --> 00:25:18,960
around a batch issue, and OT engineer may need
686
00:25:18,960 --> 00:25:21,520
to compare the source payload with what a downstream process
687
00:25:21,520 --> 00:25:22,360
interpreted.
688
00:25:22,360 --> 00:25:24,080
Without a retained raw record,
689
00:25:24,080 --> 00:25:26,240
every one of those questions turns into a hunt
690
00:25:26,240 --> 00:25:28,960
through transform tables, partial logs, and somebody's memory
691
00:25:28,960 --> 00:25:30,120
that gets old quickly.
692
00:25:30,120 --> 00:25:32,280
The raw copy should preserve the original device facts
693
00:25:32,280 --> 00:25:34,000
and the technical metadata that explains
694
00:25:34,000 --> 00:25:35,200
how those facts arrived.
695
00:25:35,200 --> 00:25:37,520
Keep the device identity, keep source time,
696
00:25:37,520 --> 00:25:39,800
and cloud receipt time, where available.
697
00:25:39,800 --> 00:25:42,280
Keep the message properties that control routing.
698
00:25:42,280 --> 00:25:45,480
If the gateway assigns a sequence number, retain that too.
699
00:25:45,480 --> 00:25:48,640
Later processing can clean, normalize, and enrich the data.
700
00:25:48,640 --> 00:25:51,120
The raw record should stay closer to what the device actually
701
00:25:51,120 --> 00:25:51,720
sent.
702
00:25:51,720 --> 00:25:53,560
From there, a prepared operational data set
703
00:25:53,560 --> 00:25:56,240
can move into Microsoft fabric through a pipeline
704
00:25:56,240 --> 00:25:57,560
that fits your data design.
705
00:25:57,560 --> 00:25:59,200
I'm being deliberate with that wording,
706
00:25:59,200 --> 00:26:01,280
because fabric isn't a magic destination
707
00:26:01,280 --> 00:26:02,920
where raw machine messages somehow
708
00:26:02,920 --> 00:26:05,320
become trusted production facts.
709
00:26:05,320 --> 00:26:07,920
Before data reaches reporting or wider analysis,
710
00:26:07,920 --> 00:26:10,760
you need to decide how to handle device clock errors,
711
00:26:10,760 --> 00:26:15,160
later rivals, duplicate messages, unit conversions, missing
712
00:26:15,160 --> 00:26:18,320
values, and changes to the gateway payload.
713
00:26:18,320 --> 00:26:20,840
You also need to connect the signal to the asset model
714
00:26:20,840 --> 00:26:23,800
and where the use case needs it to mess production context.
715
00:26:23,800 --> 00:26:25,920
That prepared layer might turn a raw message
716
00:26:25,920 --> 00:26:28,280
into a statement such as press 12 ran
717
00:26:28,280 --> 00:26:30,840
during this operation window with these process readings
718
00:26:30,840 --> 00:26:32,760
against this work order reference.
719
00:26:32,760 --> 00:26:34,200
The transformation needs clear rules
720
00:26:34,200 --> 00:26:35,760
and those rules need an owner.
721
00:26:35,760 --> 00:26:37,120
Otherwise, a clean looking table can
722
00:26:37,120 --> 00:26:39,160
hide a fairly messy interpretation.
723
00:26:39,160 --> 00:26:41,400
Once fabric receives prepared operational data,
724
00:26:41,400 --> 00:26:43,360
power BI and other consumers can use a model
725
00:26:43,360 --> 00:26:44,800
built for their question.
726
00:26:44,800 --> 00:26:47,840
A production engineer may inspect patterns by asset and shift.
727
00:26:47,840 --> 00:26:50,640
An energy analyst may compare consumption across lines,
728
00:26:50,640 --> 00:26:53,720
a maintenance team may review trend history around a known fault.
729
00:26:53,720 --> 00:26:56,240
Each view starts from the same production story.
730
00:26:56,240 --> 00:26:58,120
Notice what the route itself has done.
731
00:26:58,120 --> 00:27:00,120
It has moved and preserved machine facts.
732
00:27:00,120 --> 00:27:02,960
It has given several consumers a controlled way to read them.
733
00:27:02,960 --> 00:27:06,040
It has supported both current processing and later investigation,
734
00:27:06,040 --> 00:27:09,080
but it hasn't decided whether anyone needs immediate attention
735
00:27:09,080 --> 00:27:10,920
that requires a separate response signal
736
00:27:10,920 --> 00:27:12,680
with a clear reason for action.
737
00:27:12,680 --> 00:27:15,600
Why ordering matters more than it sounds?
738
00:27:15,600 --> 00:27:18,120
A production record only works if you can place facts
739
00:27:18,120 --> 00:27:19,240
in the right sequence.
740
00:27:19,240 --> 00:27:21,240
That sounds obvious, but it's one of those details
741
00:27:21,240 --> 00:27:23,200
that stays invisible until someone tries
742
00:27:23,200 --> 00:27:25,640
to explain a stop, a quality issue, or a count
743
00:27:25,640 --> 00:27:27,080
that doesn't match the MES.
744
00:27:27,080 --> 00:27:28,360
Take a press cycle.
745
00:27:28,360 --> 00:27:31,120
The gateway reports that the press entered run state.
746
00:27:31,120 --> 00:27:33,200
Then it reports cycle counts increasing.
747
00:27:33,200 --> 00:27:35,560
A process value moves outside its normal band.
748
00:27:35,560 --> 00:27:37,080
Finally, the press reports are stopped.
749
00:27:37,080 --> 00:27:38,520
If you can preserve that sequence,
750
00:27:38,520 --> 00:27:41,080
a later consumer can ask sensible questions.
751
00:27:41,080 --> 00:27:43,120
Did the process condition appear before the stop?
752
00:27:43,120 --> 00:27:45,640
Did the stop happen during the active operation?
753
00:27:45,640 --> 00:27:48,560
And did production resume after an operator action?
754
00:27:48,560 --> 00:27:49,840
The order carries meaning.
755
00:27:49,840 --> 00:27:51,440
Now imagine the same messages arrive
756
00:27:51,440 --> 00:27:53,960
in a different order downstream, a stop arrives first,
757
00:27:53,960 --> 00:27:55,680
then a cycle count arrives that actually occurred
758
00:27:55,680 --> 00:27:58,520
before the stop, then the process value turns up.
759
00:27:58,520 --> 00:28:00,320
A simple state calculation may conclude
760
00:28:00,320 --> 00:28:02,280
that the press produced parts while stopped,
761
00:28:02,280 --> 00:28:04,520
or that the quality condition happened after production
762
00:28:04,520 --> 00:28:07,000
ended, neither conclusion needs to be true.
763
00:28:07,000 --> 00:28:08,440
The system only lost the sequence.
764
00:28:08,440 --> 00:28:10,360
Downtime windows are a good example.
765
00:28:10,360 --> 00:28:12,520
Many teams calculate downtime by taking the time
766
00:28:12,520 --> 00:28:14,760
between a machine entering a stop state
767
00:28:14,760 --> 00:28:16,480
and returning to a running state.
768
00:28:16,480 --> 00:28:18,760
That logic seems simple until you include
769
00:28:18,760 --> 00:28:20,880
late messages a gateway reconnect
770
00:28:20,880 --> 00:28:23,400
or repeated delivery of the same state change.
771
00:28:23,400 --> 00:28:26,280
If the consumer treats cloud arrival time as machine time,
772
00:28:26,280 --> 00:28:27,960
it can create a downtime window
773
00:28:27,960 --> 00:28:30,440
that belongs to the network rather than the machine.
774
00:28:30,440 --> 00:28:33,040
If it accepts every duplicate as a new state change,
775
00:28:33,040 --> 00:28:35,600
it can split one stop into several smaller stops.
776
00:28:35,600 --> 00:28:38,160
Then somebody asks why OEE dropped on a line
777
00:28:38,160 --> 00:28:40,320
that operators insist ran normally,
778
00:28:40,320 --> 00:28:42,800
and the answer sits buried in message handling logic.
779
00:28:42,800 --> 00:28:45,200
That's why I'd keep several time concepts separate.
780
00:28:45,200 --> 00:28:47,040
The source timestamp tells you when the device
781
00:28:47,040 --> 00:28:48,840
or gateway observe the condition.
782
00:28:48,840 --> 00:28:52,040
The IoT Hub receipt time tells you when Azure accepted the message.
783
00:28:52,040 --> 00:28:55,040
Processing time tells you when a downstream service handled it.
784
00:28:55,040 --> 00:28:57,680
Those times can differ, especially when an edge device buffers data
785
00:28:57,680 --> 00:28:59,960
during a network loss and sends it after reconnecting.
786
00:28:59,960 --> 00:29:02,160
None of that means source time always wins.
787
00:29:02,160 --> 00:29:03,440
Device clocks can drift.
788
00:29:03,440 --> 00:29:06,080
Gateways can stamp a message after collecting the data.
789
00:29:06,080 --> 00:29:08,720
Sometimes the machine controller has the most trustworthy clock
790
00:29:08,720 --> 00:29:10,160
and sometimes it doesn't.
791
00:29:10,160 --> 00:29:12,640
You need an explicit rule for which timestamp supports
792
00:29:12,640 --> 00:29:16,120
each operational question rather than letting whichever field appears
793
00:29:16,120 --> 00:29:19,080
first in a data set to decide the story.
794
00:29:19,080 --> 00:29:21,200
Partitioning deserves the same level of care.
795
00:29:21,200 --> 00:29:22,880
A routed path can maintain message order,
796
00:29:22,880 --> 00:29:25,560
but order has to relate to the unit of work that matters.
797
00:29:25,560 --> 00:29:28,320
If messages from one asset land across partitions
798
00:29:28,320 --> 00:29:30,200
without a consistent partition choice,
799
00:29:30,200 --> 00:29:33,520
a consumer may not see the full sequence as one ordered stream.
800
00:29:33,520 --> 00:29:35,320
For a machine-level state reconstruction,
801
00:29:35,320 --> 00:29:37,480
you usually want a stable way to keep that machines
802
00:29:37,480 --> 00:29:38,880
related messages together.
803
00:29:38,880 --> 00:29:41,640
That doesn't mean you put an entire plant through one partition.
804
00:29:41,640 --> 00:29:44,080
You still need throughput and parallel processing.
805
00:29:44,080 --> 00:29:47,160
It means you choose a partition approach that matches the question.
806
00:29:47,160 --> 00:29:50,600
A single asset, a gateway, or another defined source boundary
807
00:29:50,600 --> 00:29:52,080
may provide the right grouping,
808
00:29:52,080 --> 00:29:53,960
depending on how the data arrives.
809
00:29:53,960 --> 00:29:55,320
Then there's duplicate delivery.
810
00:29:55,320 --> 00:29:58,400
IoT hub message routing follows in at least one delivery model.
811
00:29:58,400 --> 00:30:00,160
In plain English, a downstream endpoint
812
00:30:00,160 --> 00:30:01,960
can receive the same message more than once.
813
00:30:01,960 --> 00:30:03,520
That protects against silent loss,
814
00:30:03,520 --> 00:30:05,360
which is the right bias for industrial data,
815
00:30:05,360 --> 00:30:08,200
but it means every consumer needs a duplicate strategy.
816
00:30:08,200 --> 00:30:09,600
A sequence number can help.
817
00:30:09,600 --> 00:30:11,720
A device-generated message ID can help.
818
00:30:11,720 --> 00:30:14,680
A composite identity built from device ID, source time,
819
00:30:14,680 --> 00:30:18,280
and sequence can help if the source can create it consistently.
820
00:30:18,280 --> 00:30:21,080
The exact method depends on the equipment and gateway software,
821
00:30:21,080 --> 00:30:22,800
but the principle stays the same.
822
00:30:22,800 --> 00:30:25,400
A consumer must recognize whether it has already processed
823
00:30:25,400 --> 00:30:26,520
this production fact.
824
00:30:26,520 --> 00:30:29,800
Don't confuse ordered delivery with exactly once business outcomes.
825
00:30:29,800 --> 00:30:31,320
Even when messages arrive in order,
826
00:30:31,320 --> 00:30:33,560
a consumer might crash after updating a table
827
00:30:33,560 --> 00:30:35,440
but before recording its checkpoint.
828
00:30:35,440 --> 00:30:38,400
When it restarts, it may process the message again.
829
00:30:38,400 --> 00:30:41,280
Or two downstream systems may interpret the same machine
830
00:30:41,280 --> 00:30:44,040
state differently because their business rules differ.
831
00:30:44,040 --> 00:30:46,440
Transport order can give you a stable foundation.
832
00:30:46,440 --> 00:30:49,040
It can't remove the need for competent processing,
833
00:30:49,040 --> 00:30:51,480
clear data ownership, and careful state logic.
834
00:30:51,480 --> 00:30:52,800
That's not a flaw in the platform.
835
00:30:52,800 --> 00:30:54,400
It's what reliable distributed systems
836
00:30:54,400 --> 00:30:57,760
look like once they meet physical equipment and imperfect networks.
837
00:30:57,760 --> 00:30:59,760
For telemetry, this work earns its place
838
00:30:59,760 --> 00:31:02,880
because the sequence itself supports analysis, traceability,
839
00:31:02,880 --> 00:31:04,080
and production facts.
840
00:31:04,080 --> 00:31:07,000
But some messages don't need a consumer to reconstruct history.
841
00:31:07,000 --> 00:31:09,240
They need a team or system to react to a change,
842
00:31:09,240 --> 00:31:12,560
and that takes us from data transport into response signals.
843
00:31:12,560 --> 00:31:15,720
What Azure Event Grid actually handles?
844
00:31:15,720 --> 00:31:18,000
Event Grid handles the moment when something changes
845
00:31:18,000 --> 00:31:19,760
and other systems need to know about it.
846
00:31:19,760 --> 00:31:21,760
It uses a published subscribe model,
847
00:31:21,760 --> 00:31:24,240
which sounds abstract until you put it in a plant,
848
00:31:24,240 --> 00:31:26,600
a device, an Azure service, or an application
849
00:31:26,600 --> 00:31:29,040
publishes a notice, systems that subscribe
850
00:31:29,040 --> 00:31:32,000
to that type of notice, receive it, and decide what to do.
851
00:31:32,000 --> 00:31:33,640
They don't need to sit there reading a long stream
852
00:31:33,640 --> 00:31:35,920
and waiting for their one relevant message.
853
00:31:35,920 --> 00:31:37,680
Event Grid pushes the event to them,
854
00:31:37,680 --> 00:31:39,800
that fits work which starts because of a change.
855
00:31:39,800 --> 00:31:41,720
A device gets registered, a device disconnects,
856
00:31:41,720 --> 00:31:42,920
a device comes back online,
857
00:31:42,920 --> 00:31:45,760
an application detects a condition that needs review.
858
00:31:45,760 --> 00:31:47,800
Those events can go to separate handlers,
859
00:31:47,800 --> 00:31:50,440
and each handler can stay focused on its own job
860
00:31:50,440 --> 00:31:54,280
instead of becoming another consumer of the full telemetry flow.
861
00:31:54,280 --> 00:31:56,040
Think of a maintenance support process.
862
00:31:56,040 --> 00:31:58,200
It may care when a gateway disconnects,
863
00:31:58,200 --> 00:32:01,040
because that means the plant has lost cloud visibility
864
00:32:01,040 --> 00:32:02,280
for a part of the line.
865
00:32:02,280 --> 00:32:04,320
An asset management process may care
866
00:32:04,320 --> 00:32:06,040
when a device gets created because it needs
867
00:32:06,040 --> 00:32:08,160
to check ownership and site assignment.
868
00:32:08,160 --> 00:32:10,880
A security process may care when a device gets deleted,
869
00:32:10,880 --> 00:32:12,800
because it may need to remove related access
870
00:32:12,800 --> 00:32:14,080
or inspect what changed.
871
00:32:14,080 --> 00:32:16,040
Those processes don't need the same destination.
872
00:32:16,040 --> 00:32:18,200
With Event Grid, one event source can publish a notice,
873
00:32:18,200 --> 00:32:20,160
and many subscribers can receive it.
874
00:32:20,160 --> 00:32:22,120
And Azure Function can run a small piece of code
875
00:32:22,120 --> 00:32:24,680
to validate a condition or update a record.
876
00:32:24,680 --> 00:32:27,000
A logic app can start a workflow that sends a message,
877
00:32:27,000 --> 00:32:29,880
creates a task, or connects to an approved business system.
878
00:32:29,880 --> 00:32:32,720
A web hook can notify a service outside Azure,
879
00:32:32,720 --> 00:32:34,520
where that makes sense for the architecture.
880
00:32:34,520 --> 00:32:37,520
Each subscriber gets a copy of the event intended for it.
881
00:32:37,520 --> 00:32:38,960
That fan-out model matters
882
00:32:38,960 --> 00:32:41,080
when different teams own different reactions.
883
00:32:41,080 --> 00:32:43,000
The team that manages device identity
884
00:32:43,000 --> 00:32:44,400
shouldn't need to wait for the team
885
00:32:44,400 --> 00:32:46,080
that owns maintenance workflows.
886
00:32:46,080 --> 00:32:47,920
The maintenance process shouldn't need access
887
00:32:47,920 --> 00:32:49,400
to the provisioning logic.
888
00:32:49,400 --> 00:32:52,480
Event Grid lets them react independently to the same change
889
00:32:52,480 --> 00:32:55,160
with each subscription defining what it wants to receive.
890
00:32:55,160 --> 00:32:58,640
But Independence also means the subscribers must act defensively.
891
00:32:58,640 --> 00:33:01,400
Event Grid delivery follows and at least once model,
892
00:33:01,400 --> 00:33:03,440
a handler can receive an event more than once,
893
00:33:03,440 --> 00:33:06,360
so it needs to process the event safely if a retry occurs.
894
00:33:06,360 --> 00:33:09,360
It also doesn't promise that events arrive in the order they happened.
895
00:33:09,360 --> 00:33:11,720
A handler can't blindly assume that a connection notice
896
00:33:11,720 --> 00:33:14,040
gives it the full and final current state of a device.
897
00:33:14,040 --> 00:33:16,560
For this kind of work, that is usually manageable.
898
00:33:16,560 --> 00:33:19,080
A disconnect notification should often trigger a check,
899
00:33:19,080 --> 00:33:20,520
not an irreversible conclusion.
900
00:33:20,520 --> 00:33:22,360
The handler can inspect current device state,
901
00:33:22,360 --> 00:33:25,720
query a registry, or test whether the device has already returned.
902
00:33:25,720 --> 00:33:28,080
It can record the event for audit purposes,
903
00:33:28,080 --> 00:33:30,680
then decide whether the condition still needs action.
904
00:33:30,680 --> 00:33:33,440
That approach fits an event notification model very well.
905
00:33:33,440 --> 00:33:35,240
The event itself needs enough information
906
00:33:35,240 --> 00:33:36,960
to point the handler in the right direction.
907
00:33:36,960 --> 00:33:39,760
It has a type that tells the subscriber what kind of change occurred.
908
00:33:39,760 --> 00:33:43,800
It includes a subject which commonly identifies the affected resource or device.
909
00:33:43,800 --> 00:33:48,200
It also carries event data that the handler can inspect before it starts its work.
910
00:33:48,200 --> 00:33:52,280
For an industrial solution, the device ID gives you a technical starting point.
911
00:33:52,280 --> 00:33:55,320
Your own asset model then connects that device to a gateway,
912
00:33:55,320 --> 00:33:58,280
machine, cell, line, site, and owner.
913
00:33:58,280 --> 00:33:59,640
Event Grid carries the notice.
914
00:33:59,640 --> 00:34:02,040
It doesn't hold the whole factory model in its head.
915
00:34:02,040 --> 00:34:04,760
That distinction protects you from a common design mistake.
916
00:34:04,760 --> 00:34:07,080
A subscriber receives a device disconnected event
917
00:34:07,080 --> 00:34:09,240
then assumes it knows that a press stopped producing.
918
00:34:09,240 --> 00:34:11,600
It doesn't. It knows the device connection changed.
919
00:34:11,600 --> 00:34:14,640
Any production conclusion must come from the right mix of machine state,
920
00:34:14,640 --> 00:34:16,760
MES facts, and asset context.
921
00:34:16,760 --> 00:34:20,240
Event Grid is a notification fabric, not an industrial historian.
922
00:34:20,240 --> 00:34:22,480
It is built to distribute notices in near real time,
923
00:34:22,480 --> 00:34:24,560
so reactive systems can start their own work.
924
00:34:24,560 --> 00:34:26,480
It doesn't replace a retained machine record
925
00:34:26,480 --> 00:34:30,320
and it isn't the place to reconstruct hours of sensor behavior after a quality problem.
926
00:34:30,320 --> 00:34:33,080
You can think of it as the part of the architecture that says,
927
00:34:33,080 --> 00:34:34,840
"A known change occurred."
928
00:34:34,840 --> 00:34:37,120
These systems ask to hear about it.
929
00:34:37,120 --> 00:34:39,680
The next action can be a technical check, a workflow,
930
00:34:39,680 --> 00:34:42,360
a record update, or a message to an accountable person.
931
00:34:42,360 --> 00:34:44,680
That makes it a strong fit for event-driven integration,
932
00:34:44,680 --> 00:34:47,760
particularly when many separate handlers need the same notice
933
00:34:47,760 --> 00:34:50,120
without tight coupling between their applications.
934
00:34:50,120 --> 00:34:53,560
IoT Hub can publish certain device and telemetry related events
935
00:34:53,560 --> 00:34:55,200
into that notification fabric.
936
00:34:55,200 --> 00:34:57,200
The details of which events it publishes
937
00:34:57,200 --> 00:34:59,280
and the limits that come with that integration
938
00:34:59,280 --> 00:35:02,480
shape whether event grid belongs in a specific manufacturing design.
939
00:35:02,480 --> 00:35:05,240
IoT Hub Event Grid integration.
940
00:35:05,240 --> 00:35:08,640
IoT Hub can publish selected device events into Event Grid,
941
00:35:08,640 --> 00:35:11,280
which gives you a direct bridge from connected equipment
942
00:35:11,280 --> 00:35:13,800
into a wider set of reactive Azure workflows.
943
00:35:13,800 --> 00:35:17,400
The event types matter because they tell you what IoT Hub is actually announcing.
944
00:35:17,400 --> 00:35:20,280
It can publish device created and device deleted events.
945
00:35:20,280 --> 00:35:23,240
It can publish device connected and device disconnected events.
946
00:35:23,240 --> 00:35:26,000
And it can publish device telemetry events
947
00:35:26,000 --> 00:35:28,120
when a device sends telemetry to the hub.
948
00:35:28,120 --> 00:35:30,760
Each event includes a subject that identifies the device.
949
00:35:30,760 --> 00:35:33,840
In practice, that subject follows the device identity path
950
00:35:33,840 --> 00:35:36,360
in the form of devices thus push device it.
951
00:35:36,360 --> 00:35:38,840
That might sound like a small implementation detail,
952
00:35:38,840 --> 00:35:42,600
but it gives subscription rules a clean way to focus on a group of devices
953
00:35:42,600 --> 00:35:45,640
without asking every handler to inspect every event.
954
00:35:45,640 --> 00:35:48,440
Say your plant name's gateways by site and production area.
955
00:35:48,440 --> 00:35:52,760
A subscription can focus on devices whose IDs begin with a site or line prefix.
956
00:35:52,760 --> 00:35:56,280
Or a handler can receive a device event, take the device ID
957
00:35:56,280 --> 00:35:59,360
and use the asset model to find the machine, cell, owner,
958
00:35:59,360 --> 00:36:01,760
and support group linked to that gateway.
959
00:36:01,760 --> 00:36:03,880
The event carries the technical identity.
960
00:36:03,880 --> 00:36:07,080
Your manufacturing model supplies the operational meaning.
961
00:36:07,080 --> 00:36:10,040
Telemetry events through Event Grid need a little more discipline
962
00:36:10,040 --> 00:36:11,440
than lifecycle events.
963
00:36:11,440 --> 00:36:14,520
IoT Hub requires the telemetry payload to contain valid JSON.
964
00:36:14,520 --> 00:36:17,840
The message content type must be set to application JSON.
965
00:36:17,840 --> 00:36:21,120
And the content encoding must be UTF8.
966
00:36:21,120 --> 00:36:22,960
If those fields aren't set correctly,
967
00:36:22,960 --> 00:36:26,960
IoT Hub can't present the telemetry payload to Event Grid in the expected structure.
968
00:36:26,960 --> 00:36:29,720
That can seem annoyingly strict when you first meet it,
969
00:36:29,720 --> 00:36:31,440
but it forces a useful question.
970
00:36:31,440 --> 00:36:34,400
Are your gateways publishing a defined message contract?
971
00:36:34,400 --> 00:36:38,600
Or are they just forwarding whatever shape of data happened to come from the equipment?
972
00:36:38,600 --> 00:36:41,200
For a manufacturing design, define that contract early.
973
00:36:41,200 --> 00:36:44,520
A telemetry message might include asset ID, source time sequence,
974
00:36:44,520 --> 00:36:46,040
signal type, and the actual readings.
975
00:36:46,040 --> 00:36:48,920
You may also use message properties for details such as site,
976
00:36:48,920 --> 00:36:50,840
line, or data classification.
977
00:36:50,840 --> 00:36:53,000
When you later create Event Grid subscriptions,
978
00:36:53,000 --> 00:36:57,600
those fields give you ways to direct only the relevant notices to the systems that asked for them.
979
00:36:57,600 --> 00:37:00,320
This can work well when the same event needs broad fan out.
980
00:37:00,320 --> 00:37:02,920
Imagine that a new industrial gateway is registered.
981
00:37:02,920 --> 00:37:06,520
The asset onboarding process needs to create or verify an equipment link.
982
00:37:06,520 --> 00:37:09,560
A security workflow needs to verify its owner and access status.
983
00:37:09,560 --> 00:37:11,920
The operation support team may need to confirm
984
00:37:11,920 --> 00:37:14,480
that the device belongs to the right plant and network zone.
985
00:37:14,480 --> 00:37:17,400
Those are separate reactions to one device created event.
986
00:37:17,400 --> 00:37:20,640
With Event Grid, each team can subscribe through its own handler.
987
00:37:20,640 --> 00:37:22,520
One pass may call in Azure function.
988
00:37:22,520 --> 00:37:23,800
Another may start a logic app.
989
00:37:23,800 --> 00:37:26,200
A third may send a web hook to a system outside Azure
990
00:37:26,200 --> 00:37:28,880
if that system belongs in the approved integration design.
991
00:37:28,880 --> 00:37:32,200
The same pattern applies to telemetry events, but you need to be selective.
992
00:37:32,200 --> 00:37:34,400
You might publish a filtered telemetry event
993
00:37:34,400 --> 00:37:37,280
when a gateway reports a defined machine condition.
994
00:37:37,280 --> 00:37:39,560
Several independent systems can then receive it.
995
00:37:39,560 --> 00:37:42,080
A maintenance application may record the condition.
996
00:37:42,080 --> 00:37:43,960
A workflow may open a review task.
997
00:37:43,960 --> 00:37:47,120
A data quality service may check whether related signals also arrived.
998
00:37:47,120 --> 00:37:49,560
That doesn't mean Event Grid should receive every sensor message
999
00:37:49,560 --> 00:37:51,920
simply because it can publish telemetry events.
1000
00:37:51,920 --> 00:37:55,160
A fan out path has a purpose when several systems need a prompt notice.
1001
00:37:55,160 --> 00:37:57,320
It becomes noise when routine measurements go everywhere
1002
00:37:57,320 --> 00:37:59,480
and every team has to decide which ones matter.
1003
00:37:59,480 --> 00:38:01,680
The scale difference gives Event Grid a clear place
1004
00:38:01,680 --> 00:38:03,680
in large integration designs.
1005
00:38:03,680 --> 00:38:07,880
IoT Hub message routing on paid tiers supports a limited set of custom endpoints
1006
00:38:07,880 --> 00:38:11,280
while Event Grid supports up to 500 endpoints per IoT Hub.
1007
00:38:11,280 --> 00:38:14,120
That wider subscriber model helps when separate applications,
1008
00:38:14,120 --> 00:38:17,320
teams and external services need their own event subscriptions.
1009
00:38:17,320 --> 00:38:19,600
Still, more endpoints don't create more meaning.
1010
00:38:19,600 --> 00:38:21,960
A larger fan out capability solves distribution.
1011
00:38:21,960 --> 00:38:24,840
It doesn't tell you whether a device disconnect affects production,
1012
00:38:24,840 --> 00:38:26,840
whether a temperature value needs action,
1013
00:38:26,840 --> 00:38:28,880
or whether a workflow should change a plan.
1014
00:38:28,880 --> 00:38:30,920
Those decisions still depend on the event contract,
1015
00:38:30,920 --> 00:38:33,400
the current state of the equipment and context from systems
1016
00:38:33,400 --> 00:38:35,280
such as the MES and asset model.
1017
00:38:35,280 --> 00:38:37,320
So use the integration for what it does well.
1018
00:38:37,320 --> 00:38:40,040
Publish clear device and telemetry related notices
1019
00:38:40,040 --> 00:38:42,520
to systems that need to react independently.
1020
00:38:42,520 --> 00:38:44,000
Keep the message shape deliberate.
1021
00:38:44,000 --> 00:38:45,800
Keep device identity consistent.
1022
00:38:45,800 --> 00:38:48,160
And let the subscriber receive a notice that starts work
1023
00:38:48,160 --> 00:38:51,000
rather than pretending it received a complete account of the factory.
1024
00:38:51,000 --> 00:38:54,160
More subscribers solve one problem, they don't solve every problem.
1025
00:38:54,160 --> 00:38:57,040
Event Grid does not promise order.
1026
00:38:57,040 --> 00:38:59,280
There's one constraint you need to design around
1027
00:38:59,280 --> 00:39:02,160
before you use Event Grid for an industrial workflow.
1028
00:39:02,160 --> 00:39:05,400
Event Grid doesn't guarantee that subscribers receive events
1029
00:39:05,400 --> 00:39:06,680
in the order they occurred.
1030
00:39:06,680 --> 00:39:07,640
That's not an edge case.
1031
00:39:07,640 --> 00:39:10,480
It changes what an event handler is allowed to conclude.
1032
00:39:10,480 --> 00:39:13,400
Picture a gateway that briefly loses its cloud connection,
1033
00:39:13,400 --> 00:39:16,600
reconnects, then loses it again while the plant network team
1034
00:39:16,600 --> 00:39:18,160
works through a switch issue.
1035
00:39:18,160 --> 00:39:20,520
IoT have may publish connection and disconnection events
1036
00:39:20,520 --> 00:39:21,920
for those state changes.
1037
00:39:21,920 --> 00:39:24,320
A subscriber could receive them in an awkward sequence.
1038
00:39:24,320 --> 00:39:27,280
It might see a connected event after a later disconnected event
1039
00:39:27,280 --> 00:39:28,960
or receive one notice after a delay
1040
00:39:28,960 --> 00:39:31,200
while newer notices have already arrived.
1041
00:39:31,200 --> 00:39:34,120
If the handler treats each incoming event as the final truth,
1042
00:39:34,120 --> 00:39:36,400
it can create the wrong operational state.
1043
00:39:36,400 --> 00:39:38,440
You can end up with a support record that says,
1044
00:39:38,440 --> 00:39:41,240
the gateway is online when it has already dropped again.
1045
00:39:41,240 --> 00:39:43,120
Or you can send an escalation to a line team
1046
00:39:43,120 --> 00:39:44,720
after the connection recovered.
1047
00:39:44,720 --> 00:39:46,600
Neither outcome means Event Grid failed.
1048
00:39:46,600 --> 00:39:49,600
The handler assumed more than the delivery model promised.
1049
00:39:49,600 --> 00:39:51,080
So the right pattern is simple.
1050
00:39:51,080 --> 00:39:53,000
Treat the event as a prompt to check.
1051
00:39:53,000 --> 00:39:55,960
A device disconnected event can start an Azure function
1052
00:39:55,960 --> 00:39:57,840
that function reads the current device state
1053
00:39:57,840 --> 00:40:00,240
from the device registry or checks the relevant device
1054
00:40:00,240 --> 00:40:01,960
twin then compares the current state
1055
00:40:01,960 --> 00:40:03,320
with the event it received.
1056
00:40:03,320 --> 00:40:04,920
If the device already reconnects,
1057
00:40:04,920 --> 00:40:06,880
the function may record the short disruption
1058
00:40:06,880 --> 00:40:08,440
and close the technical check.
1059
00:40:08,440 --> 00:40:10,320
If the device still appears unavailable,
1060
00:40:10,320 --> 00:40:12,440
it can raise the issue through the support path
1061
00:40:12,440 --> 00:40:14,040
your plant has defined.
1062
00:40:14,040 --> 00:40:15,400
The event starts the investigation.
1063
00:40:15,400 --> 00:40:16,560
It doesn't settle it.
1064
00:40:16,560 --> 00:40:18,680
This also protects you from duplicate delivery.
1065
00:40:18,680 --> 00:40:21,000
Event Grid uses at least one delivery,
1066
00:40:21,000 --> 00:40:23,960
which means a subscriber can receive the same event more than once.
1067
00:40:23,960 --> 00:40:26,280
Your function or workflow needs a way to recognize
1068
00:40:26,280 --> 00:40:28,200
that it already created the support case,
1069
00:40:28,200 --> 00:40:29,880
already updated the asset status
1070
00:40:29,880 --> 00:40:32,000
or already notified the responsible team.
1071
00:40:32,000 --> 00:40:34,680
I'd import and see is the technical word for this.
1072
00:40:34,680 --> 00:40:37,440
In plain terms, processing the same event twice
1073
00:40:37,440 --> 00:40:41,320
should not create two tickets, two emails and two conflicting status changes.
1074
00:40:41,320 --> 00:40:44,400
A good handler stores enough information to make that decision.
1075
00:40:44,400 --> 00:40:47,440
It may use the event ID as a processed record reference.
1076
00:40:47,440 --> 00:40:50,640
It may keep the latest known state and timestamp for the device.
1077
00:40:50,640 --> 00:40:53,680
For lifecycle changes such as device creation or deletion,
1078
00:40:53,680 --> 00:40:55,720
it can inspect the versioning information
1079
00:40:55,720 --> 00:40:59,000
and current registry state before it changes anything.
1080
00:40:59,000 --> 00:41:01,520
That might sound like extra work for a simple notification,
1081
00:41:01,520 --> 00:41:03,840
but it's the work that lets notifications stay simple
1082
00:41:03,840 --> 00:41:05,240
when conditions become messy,
1083
00:41:05,240 --> 00:41:08,640
which they always do around networks, gateways and real equipment.
1084
00:41:08,640 --> 00:41:10,440
There's another reason to check current state
1085
00:41:10,440 --> 00:41:12,560
rather than trust the arrival order.
1086
00:41:12,560 --> 00:41:15,880
A device event describes a cloud-facing device connection.
1087
00:41:15,880 --> 00:41:18,640
It does not automatically describe the health of the machine,
1088
00:41:18,640 --> 00:41:21,240
the control network or the production operation.
1089
00:41:21,240 --> 00:41:24,960
A gateway may reconnect while its PLC data collection has failed.
1090
00:41:24,960 --> 00:41:28,320
A gateway may disconnect while the press continues to run locally,
1091
00:41:28,320 --> 00:41:30,240
and a device created event may arrive
1092
00:41:30,240 --> 00:41:32,120
before your onboarding workflow finishes
1093
00:41:32,120 --> 00:41:35,520
assigning the correct plant, owner and asset relationship.
1094
00:41:35,520 --> 00:41:37,120
The event tells you where to begin.
1095
00:41:37,120 --> 00:41:38,720
Your systems of record tell you what to do.
1096
00:41:38,720 --> 00:41:42,120
For that reason, I design event grid subscribers around a current state check
1097
00:41:42,120 --> 00:41:43,520
and a clear action boundary.
1098
00:41:43,520 --> 00:41:45,320
A handler receives a notification.
1099
00:41:45,320 --> 00:41:47,720
It validates the device identity and event type.
1100
00:41:47,720 --> 00:41:50,320
It checks the latest state from the right source.
1101
00:41:50,320 --> 00:41:53,720
Then it either records, escalates or closes the condition.
1102
00:41:53,720 --> 00:41:56,520
That approach also keeps your event contracts honest.
1103
00:41:56,520 --> 00:41:58,120
The notification can say,
1104
00:41:58,120 --> 00:42:01,520
a connection state change was reported for this device at this time.
1105
00:42:01,520 --> 00:42:04,120
It should not claim this production line has stopped,
1106
00:42:04,120 --> 00:42:06,120
unless another component has checked the machine
1107
00:42:06,120 --> 00:42:08,720
and MES context required to make that statement.
1108
00:42:08,720 --> 00:42:11,720
Event grid works very well for this kind of reactive flow.
1109
00:42:11,720 --> 00:42:14,720
It can tell separate systems that something worth checking occurred
1110
00:42:14,720 --> 00:42:18,120
and those systems can act without consuming the entire machine stream.
1111
00:42:18,120 --> 00:42:21,920
Just don't ask it to rebuild operational history from arrival order.
1112
00:42:21,920 --> 00:42:23,320
Use it to start the next check.
1113
00:42:23,320 --> 00:42:26,120
Connection events bring one more manufacturing-specific limitation
1114
00:42:26,120 --> 00:42:29,720
that teams often miss until they try to use them as a source for uptime.
1115
00:42:29,720 --> 00:42:31,920
The 60-second connection state trap.
1116
00:42:31,920 --> 00:42:35,520
A connection event can help you spot a cloud communication issue.
1117
00:42:35,520 --> 00:42:38,320
It cannot give you a precise uptime record for a machine.
1118
00:42:38,320 --> 00:42:41,120
IoT Hub attempts to report device connection state changes
1119
00:42:41,120 --> 00:42:44,120
but it reports changes no more often than every 60 seconds.
1120
00:42:44,120 --> 00:42:47,920
That means a gateway can disconnect, reconnect and disconnect again
1121
00:42:47,920 --> 00:42:48,920
within that window,
1122
00:42:48,920 --> 00:42:51,720
while the event stream doesn't show every change in between.
1123
00:42:51,720 --> 00:42:53,520
You may see several connected events
1124
00:42:53,520 --> 00:42:55,920
without a matching, disconnected event.
1125
00:42:55,920 --> 00:42:58,920
That can look strange until you understand the reporting interval.
1126
00:42:58,920 --> 00:43:02,720
The cloud received state notifications at the times it could report them.
1127
00:43:02,720 --> 00:43:05,120
It didn't record every movement in the network path
1128
00:43:05,120 --> 00:43:07,520
for a factory team that distinction matters a lot.
1129
00:43:07,520 --> 00:43:09,720
Imagine an industrial gateway on a press line
1130
00:43:09,720 --> 00:43:13,320
loses its Wi-Fi or wired uplink for 20 seconds during a shift change.
1131
00:43:13,320 --> 00:43:14,720
The gateway reconnects quickly.
1132
00:43:14,720 --> 00:43:17,320
It may buffer its telemetry locally and send it later.
1133
00:43:17,320 --> 00:43:20,320
The press may never stop. Operators may never notice anything.
1134
00:43:20,320 --> 00:43:22,920
Yet a cloud support system that treats connection events
1135
00:43:22,920 --> 00:43:26,120
as machine uptime could record a production disruption that didn't happen.
1136
00:43:26,120 --> 00:43:27,320
Now reverse the scenario.
1137
00:43:27,320 --> 00:43:29,520
The gateway remains connected to IoT Hub
1138
00:43:29,520 --> 00:43:33,120
but the OPC-UA connection between the gateway and the PLC fails.
1139
00:43:33,120 --> 00:43:34,920
The cloud sees a connected device.
1140
00:43:34,920 --> 00:43:37,120
The machine may still run or it may stop.
1141
00:43:37,120 --> 00:43:40,120
But either way the gateway can't collect the signals you expected.
1142
00:43:40,120 --> 00:43:41,520
The cloud connection looks healthy.
1143
00:43:41,520 --> 00:43:42,720
The data collection isn't.
1144
00:43:42,720 --> 00:43:44,320
That is why device connection state,
1145
00:43:44,320 --> 00:43:48,120
data collection health and machine production state need separate definitions.
1146
00:43:48,120 --> 00:43:50,520
They're related but they answer different questions.
1147
00:43:50,520 --> 00:43:55,320
Device connection state asks whether the IoT client maintained its connection to IoT Hub.
1148
00:43:55,320 --> 00:43:59,120
Data collection health asks whether the gateway can still read the PLC,
1149
00:43:59,120 --> 00:44:01,920
sensors or edge applications that supply its messages.
1150
00:44:01,920 --> 00:44:04,920
Production state asks whether the asset is producing, stopped,
1151
00:44:04,920 --> 00:44:08,120
inset up, starved for material, blocked downstream,
1152
00:44:08,120 --> 00:44:10,120
or running a planned maintenance task.
1153
00:44:10,120 --> 00:44:12,520
A single connected flag cannot answer all three.
1154
00:44:12,520 --> 00:44:15,520
If it could, factories would be much easier places to model
1155
00:44:15,520 --> 00:44:18,320
and Excel would finally be out of a job that hasn't happened.
1156
00:44:18,320 --> 00:44:20,320
Protocol choice also affects these events.
1157
00:44:20,320 --> 00:44:22,920
IoT Hub connection and disconnection notifications apply
1158
00:44:22,920 --> 00:44:25,920
to devices connecting through MQTT or AMQP,
1159
00:44:25,920 --> 00:44:28,120
including those protocols over web sockets.
1160
00:44:28,120 --> 00:44:31,120
Devices that only communicate through HTTPS don't trigger
1161
00:44:31,120 --> 00:44:33,120
those connection state notifications.
1162
00:44:33,120 --> 00:44:34,720
That doesn't make HTTPS wrong.
1163
00:44:34,720 --> 00:44:37,120
It just means you shouldn't design a monitoring process
1164
00:44:37,120 --> 00:44:39,320
that expects connection events from a device
1165
00:44:39,320 --> 00:44:41,920
that communicates only through HTTPS.
1166
00:44:41,920 --> 00:44:44,720
The delivery behavior follows the protocol and connection model.
1167
00:44:44,720 --> 00:44:47,320
Before you build a workflow around disconnect events,
1168
00:44:47,320 --> 00:44:49,120
ask a few direct questions,
1169
00:44:49,120 --> 00:44:52,320
which devices use persistent MQTT or AMQP connections,
1170
00:44:52,320 --> 00:44:55,120
which devices batch messages through HTTPS?
1171
00:44:55,120 --> 00:44:57,520
Does the gateway buffer when the plant network fails?
1172
00:44:57,520 --> 00:44:59,120
Does it publish a health signal that confirms
1173
00:44:59,120 --> 00:45:00,720
it can still read the PLC?
1174
00:45:00,720 --> 00:45:03,720
And which system owns the machine state that production uses?
1175
00:45:03,720 --> 00:45:07,120
Those answers shape the design more than the name of the Azure service.
1176
00:45:07,120 --> 00:45:09,320
A better industrial pattern uses connection events
1177
00:45:09,320 --> 00:45:11,520
as one input into a technical health check.
1178
00:45:11,520 --> 00:45:13,920
A disconnect notice may tell the support process
1179
00:45:13,920 --> 00:45:16,520
to inspect the gateway, its last telemetry time,
1180
00:45:16,520 --> 00:45:18,120
and the status of nearby devices.
1181
00:45:18,120 --> 00:45:19,720
If a whole group disappears together,
1182
00:45:19,720 --> 00:45:22,520
the likely cause may sit in a network segment or switch.
1183
00:45:22,520 --> 00:45:24,920
If one gateway drops while others remain online,
1184
00:45:24,920 --> 00:45:26,920
the team can narrow the investigation,
1185
00:45:26,920 --> 00:45:29,120
but don't label the machine as down yet.
1186
00:45:29,120 --> 00:45:31,720
For production availability, use machine state
1187
00:45:31,720 --> 00:45:33,520
from the control layer or MES
1188
00:45:33,520 --> 00:45:36,920
with the reason codes and operational rules your plant already uses.
1189
00:45:36,920 --> 00:45:39,920
For OT network health, use the monitoring tools and signals
1190
00:45:39,920 --> 00:45:41,520
that can see the relevant network path.
1191
00:45:41,520 --> 00:45:45,520
For cloud connectivity, use IoT Hub events and message arrival patterns.
1192
00:45:45,520 --> 00:45:48,120
Each source provides a different part of the operational picture.
1193
00:45:48,120 --> 00:45:50,520
That separation also helps when an event arrives
1194
00:45:50,520 --> 00:45:52,520
after the device has already returned.
1195
00:45:52,520 --> 00:45:54,120
The workflow doesn't need to panic.
1196
00:45:54,120 --> 00:45:55,720
It can check current cloud state,
1197
00:45:55,720 --> 00:45:57,520
inspect recent telemetry,
1198
00:45:57,520 --> 00:46:00,120
and record whether the incident still needs a person.
1199
00:46:00,120 --> 00:46:03,120
A short cloud gap may matter for data completeness
1200
00:46:03,120 --> 00:46:05,120
even when it has no effect on output.
1201
00:46:05,120 --> 00:46:07,520
A connected gateway proves one narrow fact
1202
00:46:07,520 --> 00:46:09,320
that gateway can talk to IoT Hub.
1203
00:46:09,320 --> 00:46:11,120
It doesn't prove the machine is healthy,
1204
00:46:11,120 --> 00:46:13,920
the PLC link works, or the line produced good parts.
1205
00:46:13,920 --> 00:46:16,520
Device lifecycle events and provisioning workflows.
1206
00:46:17,720 --> 00:46:19,520
A device lifecycle event often starts
1207
00:46:19,520 --> 00:46:22,120
before the gateway sends a single useful machine signal.
1208
00:46:22,120 --> 00:46:25,120
Picture a new gateway being prepared for a packaging cell.
1209
00:46:25,120 --> 00:46:27,120
Someone registers it in IoT Hub,
1210
00:46:27,120 --> 00:46:29,520
perhaps as part of an automated provisioning process
1211
00:46:29,520 --> 00:46:31,720
or a controlled engineering handover.
1212
00:46:31,720 --> 00:46:35,720
That registration can publish a device-created event into event grid
1213
00:46:35,720 --> 00:46:37,920
and the event can start the administrative work
1214
00:46:37,920 --> 00:46:40,120
that turns a technical device identity
1215
00:46:40,120 --> 00:46:42,520
into something the plant can support.
1216
00:46:42,520 --> 00:46:44,520
That work needs ownership.
1217
00:46:44,520 --> 00:46:46,920
An Azure function might check whether the device ID
1218
00:46:46,920 --> 00:46:49,520
follows the naming rule for that site and line.
1219
00:46:49,520 --> 00:46:51,120
It can create a pending asset record,
1220
00:46:51,120 --> 00:46:53,120
attach the device to an approved support group
1221
00:46:53,120 --> 00:46:55,320
and flag the record if the expected machine link
1222
00:46:55,320 --> 00:46:56,520
doesn't exist yet.
1223
00:46:56,520 --> 00:46:58,320
A logic app might then create a task
1224
00:46:58,320 --> 00:47:00,720
for the local OTE engineer to confirm the installation
1225
00:47:00,720 --> 00:47:03,320
and for the security team to check the access process.
1226
00:47:03,320 --> 00:47:04,920
None of that needs a telemetry stream.
1227
00:47:04,920 --> 00:47:06,920
It needs a prompt, traceable workflow.
1228
00:47:06,920 --> 00:47:09,120
The same event can reach more than one subscriber,
1229
00:47:09,120 --> 00:47:11,320
but each subscriber should own a narrow step.
1230
00:47:11,320 --> 00:47:13,920
The asset process manages the equipment relationship.
1231
00:47:13,920 --> 00:47:16,120
The identity process manages access.
1232
00:47:16,120 --> 00:47:18,720
The service process confirms that someone owns the device
1233
00:47:18,720 --> 00:47:20,320
after it reaches the plant.
1234
00:47:20,320 --> 00:47:21,720
If one workflow fails,
1235
00:47:21,720 --> 00:47:23,720
it shouldn't leave the whole onboarding process
1236
00:47:23,720 --> 00:47:25,920
as a mystery spread across email threads.
1237
00:47:25,920 --> 00:47:27,520
That is where event grid fits well.
1238
00:47:27,520 --> 00:47:30,320
It distributes the notice that a device record changed.
1239
00:47:30,320 --> 00:47:32,720
Your downstream services then check the details they need
1240
00:47:32,720 --> 00:47:35,520
and update their own records through defined interfaces.
1241
00:47:35,520 --> 00:47:38,520
The event doesn't need to contain the whole asset structure.
1242
00:47:38,520 --> 00:47:40,720
It only needs enough information to identify the device
1243
00:47:40,720 --> 00:47:42,320
and start the right process.
1244
00:47:42,320 --> 00:47:44,720
A deletion event deserves the same care.
1245
00:47:44,720 --> 00:47:47,920
When somebody deletes a device from IoT Hub that may be planned,
1246
00:47:47,920 --> 00:47:49,920
a gateway may have reached end of life,
1247
00:47:49,920 --> 00:47:51,120
moved to another plant,
1248
00:47:51,120 --> 00:47:52,920
or been replaced during a line upgrade.
1249
00:47:52,920 --> 00:47:55,920
It may also be an unexpected change that needs review.
1250
00:47:55,920 --> 00:47:57,720
A device deleted event can start,
1251
00:47:57,720 --> 00:47:59,920
access cleanup, remove stale monitoring links,
1252
00:47:59,920 --> 00:48:01,920
and ask the asset owner to confirm
1253
00:48:01,920 --> 00:48:03,520
whether the physical equipment changed
1254
00:48:03,520 --> 00:48:05,120
or only its cloud registration.
1255
00:48:05,120 --> 00:48:08,120
Don't let deletion quietly become housekeeping.
1256
00:48:08,120 --> 00:48:09,520
If the cloud identity disappears
1257
00:48:09,520 --> 00:48:11,920
while the gateway still sits beside a running machine,
1258
00:48:11,920 --> 00:48:13,920
the plant may lose data without anyone noticing
1259
00:48:13,920 --> 00:48:15,520
until a report turns blank.
1260
00:48:15,520 --> 00:48:17,720
A deletion workflow can check whether the asset still
1261
00:48:17,720 --> 00:48:19,720
appears active in the maintenance system,
1262
00:48:19,720 --> 00:48:21,720
whether its last message arrived recently,
1263
00:48:21,720 --> 00:48:23,720
and whether a replacement device registration
1264
00:48:23,720 --> 00:48:26,320
follows within the expected change process.
1265
00:48:26,320 --> 00:48:27,720
The action should match the risk.
1266
00:48:27,720 --> 00:48:30,520
That doesn't mean every life cycle event needs a human approval.
1267
00:48:30,520 --> 00:48:32,120
Many checks can run automatically.
1268
00:48:32,120 --> 00:48:34,120
A function can validate that a new device belongs
1269
00:48:34,120 --> 00:48:35,320
to an allowed site.
1270
00:48:35,320 --> 00:48:36,720
It can create a pending record.
1271
00:48:36,720 --> 00:48:39,320
It can add standard tags or send a configuration request
1272
00:48:39,320 --> 00:48:41,720
through the controlled device management path.
1273
00:48:41,720 --> 00:48:43,920
Human review enters where the workflow changes.
1274
00:48:43,920 --> 00:48:46,520
Equipment ownership, network access, production use,
1275
00:48:46,520 --> 00:48:48,720
or a record that people rely on during an incident.
1276
00:48:48,720 --> 00:48:50,320
There's one small but important warning
1277
00:48:50,320 --> 00:48:52,320
around device created events.
1278
00:48:52,320 --> 00:48:53,620
The event can include twin data,
1279
00:48:53,620 --> 00:48:55,420
but Microsoft documents that this twin data
1280
00:48:55,420 --> 00:48:56,820
starts as default configuration
1281
00:48:56,820 --> 00:48:58,920
and shouldn't be treated as proof of the device's actual
1282
00:48:58,920 --> 00:49:02,320
authentication details or other newly created device properties.
1283
00:49:02,320 --> 00:49:03,820
That means your onboarding workflow
1284
00:49:03,820 --> 00:49:06,420
should verify the device through the proper device registry
1285
00:49:06,420 --> 00:49:09,120
or management interface before it trusts those details.
1286
00:49:09,120 --> 00:49:10,120
It may feel fussy,
1287
00:49:10,120 --> 00:49:11,620
but it avoids a common failure.
1288
00:49:11,620 --> 00:49:13,220
A workflow reads a creation event,
1289
00:49:13,220 --> 00:49:16,020
assumes the included twin represents the final device state
1290
00:49:16,020 --> 00:49:19,220
then assigns controls or access based on incomplete information.
1291
00:49:19,220 --> 00:49:20,220
A few minutes later,
1292
00:49:20,220 --> 00:49:22,820
the device set up changes and the records no longer match.
1293
00:49:22,820 --> 00:49:24,820
Now the team has two versions of device truth
1294
00:49:24,820 --> 00:49:26,520
which is exactly the sort of problem
1295
00:49:26,520 --> 00:49:29,520
that turns a simple gateway rollout into a week of calls.
1296
00:49:29,520 --> 00:49:31,820
Treat the life cycle event as a reliable notice
1297
00:49:31,820 --> 00:49:33,020
that something changed.
1298
00:49:33,020 --> 00:49:35,020
Then query, the authoritative source
1299
00:49:35,020 --> 00:49:37,220
for the current state needed by the next step.
1300
00:49:37,220 --> 00:49:39,220
This is why Spass action driven workflows
1301
00:49:39,220 --> 00:49:40,720
belong naturally in event grid.
1302
00:49:40,720 --> 00:49:43,920
A device gets created, deleted, connected, or disconnected.
1303
00:49:43,920 --> 00:49:45,820
A focused process receives that notice
1304
00:49:45,820 --> 00:49:47,320
and performs a defined follow-up.
1305
00:49:47,320 --> 00:49:49,320
It doesn't need to read thousands of measurements
1306
00:49:49,320 --> 00:49:51,820
just to learn that a record needs attention.
1307
00:49:51,820 --> 00:49:53,620
The details of that follow-up still depend
1308
00:49:53,620 --> 00:49:55,120
on how you filter the message flow
1309
00:49:55,120 --> 00:49:58,620
because filtering decides where the first piece of logic lives.
1310
00:49:58,620 --> 00:49:59,720
Filtering.
1311
00:49:59,720 --> 00:50:02,620
Select the right messages before they spread.
1312
00:50:02,620 --> 00:50:05,920
Filtering decides where you want the first decision
1313
00:50:05,920 --> 00:50:07,320
about a message to happen.
1314
00:50:07,320 --> 00:50:08,820
That sounds like a small conflict choice,
1315
00:50:08,820 --> 00:50:11,220
but it shapes how much traffic reaches each system,
1316
00:50:11,220 --> 00:50:12,720
how much code each team writes,
1317
00:50:12,720 --> 00:50:14,920
and how easily you can explain why a message
1318
00:50:14,920 --> 00:50:16,720
reached a given destination.
1319
00:50:16,720 --> 00:50:18,320
For IoT Hub message routing,
1320
00:50:18,320 --> 00:50:20,720
the question usually starts with the stream itself
1321
00:50:20,720 --> 00:50:23,220
which device messages belong in this path.
1322
00:50:23,220 --> 00:50:25,320
A root can filter using application properties
1323
00:50:25,320 --> 00:50:28,520
that the gateway adds, system properties supplied by IoT Hub,
1324
00:50:28,520 --> 00:50:30,020
fields in the message body,
1325
00:50:30,020 --> 00:50:32,420
and device twin tags or twin properties.
1326
00:50:32,420 --> 00:50:34,820
That gives you several ways to separate industrial traffic
1327
00:50:34,820 --> 00:50:37,120
before it reaches the downstream endpoint.
1328
00:50:37,120 --> 00:50:40,320
Say a gateway reads signals from several machines online too.
1329
00:50:40,320 --> 00:50:42,920
It may send production state, energy use diagnostics
1330
00:50:42,920 --> 00:50:44,220
and condition monitoring signals
1331
00:50:44,220 --> 00:50:45,720
through the same device connection.
1332
00:50:45,720 --> 00:50:47,520
You don't want to rely on the receiving system
1333
00:50:47,520 --> 00:50:50,320
to guess which message belongs to which process,
1334
00:50:50,320 --> 00:50:52,920
put the useful classification close to the source.
1335
00:50:52,920 --> 00:50:55,020
The gateway might add a message property
1336
00:50:55,020 --> 00:50:57,020
that identifies the signal class as energy,
1337
00:50:57,020 --> 00:50:58,320
production, or diagnostic.
1338
00:50:58,320 --> 00:51:00,720
It might include an asset ID and a source timestamp
1339
00:51:00,720 --> 00:51:01,520
in the payload.
1340
00:51:01,520 --> 00:51:03,320
The device twin may hold a managed tag
1341
00:51:03,320 --> 00:51:06,720
that identifies the gateway's plant, line, or support domain.
1342
00:51:06,720 --> 00:51:08,820
IoT Hub routing can then use those fields
1343
00:51:08,820 --> 00:51:11,620
to send only line two energy messages to an energy stream
1344
00:51:11,620 --> 00:51:14,020
while production state messages follow the path
1345
00:51:14,020 --> 00:51:15,820
built for production analysis.
1346
00:51:15,820 --> 00:51:18,020
That keeps the first split clear invisible.
1347
00:51:18,020 --> 00:51:19,320
Body queries can help too,
1348
00:51:19,320 --> 00:51:21,320
provided your message format stays stable.
1349
00:51:21,320 --> 00:51:23,320
For example, you may root messages only
1350
00:51:23,320 --> 00:51:26,120
when a state field indicates a defined operating condition
1351
00:51:26,120 --> 00:51:29,220
or only when a payload identifies a particular signal group.
1352
00:51:29,220 --> 00:51:31,420
But I'd be careful about putting all business logic
1353
00:51:31,420 --> 00:51:32,920
into a routing query.
1354
00:51:32,920 --> 00:51:34,920
Roting should make broad traffic decisions.
1355
00:51:34,920 --> 00:51:36,820
It should not become a hidden rule engine
1356
00:51:36,820 --> 00:51:39,320
that nobody on the plant side can explain or maintain.
1357
00:51:39,320 --> 00:51:41,120
Twin tags help when the classification
1358
00:51:41,120 --> 00:51:43,920
belongs to the managed device rather than each message.
1359
00:51:43,920 --> 00:51:46,320
If a gateway always belongs to a given site,
1360
00:51:46,320 --> 00:51:49,520
a twin tag can avoid repeating that fact in every payload.
1361
00:51:49,520 --> 00:51:51,320
But if a gateway observes multiple machines
1362
00:51:51,320 --> 00:51:53,020
or can move between production areas,
1363
00:51:53,020 --> 00:51:54,620
you need governance around those tags.
1364
00:51:54,620 --> 00:51:57,320
A tag that says line two while the gateway now sits
1365
00:51:57,320 --> 00:52:00,520
on line four creates a clean root carrying the wrong data.
1366
00:52:00,520 --> 00:52:02,220
Cleanly wrong still counts as wrong.
1367
00:52:02,220 --> 00:52:04,620
Event grid filtering answers a different question.
1368
00:52:04,620 --> 00:52:06,720
Instead of choosing a data path for a stream,
1369
00:52:06,720 --> 00:52:09,820
it chooses which subscribers should receive a notification.
1370
00:52:09,820 --> 00:52:12,620
An event grid subscription can filter by event type,
1371
00:52:12,620 --> 00:52:15,120
by subject and by fields in the event data.
1372
00:52:15,120 --> 00:52:18,320
For IoT Hub events, the subject identifies the device path.
1373
00:52:18,320 --> 00:52:21,120
That lets a subscriber focus on a defined device set,
1374
00:52:21,120 --> 00:52:24,620
such as gateways from one side, one line, or one naming group.
1375
00:52:24,620 --> 00:52:26,520
A device support workflow might subscribe
1376
00:52:26,520 --> 00:52:27,820
only to disconnect the events
1377
00:52:27,820 --> 00:52:29,620
from a group of production gateways.
1378
00:52:29,620 --> 00:52:33,120
An onboarding service might subscribe only to device-created events.
1379
00:52:33,120 --> 00:52:36,020
A telemetry-triggered workflow might filter for messages
1380
00:52:36,020 --> 00:52:37,920
whose data matches a narrow condition
1381
00:52:37,920 --> 00:52:39,720
that its owner has agreed to handle.
1382
00:52:39,720 --> 00:52:41,020
The difference matters.
1383
00:52:41,020 --> 00:52:43,020
IoT Hub root filtering decides
1384
00:52:43,020 --> 00:52:45,920
which operational messages should enter this data path.
1385
00:52:45,920 --> 00:52:48,220
Event grid subscription filtering decides
1386
00:52:48,220 --> 00:52:50,520
which notices should this handler receive.
1387
00:52:50,520 --> 00:52:52,520
Both reduce noise, they do different jobs.
1388
00:52:52,520 --> 00:52:55,020
If you use event grid filters as the main way
1389
00:52:55,020 --> 00:52:57,020
to segment continuous machine data,
1390
00:52:57,020 --> 00:52:59,020
each subscriber can end up with its own version
1391
00:52:59,020 --> 00:53:00,720
of what counts as relevant.
1392
00:53:00,720 --> 00:53:02,820
One handler filters by device prefix,
1393
00:53:02,820 --> 00:53:04,720
another filters by a payload field.
1394
00:53:04,720 --> 00:53:06,220
A third ignores a message property
1395
00:53:06,220 --> 00:53:08,020
and recreates the logic in code.
1396
00:53:08,020 --> 00:53:11,120
Over time, the architecture spreads basic data selection
1397
00:53:11,120 --> 00:53:12,820
across too many places.
1398
00:53:12,820 --> 00:53:14,920
That makes changes harder than they need to be.
1399
00:53:14,920 --> 00:53:17,020
A shared telemetry root should carry an explicit
1400
00:53:17,020 --> 00:53:18,520
governed class of data.
1401
00:53:18,520 --> 00:53:20,920
Event subscribers should receive a focused signal
1402
00:53:20,920 --> 00:53:22,820
that starts a specific reaction.
1403
00:53:22,820 --> 00:53:25,720
That boundary lets the data team own stream rules
1404
00:53:25,720 --> 00:53:28,120
while the workflow team owns response rules
1405
00:53:28,120 --> 00:53:31,720
without either group quietly rewriting the other groups logic.
1406
00:53:31,720 --> 00:53:34,620
Filtering also needs a versioning plan, payloads change,
1407
00:53:34,620 --> 00:53:36,020
new signal types appear,
1408
00:53:36,020 --> 00:53:38,320
a gateway firmware update may rename a field,
1409
00:53:38,320 --> 00:53:39,620
add a nested structure,
1410
00:53:39,620 --> 00:53:41,320
or change how it reports a state.
1411
00:53:41,320 --> 00:53:43,720
If roots and subscriptions depend on those fields,
1412
00:53:43,720 --> 00:53:45,420
test the change against the whole path
1413
00:53:45,420 --> 00:53:46,920
before it reaches production.
1414
00:53:46,920 --> 00:53:50,120
Otherwise, a message can still reach IoT Hub successfully
1415
00:53:50,120 --> 00:53:52,920
while disappearing from the root or subscription that matters.
1416
00:53:52,920 --> 00:53:54,820
From the device side, everything looks fine.
1417
00:53:54,820 --> 00:53:56,920
From the business side, the data has vanished
1418
00:53:56,920 --> 00:53:59,120
and filtering cannot repair missing meaning.
1419
00:53:59,120 --> 00:54:02,020
It can separate line two energy telemetry from press diagnostics.
1420
00:54:02,020 --> 00:54:04,320
It cannot tell you whether a high energy reading came
1421
00:54:04,320 --> 00:54:06,420
during normal production, a setup run,
1422
00:54:06,420 --> 00:54:08,820
a maintenance task, or an abnormal condition
1423
00:54:08,820 --> 00:54:10,920
that context has to come from somewhere else.
1424
00:54:10,920 --> 00:54:14,720
Context comes from ERP, MES, and the asset model.
1425
00:54:14,720 --> 00:54:19,220
A device ID tells you where a message came from.
1426
00:54:19,220 --> 00:54:21,520
It doesn't tell you what the message means to production.
1427
00:54:21,520 --> 00:54:24,420
That gap appears the moment somebody asks a practical question.
1428
00:54:24,420 --> 00:54:27,020
A gateway reports that a press has stopped sending data.
1429
00:54:27,020 --> 00:54:29,120
Is that press in production? Is it being set up?
1430
00:54:29,120 --> 00:54:30,420
Is it under planned maintenance?
1431
00:54:30,420 --> 00:54:32,920
Does it run an order that another machine can take over?
1432
00:54:32,920 --> 00:54:35,620
Or is it the only resource qualified for that operation?
1433
00:54:35,620 --> 00:54:37,720
You won't find those answers in a device ID.
1434
00:54:37,720 --> 00:54:40,120
The MES, your manufacturing execution system,
1435
00:54:40,120 --> 00:54:42,220
supplies much of the shop floor context.
1436
00:54:42,220 --> 00:54:44,720
It can connect an asset signal to an active operation,
1437
00:54:44,720 --> 00:54:47,920
a work order, a recipe, or rooting step, a shift,
1438
00:54:47,920 --> 00:54:50,120
and the production status recorded by the people
1439
00:54:50,120 --> 00:54:51,720
in systems running the line.
1440
00:54:51,720 --> 00:54:54,220
That link turns a raw state such as stopped
1441
00:54:54,220 --> 00:54:55,320
into something more useful.
1442
00:54:55,320 --> 00:54:57,920
For example, the same machine state may mean
1443
00:54:57,920 --> 00:55:00,920
very different things depending on the MES record.
1444
00:55:00,920 --> 00:55:04,320
A stop press during a planned die change is expected.
1445
00:55:04,320 --> 00:55:07,220
A stopped press during a high priority production order
1446
00:55:07,220 --> 00:55:08,420
might need attention.
1447
00:55:08,420 --> 00:55:11,520
A stopped press after the order completed may not affect output at all.
1448
00:55:11,520 --> 00:55:12,720
The sensor value didn't change.
1449
00:55:12,720 --> 00:55:15,420
The context changed. ERP brings another layer.
1450
00:55:15,420 --> 00:55:17,620
It usually knows the demand side of the story.
1451
00:55:17,620 --> 00:55:19,520
Customer commitments, material availability,
1452
00:55:19,520 --> 00:55:22,020
order dates, inventory positions, and the wider plan.
1453
00:55:22,020 --> 00:55:24,720
If a production issue lasts long enough to threaten output,
1454
00:55:24,720 --> 00:55:28,120
ERP context helps the business understand where the pressure lands.
1455
00:55:28,120 --> 00:55:30,620
But ERP shouldn't pretend to know the current machine condition
1456
00:55:30,620 --> 00:55:32,420
just because it contains a production order.
1457
00:55:32,420 --> 00:55:35,520
Its plan describes intent, the MES and the machine layer
1458
00:55:35,520 --> 00:55:36,820
describe execution.
1459
00:55:36,820 --> 00:55:40,320
Those systems need to connect, but they shouldn't overwrite each other's role.
1460
00:55:40,320 --> 00:55:42,220
Think about a plan on looking at a late order.
1461
00:55:42,220 --> 00:55:45,220
ERP may show a promised delivery date and the order quantity.
1462
00:55:45,220 --> 00:55:49,520
MES may show that the operation started, paused, or completed only in part.
1463
00:55:49,520 --> 00:55:52,520
Telemetry may show rising vibration before a stop.
1464
00:55:52,520 --> 00:55:56,320
The asset model may show which other machines can perform the same process.
1465
00:55:56,320 --> 00:55:59,020
Only together do those facts support a sound decision.
1466
00:55:59,020 --> 00:56:02,320
The asset model provides the physical structure that links technical devices
1467
00:56:02,320 --> 00:56:03,820
to equipment people recognize.
1468
00:56:03,820 --> 00:56:05,620
A gateway belongs to a network zone.
1469
00:56:05,620 --> 00:56:07,220
It reads one or more PLCs.
1470
00:56:07,220 --> 00:56:08,620
Those PLCs relate to machines.
1471
00:56:08,620 --> 00:56:11,320
Machines sit in cells, lines, departments, and sites.
1472
00:56:11,320 --> 00:56:15,920
Without that model, every downstream system ends up rebuilding its own mapping table.
1473
00:56:15,920 --> 00:56:17,820
One team calls the asset press 12.
1474
00:56:17,820 --> 00:56:20,220
Another knows the gateway as GW-037.
1475
00:56:20,220 --> 00:56:21,820
The MES uses a resource code.
1476
00:56:21,820 --> 00:56:24,520
The maintenance system uses a different equipment number.
1477
00:56:24,520 --> 00:56:27,620
Then a simple question such as, "Which line lost data?"
1478
00:56:27,620 --> 00:56:30,120
Becomes a joint problem with human consequences.
1479
00:56:30,120 --> 00:56:32,520
That's not unusual. It's just expensive confusion.
1480
00:56:32,520 --> 00:56:35,820
In practical terms, the asset model should give you stable relationships.
1481
00:56:35,820 --> 00:56:37,720
This device reports for this gateway.
1482
00:56:37,720 --> 00:56:39,320
This gateway collects from this machine.
1483
00:56:39,320 --> 00:56:41,120
This machine belongs to this line.
1484
00:56:41,120 --> 00:56:43,420
This line can produce these product families.
1485
00:56:43,420 --> 00:56:45,320
Under these approved process rules,
1486
00:56:45,320 --> 00:56:47,820
you don't need to put every relationship into IoT Hub.
1487
00:56:47,820 --> 00:56:51,920
You do need one governed place where those relationships can be maintained and consumed.
1488
00:56:51,920 --> 00:56:55,820
For some manufacturers, a structured asset register plus MES master data is enough.
1489
00:56:55,820 --> 00:56:58,420
Others need a digital twin or a knowledge graph
1490
00:56:58,420 --> 00:57:01,120
because they need to reason across changing relationships.
1491
00:57:01,120 --> 00:57:07,920
Assets, sensors, processes, materials, workers, quality checks, and production constraints.
1492
00:57:07,920 --> 00:57:11,020
A knowledge graph sounds like a big idea, but the plain version is simple.
1493
00:57:11,020 --> 00:57:12,620
It records how things relate.
1494
00:57:12,620 --> 00:57:15,520
Instead of storing a machine as one isolated record,
1495
00:57:15,520 --> 00:57:19,620
it can represent that the machine uses a certain tool, runs a certain process,
1496
00:57:19,620 --> 00:57:21,820
receives data from a certain gateway,
1497
00:57:21,820 --> 00:57:24,020
and supports a defined set of operations.
1498
00:57:24,020 --> 00:57:28,320
When something changes, systems can follow those links rather than asking people to remember them.
1499
00:57:28,320 --> 00:57:31,020
That's useful when the question moves beyond technical support.
1500
00:57:31,020 --> 00:57:33,620
A device disconnected notice may point to a gateway.
1501
00:57:33,620 --> 00:57:36,020
The asset model can identify the machine in line.
1502
00:57:36,020 --> 00:57:38,720
MES can show the active work order and current operation.
1503
00:57:38,720 --> 00:57:42,120
ERP can show whether delayed output affects downstream demand.
1504
00:57:42,120 --> 00:57:44,120
The event still doesn't decide anything by itself,
1505
00:57:44,120 --> 00:57:46,520
but it reaches a system that can ask better questions.
1506
00:57:46,520 --> 00:57:47,320
And that's the point.
1507
00:57:47,320 --> 00:57:51,020
Integration isn't only about moving messages between Azure services.
1508
00:57:51,020 --> 00:57:53,720
It's about connecting the dots between IT and OT
1509
00:57:53,720 --> 00:57:57,020
without turning every raw signal into a false business conclusion.
1510
00:57:57,020 --> 00:57:58,320
So let's return to the press line.
1511
00:57:58,320 --> 00:58:00,920
The gateway goes offline while a work order is active,
1512
00:58:00,920 --> 00:58:04,220
and now the architecture has to separate a loss of cloud visibility
1513
00:58:04,220 --> 00:58:06,020
from an actual production problem.
1514
00:58:06,020 --> 00:58:10,320
Scenario, the gateway goes offline during a work order.
1515
00:58:10,320 --> 00:58:11,620
Let's make the failure concrete.
1516
00:58:11,620 --> 00:58:14,220
Press 12 is halfway through an active work order.
1517
00:58:14,220 --> 00:58:16,020
The operator has the press running.
1518
00:58:16,020 --> 00:58:18,320
The MES shows the operation is active,
1519
00:58:18,320 --> 00:58:21,520
and the gateway sends normal process data into IoT Hub.
1520
00:58:21,520 --> 00:58:23,820
Then the gateway loses its connection to the cloud.
1521
00:58:23,820 --> 00:58:26,920
An event grid subscription receives a device disconnected event
1522
00:58:26,920 --> 00:58:28,620
and starts a support workflow.
1523
00:58:28,620 --> 00:58:29,920
That's a sensible first move.
1524
00:58:29,920 --> 00:58:32,220
The workflow creates an investigation record,
1525
00:58:32,220 --> 00:58:34,120
adds the device ID, the event time,
1526
00:58:34,120 --> 00:58:36,620
and the support owner for that production area.
1527
00:58:36,620 --> 00:58:37,920
It should not stop the press.
1528
00:58:37,920 --> 00:58:39,620
The cloud lost contact with a gateway,
1529
00:58:39,620 --> 00:58:41,320
that is all the event proves.
1530
00:58:41,320 --> 00:58:43,720
The press may still run under local PLC control.
1531
00:58:43,720 --> 00:58:45,520
The operator may still produce parts,
1532
00:58:45,520 --> 00:58:47,120
or the gateway may have failed in a way
1533
00:58:47,120 --> 00:58:48,720
that also removed data collection
1534
00:58:48,720 --> 00:58:51,020
from the line while production continues normally.
1535
00:58:51,020 --> 00:58:52,520
Those are different situations.
1536
00:58:52,520 --> 00:58:56,120
The first check should focus on the gateway and the telemetry path.
1537
00:58:56,120 --> 00:58:58,120
A function triggered by the event can inspect
1538
00:58:58,120 --> 00:59:01,320
when IoT Hub last received a valid message from that gateway.
1539
00:59:01,320 --> 00:59:03,120
It can compare the last source timestamp
1540
00:59:03,120 --> 00:59:04,520
with the cloud receipt time
1541
00:59:04,520 --> 00:59:07,120
and look for a sudden gap in production telemetry.
1542
00:59:07,120 --> 00:59:09,520
That gives the support team a better starting point.
1543
00:59:09,520 --> 00:59:12,420
If the last message is showed press 12 in a running state
1544
00:59:12,420 --> 00:59:14,120
and telemetry stops abruptly,
1545
00:59:14,120 --> 00:59:16,220
the plant has lost visibility after that point.
1546
00:59:16,220 --> 00:59:18,320
It has not yet proved that the press stopped.
1547
00:59:18,320 --> 00:59:20,820
If Buffet messages arrive when the gateway reconnects,
1548
00:59:20,820 --> 00:59:22,820
they may fill part of the timeline later
1549
00:59:22,820 --> 00:59:25,120
and the incident becomes a data completeness issue
1550
00:59:25,120 --> 00:59:26,920
rather than a production interruption.
1551
00:59:26,920 --> 00:59:28,920
That distinction saves a lot of bad escalation.
1552
00:59:28,920 --> 00:59:29,920
Now bring in the mess.
1553
00:59:29,920 --> 00:59:32,720
The MES can confirm whether the operation remained active,
1554
00:59:32,720 --> 00:59:34,320
whether the operator recorded a stop,
1555
00:59:34,320 --> 00:59:35,920
whether parts continued to count
1556
00:59:35,920 --> 00:59:38,420
and whether a reason code appeared.
1557
00:59:38,420 --> 00:59:40,820
In some plants, the MES receives production data
1558
00:59:40,820 --> 00:59:42,120
through another local path.
1559
00:59:42,120 --> 00:59:44,020
In others, the operator records the change.
1560
00:59:44,020 --> 00:59:47,220
Either way, the MES helps separate a cloud connection issue
1561
00:59:47,220 --> 00:59:49,120
from the execution state of the work order.
1562
00:59:49,120 --> 00:59:51,120
Suppose the MES still records cycles
1563
00:59:51,120 --> 00:59:53,020
and the operator has no stop reason.
1564
00:59:53,020 --> 00:59:54,520
The press likely kept running.
1565
00:59:54,520 --> 00:59:56,620
The immediate task becomes checking the gateway,
1566
00:59:56,620 --> 00:59:57,720
the OT network path,
1567
00:59:57,720 --> 01:00:00,120
and whether missing telemetry will affect traceability
1568
01:00:00,120 --> 01:00:01,820
or condition analysis.
1569
01:00:01,820 --> 01:00:04,220
Suppose instead the MES shows the operation paused
1570
01:00:04,220 --> 01:00:06,320
and the operator records an equipment fault.
1571
01:00:06,320 --> 01:00:09,220
Now the event grid notification and the telemetry gap
1572
01:00:09,220 --> 01:00:11,320
point toward a real operational issue.
1573
01:00:11,320 --> 01:00:13,620
But the planner still needs more than a disconnect event
1574
01:00:13,620 --> 01:00:15,220
before changing the production plan.
1575
01:00:15,220 --> 01:00:16,020
They need duration.
1576
01:00:16,020 --> 01:00:17,520
They need to know whether the press has stopped
1577
01:00:17,520 --> 01:00:19,920
for a few minutes, whether maintenance can restore it,
1578
01:00:19,920 --> 01:00:22,820
and whether the operation can move to another approved resource.
1579
01:00:22,820 --> 01:00:25,420
The production impact comes after the operational facts,
1580
01:00:25,420 --> 01:00:27,420
not from the first cloud notification.
1581
01:00:27,420 --> 01:00:28,720
That order of reasoning matters.
1582
01:00:28,720 --> 01:00:31,520
A planner may see that the work order sits on press 12.
1583
01:00:31,520 --> 01:00:34,520
ERP can show demand dates and downstream commitments.
1584
01:00:34,520 --> 01:00:36,520
But neither system should assume that every gateway
1585
01:00:36,520 --> 01:00:38,220
disconnect threatens delivery.
1586
01:00:38,220 --> 01:00:39,820
If the press continues locally,
1587
01:00:39,820 --> 01:00:43,020
re-planning creates churn for no operational reason.
1588
01:00:43,020 --> 01:00:44,320
If the press truly stopped,
1589
01:00:44,320 --> 01:00:46,720
waiting for a dashboard refresh may waste time.
1590
01:00:46,720 --> 01:00:49,020
The architecture needs both speed and restraint.
1591
01:00:49,020 --> 01:00:50,620
Event grid provides the fast signal
1592
01:00:50,620 --> 01:00:52,120
that starts the investigation.
1593
01:00:52,120 --> 01:00:55,120
IoT Hub telemetry provides the last known machine facts
1594
01:00:55,120 --> 01:00:57,920
and potentially the backfilled record after reconnection.
1595
01:00:57,920 --> 01:00:59,720
MES confirms what production actually did.
1596
01:00:59,720 --> 01:01:02,820
The asset model identifies the affected machine in line.
1597
01:01:02,820 --> 01:01:04,920
Only then can a planner judge whether the work order
1598
01:01:04,920 --> 01:01:06,520
faces a real constraint.
1599
01:01:06,520 --> 01:01:08,020
Consider a more awkward case.
1600
01:01:08,020 --> 01:01:09,720
The gateway reconnects after 10 minutes
1601
01:01:09,720 --> 01:01:11,620
and event grid sends a connected notice.
1602
01:01:11,620 --> 01:01:13,920
That does not close the production question either.
1603
01:01:13,920 --> 01:01:16,920
The function can confirm that cloud connectivity returned,
1604
01:01:16,920 --> 01:01:18,220
but the team still needs to check
1605
01:01:18,220 --> 01:01:20,720
whether the gateway resume PLC reads,
1606
01:01:20,720 --> 01:01:22,220
whether buffer data arrived,
1607
01:01:22,220 --> 01:01:24,820
and whether the production operation continued.
1608
01:01:24,820 --> 01:01:26,920
A green connection status can end a network incident
1609
01:01:26,920 --> 01:01:28,720
while leaving a data incident open.
1610
01:01:28,720 --> 01:01:30,720
This is why I prefer an investigation workflow
1611
01:01:30,720 --> 01:01:32,020
with clear states.
1612
01:01:32,020 --> 01:01:33,520
First cloud connection lost,
1613
01:01:33,520 --> 01:01:35,120
then telemetry path checked,
1614
01:01:35,120 --> 01:01:37,220
then production state confirmed through MES
1615
01:01:37,220 --> 01:01:38,720
or local control data.
1616
01:01:38,720 --> 01:01:41,520
Finally, impact assessed if the machine actually stopped
1617
01:01:41,520 --> 01:01:44,520
or if data lost threatens a quality or traceability record.
1618
01:01:44,520 --> 01:01:46,720
Each state has an owner and a source of evidence
1619
01:01:46,720 --> 01:01:48,820
that may sound more deliberate than firing off an alert
1620
01:01:48,820 --> 01:01:50,820
and assuming someone sorts it out.
1621
01:01:50,820 --> 01:01:52,820
But a work order doesn't care that an Azure event
1622
01:01:52,820 --> 01:01:53,720
arrived quickly.
1623
01:01:53,720 --> 01:01:56,320
It cares whether the press produced the required parts,
1624
01:01:56,320 --> 01:01:58,520
whether the process stayed within limits
1625
01:01:58,520 --> 01:02:01,320
and whether the plant can still meet the next commitment.
1626
01:02:01,320 --> 01:02:04,020
The same incident makes the architectural boundary clear.
1627
01:02:04,020 --> 01:02:07,020
Neither IoT have rooting nor event grid replaces the other
1628
01:02:07,020 --> 01:02:08,920
because one preserves the operational record
1629
01:02:08,920 --> 01:02:10,920
while the other starts the right response.
Apple Podcasts
Spotify
Youtube Music
Spreaker
Podchaser
Amazon Music


