The monitoring domain accounts for 10–15% of the AZ-104 exam. Rather than simply memorizing feature names, understanding why each tool is needed will help you retain the information far longer.
---
Azure Monitor
What is Azure Monitor?
Azure Monitor is like the dashboard of a car. While driving, you glance at the speedometer, fuel gauge, engine temperature, and other readings at a glance. If something goes wrong, a warning light comes on and a record is kept. Azure Monitor works the same way — it monitors the real-time status of every resource running in the cloud, including servers, databases, and networks, sends alerts when problems arise, and keeps records for later analysis.
Why is it necessary? In a cloud environment, dozens or even hundreds of resources run simultaneously. It is impossible for a person to check each one manually if something goes wrong. Azure Monitor acts as a '24-hour security guard' that automatically watches over all of them.
---
Metrics: A Health Report in Numbers
Metrics are data that express the state of a resource in numbers. For example: CPU usage at 75%, memory usage at 4 GB, or network inbound traffic at 100 MB per second.
Using a human analogy, metrics are like body temperature, blood pressure, and pulse. If the values are within a normal range, everything is healthy; if they fall outside the range, it is a signal that something is wrong.
Key characteristics of metrics:
Automatic collection: Most Azure services automatically collect metrics without any additional configuration. When you create a virtual machine, Azure automatically records CPU, memory, and disk usage. Retention period: Collected metric data is retained for 93 days. To keep it longer, you must export it separately. Metrics Explorer: Opening Metrics Explorer in the Azure Portal lets you visualize data for any time period as line charts, bar charts, and more. It is like viewing the trend of your step count over the past month in a smartwatch app.
Real-world scenario: A report comes in that a web server has suddenly become slow. You check the CPU usage for the past hour in Metrics Explorer and see that it spiked to 95% starting at 2 p.m. You now have a clue to begin identifying the root cause.
---
Logs: A Detailed Record of Events
If metrics are 'numbers,' logs are 'sentences.' They record what happened, when, who did it, and how — like "At 2:15 p.m., user Hong Gildong logged in" or "At 2:16 p.m., a file deletion request was denied."
It is similar to an office entry log. It keeps a detailed record of who came in at what time, which rooms they accessed, and what actions they took.
Logs by default remain inside Azure resources. To analyze them, you must send them to a Log Analytics workspace via Diagnostic Settings. Think of Diagnostic Settings as the 'delivery address that specifies where to send the logs.'
!Metrics versus Logs
Log Analytics: A Tool for Searching and Analyzing Logs
Log Analytics is a space where log data is stored and can be searched. Using a library analogy, Log Analytics is the library, logs are the books, and KQL (Kusto Query Language) is the search system used to find books.
KQL is a query language similar to SQL. It may look complex at first, but knowing the basic patterns is enough.
Example of basic KQL syntax:
Breaking this query down in plain language: — Open the Azure activity log table — Look only at entries generated within the last 24 hours — Filter to only those with an error level — Count how many there are per resource group
It is similar to using Excel's filter and aggregation features via commands.
Real-world scenario: You suspect someone deleted resources without authorization last night. By querying the deletion operation logs between 10 p.m. and midnight using KQL in Log Analytics, you can find out exactly who deleted which resource.
---
Alerts: An Alarm That Immediately Notifies You of Anomalies
Alerts are a feature that 'automatically sends a notification when a specific condition is met.' It is similar to setting an alarm on a smartphone. Just as you set 'ring an alarm every day at 7 a.m.,' an Azure alert lets you configure 'send me an email if CPU exceeds 80%.'
Alerts consist of three components:
| Component | Role | Analogy | |-----------|------|---------| | Alert Rule | Defines under what conditions a notification is sent | Setting the alarm time | | Action Group | Defines who to notify and how | What to do when the alarm goes off | | Alert Processing Rule | Suppresses or adjusts notifications during specific periods | Turning the alarm off on weekends |
An Action Group is a configuration that bundles together the methods and recipients for notifications. In addition to email, SMS, and phone calls, you can also trigger automated responses via Webhook (sending a signal to another system), Logic Apps, or Azure Functions. For example, you can configure an action that automatically adds another server when CPU exceeds 90%.
Alert Processing Rules define when notifications should not be sent. During system maintenance windows, notifications can flood in, and alert processing rules can temporarily suppress them.
Real-world scenario: You can configure an alert so that when a shopping mall server's CPU usage exceeds 80%, the entire operations team receives an email and SMS notification, while simultaneously triggering an auto-scaling script.
---
Insights: Customized Monitoring Dashboards per Service
Insights are monitoring screens optimized for specific services. While the general Metrics Explorer allows you to combine any data however you like, it can be overwhelming at first to know what to look at. Insights are 'pre-arranged custom dashboards that show only the most important things when monitoring a given service.'
Just as each auto repair shop has specialized diagnostic equipment, each service has its own optimized Insights:
| Insights | Target Service | Key Information Provided | |----------|---------------|-------------------------| | VM Insights | Virtual Machines | CPU, memory, disk performance + a dependency map showing which processes communicate with which servers | | Storage Insights | Storage Accounts | Availability, response time, transaction status | | Network Insights | Network Resources | Overall network topology and performance status |
In particular, VM Insights' dependency map is extremely useful. Because it visually shows which servers a virtual machine communicates with, it helps you quickly understand the blast radius when an incident occurs.
---
Network Watcher
What is Network Watcher?
Network Watcher is a collection of specialized tools for diagnosing network problems. If Azure Monitor shows the overall health status, Network Watcher is the 'plumbing expert' that digs deep into network piping issues.
Network problems are especially difficult because they are invisible. In situations where "the server is up but why can't I connect?", the tools in Network Watcher shine.
Key Tools in Network Watcher
IP Flow Verify
This is the tool that answers the question "Is this traffic being blocked by the firewall (NSG)?" It instantly tells you whether a packet sent from a specific IP to a specific port is allowed or blocked, and which NSG rule is responsible. It is like asking at a security gate, "Can I pass through this door with this card?"
Next Hop
This is the tool that answers the question "Where is this packet headed?" It shows the actual path that a packet takes while traveling from A to B. It is similar to a navigation app telling you the direction to the next intersection. It is useful for finding the cause when traffic is going somewhere unexpected due to a misconfigured routing table.
Connection Monitor
This continuously monitors whether the connection between two endpoints is working properly. Rather than a one-time check, it automatically tests the connection every 5 minutes and records response time and packet loss rate. You can use it like surveillance ensuring that the VPN connection between the head office and a branch is always stable.
NSG Flow Logs
This records all traffic that passes through or is blocked by an NSG (Network Security Group). Like CCTV footage, you can later query "Which IPs attempted to connect to our server last week?" It is essential for security audits and abnormal traffic analysis.
Packet Capture
This records the actual network packets going to and from a virtual machine. Like recording the content of a phone call, it saves the actual data being exchanged to a file. It is used when the deepest level of analysis is needed, and the saved file can be analyzed with tools like Wireshark.
---