A Complete Guide to Data Integration and Analytics Pipelines

Data Factory, Synapse Analytics, Event Hubs, Stream Analytics — compare AZ-305 data integration services by scenario, from Integration Runtime to lambda architecture.

Data integration and analytics pipelines account for a significant portion of the AZ-305 exam. Simply memorizing service names is not enough — you need to accurately distinguish which service fits which scenario. From on-premises ETL pipelines to real-time streaming that processes millions of events per second, this guide walks through the key services and selection criteria based on 29 exam scenario patterns.

---

 

Data Factory and ETL/ELT Pipelines

Azure Data Factory (ADF) is a fully managed ETL/ELT service that lets you build data pipelines visually without writing code. It provides over 80 connectors spanning on-premises databases like SQL Server, MySQL, and Oracle, as well as cloud storage such as Blob Storage, Azure Data Lake Storage Gen2, and Cosmos DB.

ADF has three core building blocks. Copy Activity copies data from a source to a destination and includes built-in column mapping and format conversion. Mapping Data Flow lets you define filtering, aggregation, and joins through a drag-and-drop interface — no code required — and executes them on an internal Spark cluster. Integration Runtime (IR) is the compute infrastructure that actually moves the data; there are three types: Azure IR, Self-hosted Integration Runtime (SHIR), and Azure-SSIS IR.

When you need to reach an on-premises server behind a firewall, Self-hosted Integration Runtime (SHIR) is mandatory. It is an agent installed on the on-premises server that communicates with ADF over outbound HTTPS only — no inbound ports required. On the exam, whenever you see conditions like "on-premises behind a firewall" or "local server with no internet connectivity," SHIR is the core of the correct answer.

If you have hundreds of existing SSIS packages and want to run them in Azure without rewriting them, use Azure-SSIS IR. Deploy the packages to SSISDB and run them exactly as before using the Execute SSIS Package activity in an ADF pipeline.

ADF supports schedule triggers, event triggers, and tumbling window triggers. It ships with built-in retry policies and a monitoring hub, keeping operational overhead low. For incremental ingestion, you can use either a watermark-based approach or Change Data Capture to selectively pull only new or changed data.

---

 

Synapse Analytics and Data Warehousing

Azure Synapse Analytics is a unified service that combines data warehousing and big data analytics on a single platform. Inside Synapse Analytics there are three main analytical engines.

The first is the Dedicated SQL Pool. It uses an MPP (Massively Parallel Processing) architecture that distributes queries across 60 distributed nodes in parallel. It is the optimal choice for large-scale data warehouse scenarios where hundreds of users run aggregation queries concurrently. Choosing a Hash distribution key based on columns frequently used in joins and aggregations minimizes inter-node data movement (DMS traffic) and maximizes query performance.

The second is the Serverless SQL Pool. It lets you query Parquet, CSV, and JSON files stored in Azure Data Lake Storage Gen2 directly with T-SQL, with no infrastructure to manage. You are billed per query, which makes it cost-effective for ad hoc exploratory analysis.

The third is the Synapse Spark Pool. It handles complex transformations and ML pipelines written in Python, Scala, or R, and integrates with Delta Lake for ACID transactions (upsert/delete). Auto-pause brings the cost to zero automatically when there is no active work.

Synapse Analytics workspaces also include Synapse Pipelines as a built-in feature. They share the same code base as ADF and offer over 400 connectors. Choose standalone ADF when you need an independent ETL service; choose Synapse Pipelines when you want integrated management within an existing Synapse workspace. To offload the Power BI semantic layer, add Azure Analysis Services on top of Synapse — this prevents thousands of concurrent BI users from placing load directly on the Dedicated SQL Pool.

---

 

Messaging and Events: Event Hubs, Event Grid, Service Bus

These three services look similar by name, but their design philosophies are fundamentally different.

Azure Event Hubs is a service purpose-built for ingesting high-volume event streams. It reliably buffers millions of events per second, and consumers read at their own pace, independently of one another. Its partition-based parallel architecture completely decouples the ingestion layer from the processing layer, and events are retained for a configurable retention period of up to 90 days. With Event Hubs Capture, events are automatically archived to Azure Data Lake Storage Gen2 in Apache Avro format, ready for cold-path batch analysis.

Azure Event Grid routes Azure resource events using a publish-subscribe model. It is well suited for notification events such as Blob file uploads or resource state changes, with low latency and broad integration across Azure services.

Azure Service Bus is an enterprise message broker optimized for reliable, asynchronous message delivery. It supports two patterns: Queue and Topic/Subscription.

A Queue follows the competing consumer pattern — exactly one receiver processes each message. It is the right choice for scenarios that require transactional integrity and zero message loss. The Session feature ensures that messages sharing the same SessionId are always processed in order by the same consumer instance, which is valuable for order-sensitive workflows like airline ticket reservations.

Topic/Subscription implements the Publish-Subscribe pattern. When a message is published to a topic, each subscription maintains an independent copy so that multiple consumers can receive it at their own pace. You can also add filter rules to let each subscription receive only the messages that match specific conditions.

---

 

Real-Time Analytics: Stream Analytics and Databricks

Azure Stream Analytics is a fully managed service that applies SQL-like queries to event streams for real-time analysis. It processes aggregations, filters, and joins on streams arriving from Event Hubs or IoT Hub, and instantly writes results to Power BI, SQL Database, Cosmos DB, and more. No servers to provision — you scale processing capacity with Streaming Units (SU).

Azure Databricks is a managed analytics platform built on Apache Spark. It supports large-scale batch processing, ML model training, and real-time processing via Structured Streaming — all in one place. It is the right fit when you need complex transformation logic, custom ML pipelines, or deep integration with open-source libraries.

Azure Data Explorer is designed for ultra-low-latency processing of petabyte-scale text-based logs and time-series data. It uses KQL (Kusto Query Language) and excels at scenarios where hundreds of analysts run interactive queries simultaneously — security log analysis and network traffic pattern detection are prime examples.

To summarize the selection criteria: use Stream Analytics for real-time SQL aggregation and immediate dashboard output; use Databricks for complex ML/Spark code and combined batch-plus-streaming pipelines; use Azure Data Explorer for interactive exploration of log and time-series data.

---

 

Service Comparison Tables

| Attribute | Event Hubs | Event Grid | Service Bus | |-----------|-----------|-----------|-------------| | Primary use | High-volume event stream ingestion | Event routing and notifications | Reliable async messaging | | Model | Streaming (partition-based pull) | Pub-Sub event push | Queue / Topic+Subscription | | Message retention | Up to 90 days (retained after consumption) | Expires after 24-hour retry | Deleted after successful processing | | Throughput | Millions of events/sec | Up to tens of millions of events/sec | Thousands of messages/sec | | Ordering guarantee | Ordered within a partition | No guarantee | Ordered via Session feature | | Typical scenario | IoT telemetry, audit log ingestion | File upload notifications, resource changes | Order processing, workflow messaging |

| Attribute | Azure Data Factory | Synapse Pipelines | |-----------|------------------|-----------------| | Independence | Standalone ETL service | Built into Synapse workspace | | Code base | Identical (shared) | Identical (shared) | | Connectors | 90+ | 400+ (ADF-based) | | When to choose | When you need a standalone ETL tool | When integrating within an existing Synapse environment |

---

 

Selection Criteria That Often Confuse Exam Candidates

Scenario 1 — Accessing an on-premises server behind a firewall: Self-hosted IR is required. It communicates with ADF over outbound HTTPS only, with no inbound ports opened.

Scenario 2 — Multiple consumers each receiving the same message at their own pace: Service Bus Topic/Subscription. Each subscription maintains an independent copy of the message, and filter rules allow selective receipt. Event Hubs is designed for stream ingestion, and a Queue serves only a single consumer at a time.

Scenario 3 — Migrating 200 TB over limited bandwidth: Use Azure Data Box (a physical device). Handle incremental sync after the migration with AzCopy.

Scenario 4 — Running SSIS packages in Azure without rewriting them: ADF's Azure-SSIS IR is the only way to accomplish this.

Scenario 5 — MPP aggregation vs. interactive log exploration: For concurrent aggregation by hundreds of users, choose Synapse Dedicated SQL Pool. For interactive exploration of petabyte-scale logs, choose Azure Data Explorer with KQL.

---

 

Practical Tips for Real-World Application

Data pipelines frequently call for a lambda architecture that separates the hot path (real-time) from the cold path (batch). The typical flow looks like this: Event Hubs ingests events, Stream Analytics handles hot-path real-time aggregation, and Event Hubs Capture automatically archives events to Azure Data Lake Storage Gen2 in Avro format. ADF or Synapse Pipelines then refine and tran

Back to blog list