Kinesis Deep Dive

Kinesis isn't a single service — it's a family of four services.

Amazon Kinesis Complete Guide (DEA-C01 Exam Essentials)

Introduction

Amazon Kinesis is one of the most frequently tested services in the Domain 1: Data Ingestion and Transformation (34%) section of the DEA-C01 exam.

Kinesis is not a single service but rather a collection of services designed for real-time streaming. The exam particularly focuses on the differences between each service and the architecture of Kinesis Data Streams (KDS).

This guide covers the key concepts you need to know for exam preparation.

---

The Four Kinesis Services at a Glance

Kinesis consists of four distinct services. Since each serves a clearly different purpose, the exam frequently presents scenario-based questions asking you to choose the right service.

| Service | Role | Key Concepts | | --- | --- | --- | | Kinesis Data Streams (KDS) | Capture and process streaming data in real time | Shards, real-time processing | | Kinesis Data Firehose | Automatically deliver streaming data to data stores | Auto-scaling, data loading | | Kinesis Data Analytics | Analyze stream data using SQL or Flink | Real-time analytics | | Kinesis Video Streams | Capture and process video streams | Video analytics |

Key exam points to remember:

You need to implement custom stream processing logic → Kinesis Data Streams You need to automatically load data into S3 or Redshift → Firehose You need to analyze stream data with SQL → Data Analytics

!The 4 Kinesis services compared

Kinesis Data Streams Core Architecture (Shards)

The fundamental concept in Kinesis Data Streams is the Shard.

All throughput and scalability are determined by shards.

| Aspect | Detail | | --- | --- | | Shard write throughput | 1 MB/sec or 1,000 records/sec | | Shard read throughput | 2 MB/sec | | Max record size | 1 MB | | Default retention | 24 hours | | Max retention | 7 days | | Ordering guarantee | Within a shard |

The key exam point here is how to calculate throughput.

If throughput is insufficient, you can scale by adding more shards.

---

Service Modes: Provisioned vs On-Demand

Kinesis Data Streams offers two operating modes.

| Mode | Characteristics | Use Case | | --- | --- | --- | | Provisioned | You specify the shard count manually | Predictable traffic | | On-Demand | Capacity scales automatically | Unpredictable traffic |

On the exam, the correct answer is typically determined by the traffic pattern.

Traffic is steady → Provisioned Traffic fluctuates significantly → On-Demand

---

Producers: How Data Gets In

The entities that send data to a stream are called Producers.

There are three main approaches.

| Tool | Characteristics | Use Case | | --- | --- | --- | | AWS SDK | Direct API calls | Basic data transmission | | Kinesis Producer Library (KPL) | High-throughput transmission | Large-scale streaming | | Kinesis Agent | Automatic log file transmission | Server log collection |

Two scenarios appear frequently on the exam:

Large-scale streaming data processing → KPL Automatic log collection from Linux servers → Kinesis Agent

---

Consumers: How Data Gets Out

Applications that read stream data are called Consumers.

The basic approach uses the GetRecords API.

| Aspect | Detail | | --- | --- | | Call limit | 5 calls/sec | | Max read size | 10 MB | | Throughput | 2 MB/s per shard |

When multiple consumers read data simultaneously, they share the throughput.

To solve this problem, you can use the Enhanced Fan-Out feature.

With Enhanced Fan-Out:

Each consumer gets dedicated throughput (2 MB/s) Multiple applications can read the stream simultaneously without performance degradation

---

KCL (Kinesis Client Library)

When multiple application instances need to process a stream, KCL is used.

The main roles of KCL are as follows.

| Feature | Description | | --- | --- | | Distributed processing | Multiple workers process in parallel | | Checkpoint management | Saves processing position | | Failure recovery | Resumes from last checkpoint |

Checkpoint information is stored in Amazon DynamoDB.

The key exam point here is the relationship with cost.

Frequent checkpointing increases DynamoDB costs Infrequent checkpointing means more data needs to be reprocessed after a failure

---

Lambda and Kinesis Integration

Kinesis is very commonly used together with AWS Lambda.

Lambda automatically polls the stream and processes new data.

| Feature | Description | | --- | --- | | Auto polling | Lambda reads data from the stream | | Auto scaling | Lambda scales based on shard count | | Parallel processing | Uses Parallelization Factor |

The two main ways to increase throughput are:

Increase the Parallelization Factor Enable Enhanced Fan-Out

---

Shard Splitting and Merging

When traffic increases or decreases, you can adjust shards accordingly.

| Operation | Description | | --- | --- | | Splitting | Increases shard count | | Merging | Decreases shard count |

After splitting a shard, you must pay attention to data processing order.

Back to blog list