AWS Analytics Services

Athena, Kinesis, Glue, QuickSight, EMR, OpenSearch Service, MSK, Data Exchange and Redshift - which analytics tool fits which data job.

What is it?

Analytics is turning raw data into answers. AWS offers a managed service for each stage: collecting and streaming data, storing and transforming it, querying it, and visualising it. Most services here are managed, so you do not run the servers.

  • Amazon Athena: run standard SQL directly on files in S3 with no servers to manage; you pay per amount of data scanned.
  • Amazon Kinesis: collect and process real-time streaming data. Data Streams ingests and stores streams, Data Firehose delivers streams to S3, Redshift or OpenSearch, and Managed Service for Apache Flink analyses streams with SQL or code.
  • Amazon MSK (Managed Streaming for Apache Kafka): a managed Kafka cluster for teams already using Kafka.
  • AWS Glue: serverless ETL (extract, transform, load) with a Data Catalog that stores table definitions so Athena, Redshift and EMR can find your data.
  • Amazon EMR: managed big data clusters running Apache Spark, Hadoop, Hive and Presto for large batch processing.
  • Amazon Redshift: a managed data warehouse for fast SQL analytics over large structured data, using columnar storage.
  • Amazon OpenSearch Service: search, log analytics and dashboards on top of OpenSearch.
  • Amazon QuickSight: business intelligence; build interactive dashboards and charts, with ML insights and pay-per-session pricing for readers.
  • AWS Data Exchange: find, subscribe to and use third-party datasets in the cloud.

A common pattern is a data lake: raw data lands in S3, Glue catalogs and transforms it, Athena or Redshift queries it, and QuickSight visualises the results.

Explain like I'm 10

Think of a city water system. Kinesis is the river carrying a constant flow, S3 is the reservoir, Glue is the treatment plant labelling and cleaning the water, Athena is a tap you open to take a glass now, Redshift is a bottled-water warehouse built for fast big orders, and QuickSight is the dashboard in the control room showing levels.

Examples

Query CSV files in S3 with Athena

-- Describe the files once (a table over an S3 prefix)
CREATE EXTERNAL TABLE sales (
  order_id string,
  country  string,
  amount   double
)
ROW FORMAT DELIMITED FIELDS TERMINATED BY ','
LOCATION 's3://my-analytics-bucket/sales/';

-- Then ask a question with plain SQL
SELECT country, SUM(amount) AS revenue
FROM sales
GROUP BY country
ORDER BY revenue DESC
LIMIT 5;

No cluster to launch: Athena reads the files in place and bills for the bytes scanned, so compressing or partitioning data lowers the cost.

Match the need to the service

Need                                   Service
-------------------------------------  ---------------------
SQL on files in S3, no servers         Athena
Real-time clickstream ingestion        Kinesis Data Streams
Load a stream into S3 automatically    Kinesis Data Firehose
Existing Kafka applications            MSK
Serverless ETL + data catalog          Glue
Spark/Hadoop on big datasets           EMR
Data warehouse for BI queries          Redshift
Log search and dashboards              OpenSearch Service
Business dashboards                    QuickSight
Buy third-party datasets               Data Exchange

How it works

Analytics systems separate storage from compute where they can. Data sits cheaply in S3; compute services (Athena, EMR, Redshift Spectrum) read it when needed. Streaming services buffer records in shards or partitions so consumers can read them in order and at their own pace.

Glue crawlers scan data, infer its schema and write table metadata to the Data Catalog. Redshift stores data by column and distributes it across nodes, which is why aggregations over billions of rows are fast. QuickSight connects to these sources (and SPICE, its in-memory engine) to render dashboards.

  sources        ingest            store & prepare      analyse         visualise
 +--------+   +-----------+     +-------------+   +-----------+   +-----------+
 | apps / |-->| Kinesis / |---->| S3 data lake|-->| Athena    |-->| QuickSight|
 | logs / |   | MSK /     |     |  + Glue ETL |   | Redshift  |   +-----------+
 | IoT    |   | Firehose  |     |  + Catalog  |   | EMR       |
 +--------+   +-----------+     +-------------+   | OpenSearch|
                                                  +-----------+

Why does it exist?

Companies generate far more data than a single database can usefully hold or query. Building and operating a Hadoop cluster, a Kafka cluster and a warehouse by hand needs specialist staff. Managed analytics services let a small team ask questions of large data without that operating burden.

When to use it

Use Athena for occasional SQL over S3 data, Redshift for heavy repeated BI workloads, Kinesis or MSK for real-time ingestion, Glue for ETL and cataloging, EMR for Spark/Hadoop jobs, OpenSearch for search and log exploration, QuickSight for dashboards, and Data Exchange when you need an external dataset.

When not to use it

Do not use Athena or Redshift as the transactional database behind an application (use RDS or DynamoDB). Do not use EMR for tiny jobs that Lambda or Glue could do, and do not stand up Redshift for a one-off query that Athena can answer.

Common mistakes

  • Using a data warehouse as an application database.

  • Storing Athena input as large uncompressed CSV, which makes each query scan (and cost) more than needed.

  • Confusing Kinesis Data Streams (ingest and hold records) with Firehose (deliver them to a destination).

  • Forgetting QuickSight is for visualisation, not for storing or transforming data.

  • Picking EMR when a serverless option would remove cluster management.

Practice exercises

  1. Easy:

    Write a one-line purpose for each of the nine services in this lesson without looking.

  2. Medium:

    Sketch a data pipeline for website clickstream data that ends in a daily dashboard, naming each AWS service in order.

  3. Medium:

    Explain why partitioning or compressing files lowers Athena cost.

  4. Hard:

    A team already runs Kafka producers and consumers on-premises. Compare moving them to MSK versus rewriting for Kinesis.

Interview questions

Exam-style: Which service lets you run SQL queries directly on data in S3 without provisioning servers?

Amazon Athena.

Exam-style: Which service is used to build business intelligence dashboards?

Amazon QuickSight.

Exam-style: Which service is a managed data warehouse?

Amazon Redshift.

Exam-style: Which service performs serverless ETL and maintains a data catalog?

AWS Glue.

Exam-style: Which service ingests and processes real-time streaming data?

Amazon Kinesis (Data Streams, Firehose, Managed Service for Apache Flink).

Exam-style: Which service lets you find and subscribe to third-party datasets?

AWS Data Exchange.

Exam-style: Which service runs managed Apache Spark and Hadoop clusters?

Amazon EMR.

Exam-style: A company already uses Apache Kafka and wants a managed service. Which one?

Amazon MSK.