AWS Database Services

RDS, Aurora, DynamoDB, Redshift and the specialised and migration database services.

What is it?

AWS offers purpose-built databases rather than one-size-fits-all.

  • Amazon RDS: managed relational databases (MySQL, PostgreSQL, MariaDB, Oracle, SQL Server, Db2). AWS handles backups, patching, and failover (Multi-AZ); you manage schema and queries.
  • Amazon Aurora: AWS's cloud-built relational engine compatible with MySQL and PostgreSQL, with storage replicated across AZs and fast failover.
  • Amazon DynamoDB: serverless key-value and document NoSQL database with single-digit-millisecond performance at nearly any scale.
  • Amazon Redshift: data warehouse for analytic SQL over large datasets.
  • Amazon ElastiCache: managed in-memory cache (Redis-compatible and Memcached) to speed up reads. DAX is an in-memory cache specifically for DynamoDB.
  • Amazon MemoryDB: a durable, Redis-compatible in-memory database for use as a primary store.
  • Amazon DocumentDB: MongoDB-compatible document database.
  • Amazon Neptune: graph database for highly connected data.
  • Amazon Keyspaces: managed Cassandra-compatible database.
  • Amazon Timestream: time-series data.

Migration helpers: AWS Database Migration Service (DMS) moves data between databases, often with little downtime, and AWS Schema Conversion Tool (SCT) helps convert schemas and code when changing engines (for example Oracle to PostgreSQL).

Note: Amazon QLDB (ledger database) was announced for end of support, so treat it as retired and look at alternatives when you meet it in older material. Service lineups change - verify in current docs.

EC2-hosted vs managed databases: you can install a database engine on an EC2 instance, but then you handle patching, backups, high availability, replication and scaling yourself. A managed service such as RDS, Aurora or DynamoDB automates those tasks. Self-hosting is chosen only for unsupported engines, full OS access or special licensing.

In-memory caching: Amazon ElastiCache (Redis OSS/Valkey and Memcached) caches hot data to cut latency and database load; Amazon MemoryDB is a durable Redis-compatible database that can be the primary store; DAX caches DynamoDB reads. AWS Database Migration Service moves the data and the Schema Conversion Tool / DMS Schema Conversion converts schema and code when engines differ, for example Oracle to Aurora PostgreSQL.

Explain like I'm 10

Choosing a database is like choosing storage furniture. A relational database is a filing cabinet with labeled, linked folders and strict forms. DynamoDB is a wall of numbered lockers - blazing fast to open by number, not for complicated questions. Redshift is a library's reading room set up for analysts to compare thousands of books at once. A cache is the sticky note on your monitor.

Examples

Create a DynamoDB table and use it

aws dynamodb create-table --table-name Orders \
  --attribute-definitions AttributeName=customerId,AttributeType=S AttributeName=orderId,AttributeType=S \
  --key-schema AttributeName=customerId,KeyType=HASH AttributeName=orderId,KeyType=RANGE \
  --billing-mode PAY_PER_REQUEST

aws dynamodb put-item --table-name Orders \
  --item '{"customerId":{"S":"c-1"},"orderId":{"S":"o-100"},"total":{"N":"42.5"}}'

aws dynamodb query --table-name Orders \
  --key-condition-expression "customerId = :c" \
  --expression-attribute-values '{":c":{"S":"c-1"}}'

Access patterns drive the key design: query by partition key (customer), sort by order id.

Which database?

Need                                               Choose
-------------------------------------------------  ---------------------
Familiar SQL, existing engine, managed backups     RDS
High-performance MySQL/PostgreSQL, fast failover   Aurora
Huge scale key-value, serverless, ms latency       DynamoDB
Analytics across billions of rows                  Redshift
Speed up repeated reads                            ElastiCache / DAX
Social graph, fraud rings                          Neptune
Move Oracle to PostgreSQL                          SCT + DMS

How it works

RDS runs a database engine on managed instances; Multi-AZ keeps a standby copy in another AZ and fails over automatically, while read replicas scale reads. Aurora separates compute from a shared, replicated storage layer that grows automatically. DynamoDB partitions data by key across many servers and replicates it across AZs; you pay per request or provision capacity.

Redshift stores data in columns and spreads queries across nodes, so aggregate analytics are fast. Caches sit in front of a database and answer repeat reads from memory.

  App --> ElastiCache / DAX (hot reads)
     |
     +--> RDS / Aurora   (relational, SQL, transactions)
     +--> DynamoDB       (key-value, scale-out)
     '--> Redshift       (analytics, columns)

  On-prem DB --[SCT convert schema]--> [DMS replicate] --> target DB

Why does it exist?

Running databases yourself means patching, backups, failover, and capacity planning. Managed databases remove that, and specialised engines fit different data shapes far better than forcing everything into one model.

When to use it

Pick the engine by data shape and access pattern: relational for joins and transactions, DynamoDB for predictable key-based access at scale, Redshift for analytics, caches for read acceleration.

When not to use it

Do not use DynamoDB when you need ad-hoc relational queries over many fields. Do not use RDS for petabyte-scale analytics. Do not run your own database on EC2 unless you need control a managed service cannot give.

Common mistakes

  • Using a relational database for everything out of habit, or NoSQL without defining access patterns first.

  • Believing Multi-AZ improves read throughput (read replicas do; Multi-AZ is for availability).

  • Skipping backup and retention settings review.

  • Leaving a database publicly accessible.

  • Treating DMS as a schema converter (that is SCT's job).

  • Installing a database on EC2 'to save money' and forgetting the staff time for backups, patching and failover.

Practice exercises

  1. Easy:

    Match each to a database: shopping-cart sessions with millisecond reads, a monthly sales report across years of data, a friend-of-friend recommendation.

  2. Medium:

    Explain Multi-AZ vs read replica and when you would use each.

  3. Medium:

    Design a DynamoDB key schema for 'get all orders for a customer, newest first'.

  4. Hard:

    Plan a migration from on-premises Oracle to Aurora PostgreSQL using SCT and DMS, including how to minimise cutover downtime.

Interview questions

RDS vs DynamoDB?

RDS is managed relational SQL with joins and transactions; DynamoDB is a serverless key-value/document store that scales horizontally for known access patterns.

What is DMS?

A service that migrates data between databases, including continuous replication to reduce downtime.

What does ElastiCache provide?

A managed in-memory cache (Redis-compatible/Memcached) to cut read latency and database load.

Exam-style: Which service is a petabyte-scale data warehouse?

Amazon Redshift.

Exam-style: Which service converts a database schema when migrating between engines?

AWS Schema Conversion Tool (SCT).

Exam-style: Which service reduces read latency with an in-memory cache in front of a database?

Amazon ElastiCache.

Exam-style: What is the main advantage of RDS over a database on EC2?

AWS automates patching, backups, replication and failover, reducing operational overhead.