Block, File and Object Storage

Instance store, EBS, EFS and S3: choosing a storage type, S3 classes, lifecycle rules, durability and availability.

What is it?

AWS offers three storage shapes:

  • Block storage (like a raw disk): Amazon EBS volumes attach to one EC2 instance in the same AZ (some types support multi-attach) and persist independently. Instance store is disk physically attached to the host: very fast, but data is lost when the instance stops or terminates. EBS snapshots are incremental backups stored in S3-backed storage; you can copy them across Regions.
  • File storage: Amazon EFS is a managed, elastic NFS file system for Linux that many instances can mount at once across AZs. Related options include FSx for Windows and other file systems.
  • Object storage: Amazon S3 stores objects (data + metadata) in buckets with unlimited scale, accessed over HTTP APIs. It is not a disk you mount like a drive.

S3 storage classes trade retrieval speed and access frequency for price:

  • S3 Standard: frequent access.
  • S3 Express One Zone: very low latency in a single AZ, for performance-critical data.
  • S3 Intelligent-Tiering: moves objects between tiers automatically when patterns are unknown.
  • S3 Standard-IA and One Zone-IA: infrequent access, lower storage price, retrieval fee.
  • S3 Glacier Instant Retrieval: archive with millisecond access.
  • S3 Glacier Flexible Retrieval: archive, retrieval in minutes to hours.
  • S3 Glacier Deep Archive: cheapest, retrieval within hours (long-term retention).

Lifecycle rules move or expire objects automatically. S3 has high durability (designed for eleven nines, 99.999999999%, meaning data is very unlikely to be lost), which is different from availability (how often you can read it right now), which varies by class.

More storage and data-protection services

  • Amazon FSx: fully managed third-party file systems: FSx for Windows File Server (SMB), FSx for Lustre (high-performance computing), FSx for NetApp ONTAP and FSx for OpenZFS.
  • AWS Storage Gateway: a hybrid service that connects on-premises apps to AWS storage using File, Volume or Tape Gateway types, with local caching.
  • AWS Backup: a central service to define backup policies and schedules across EC2, EBS, RDS, DynamoDB, EFS, FSx and more, with cross-Region and cross-account copy.
  • AWS Elastic Disaster Recovery (DRS): continuously replicates servers to a low-cost staging area in AWS and launches them quickly after a disaster.
  • AWS Transfer Family: managed SFTP, FTPS and FTP endpoints that read and write to S3 or EFS, so partners keep using file-transfer clients.

Explain like I'm 10

EBS is a personal hard drive on your desk - fast, yours, but one desk at a time. EFS is a shared network folder the whole office can open. S3 is a giant warehouse with labeled boxes: you hand over a box and get a ticket, you never edit a box in place. Glacier is the deep archive basement: cheap shelves, but fetching a box takes a while.

Examples

S3 lifecycle configuration

{
  "Rules": [
    {
      "ID": "age-out-logs",
      "Status": "Enabled",
      "Filter": { "Prefix": "logs/" },
      "Transitions": [
        { "Days": 30,  "StorageClass": "STANDARD_IA" },
        { "Days": 90,  "StorageClass": "GLACIER_IR" },
        { "Days": 365, "StorageClass": "DEEP_ARCHIVE" }
      ],
      "Expiration": { "Days": 2555 },
      "NoncurrentVersionExpiration": { "NoncurrentDays": 30 }
    }
  ]
}

Apply with aws s3api put-bucket-lifecycle-configuration --bucket NAME --lifecycle-configuration file://lifecycle.json. Check minimum storage durations and transition rules for each class first.

Everyday storage commands

# S3
aws s3 cp report.csv s3://zykit-demo-bucket-12345/reports/report.csv --storage-class STANDARD_IA
aws s3 sync ./site s3://zykit-demo-bucket-12345/site --delete
aws s3 ls s3://zykit-demo-bucket-12345/reports/

# EBS volume and snapshot
aws ec2 create-volume --size 20 --volume-type gp3 --availability-zone eu-west-1a
aws ec2 create-snapshot --volume-id vol-0abc1234 --description "pre-upgrade"

Which storage?

Need                                        Choose
------------------------------------------  ----------------------
Boot disk / database volume for one EC2     EBS
Scratch data, caches, can be lost           Instance store
Shared files for many Linux servers         EFS
Images, backups, data lake, static site     S3
Rarely read compliance archive              S3 Glacier Deep Archive

How it works

EBS volumes are network-attached and replicated within their AZ, so they survive an instance stop but not the loss of the AZ - snapshots protect against that. EFS replicates data across multiple AZs and grows and shrinks automatically. S3 replicates objects across multiple AZs (except One Zone classes) and offers versioning, replication to other Regions, encryption, and access controls.

Access to S3 is controlled by IAM policies, bucket policies, and Block Public Access settings. Objects can be fetched through presigned URLs that grant temporary access.

  EC2 instance
    |-- Instance store (local, ephemeral)
    |-- EBS volume (one AZ)  --snapshots--> S3-backed storage
    '-- EFS mount (shared, multi-AZ)

  Application / users --HTTP API--> S3 bucket (objects)
        lifecycle: Standard -> IA -> Glacier -> Deep Archive -> expire

Why does it exist?

Different data has different access patterns, performance needs, and value. One storage type for everything would be too slow, too expensive, or too fragile. Specialised storage types match cost to need.

When to use it

Use EBS for databases and boot volumes, EFS for shared Linux file systems, S3 for static content, backups, logs, and analytics data, lifecycle rules to cut cost as data ages, and Intelligent-Tiering when access is unpredictable.

When not to use it

Do not keep irreplaceable data on instance store. Do not use S3 as a low-latency block device for a database. Do not put hot data in deep archive classes - retrieval is slow and fees apply.

Common mistakes

  • Treating durability and availability as the same thing.

  • Leaving buckets public.

  • Forgetting EBS volumes and snapshots keep costing after instances are deleted.

  • Moving small objects to cold classes where minimum size or duration charges outweigh savings.

  • Using One Zone classes for data you cannot recreate.

  • Using Storage Gateway when a one-time bulk move is needed - that is DataSync or Snow Family territory.

Practice exercises

  1. Easy:

    Classify each as block, file, or object: EBS, EFS, S3, instance store.

  2. Medium:

    Write lifecycle rules for application logs kept 7 years with the cheapest cost and rare retrieval.

  3. Medium:

    Use the unit converter to estimate the monthly size of 1 TB written daily and what that means for storage class choice.

  4. Hard:

    Explain how you would restore a database volume after an AZ failure using snapshots, and what data could be lost.

Interview questions

Durability vs availability for S3?

Durability is the likelihood data is not lost (eleven nines); availability is the likelihood you can access it at a moment, which differs between classes.

EBS vs EFS?

EBS is block storage usually attached to a single instance in one AZ; EFS is a shared network file system accessible from many instances across AZs.

What happens to instance store data when the instance stops?

It is lost.

Exam-style: Which S3 class is cheapest for long-term archives retrieved rarely?

S3 Glacier Deep Archive.

Exam-style: Which feature automatically moves objects between storage classes over time?

S3 Lifecycle rules (or S3 Intelligent-Tiering for unpredictable access).

Exam-style: Which service gives on-premises applications access to cloud storage with local caching?

AWS Storage Gateway.

Which service centralises backup policies across many AWS services?

AWS Backup.

Which service provides managed SFTP into Amazon S3?

AWS Transfer Family.