Practice this topic in a realistic system design interview
Every database, virtual machine, message broker, and search engine needs somewhere to store its data. Many of these systems rely on block storage, a fundamental storage model designed for fast, efficient access to data.
In this chapter, we will look at what block storage is, how it works under the hood, and why databases, virtual machines, and other performance-sensitive systems depend on it.
Suppose we are running a PostgreSQL database for an e-commerce application. Every second, the database reads and updates thousands of small pieces of data. It changes one customer's balance, inserts one order row, and updates one index entry. Each of these operations touches a few kilobytes somewhere in the middle of a very large set of files.
Could we store those database files in object storage? That would not work well.
What the database needs is storage that behaves like a disk. It should allow fast reads and writes at any position, and it should let us update small pieces of data in place. That is what block storage provides.
Block storage stores data as a sequence of fixed-size blocks. Each block has a numeric address, and nothing else. There are no file names, no folders, and no metadata describing what the data means.
The storage device only understands two basic operations:
A common block size is 4 KB. So a 1 TB volume is simply a very long array of roughly 268 million blocks.
The hard drive or SSD in your laptop is block storage. In the cloud, services like Amazon EBS, Google Cloud Persistent Disk, and Azure Managed Disks provide block storage as virtual disks that you attach to servers.
If block storage has no concept of files, how does a database work with files? The answer is the file system.
When we attach a new block storage volume to a server, the operating system sees it as a raw disk. Before we can store files on it, we format it with a file system, such as ext4 or XFS on Linux, or NTFS on Windows. The file system keeps track of which blocks belong to which file.
Now let's follow a single write. Suppose PostgreSQL wants to update one 8 KB page inside the orders table file.
orders file.Nothing else on the volume is touched. This is the key strength of block storage. It supports small, random, in-place updates, which is the access pattern databases rely on.
Some databases go one step further. They skip parts of the file system's caching (for example with O_DIRECT) and manage their own memory buffers, because they know their access patterns better than the operating system does. Either way, the block device underneath sees the same thing: reads and writes to numbered blocks.
Block storage comes in two main forms. The first is local storage: a disk that is physically attached to the server, such as an NVMe SSD inside the machine.
Because the data does not cross a network, local storage offers very low latency, often well under a millisecond. It can also deliver very high IOPS and throughput.
But local storage has an important limitation: the data is tied to that physical machine. If the server fails, the data may be lost or stuck on that machine.
In the cloud, local disks are often ephemeral. For example, the data on an AWS instance store disk is lost when the instance is stopped or terminated, or when the underlying hardware fails. It only survives a normal reboot.
So local storage is a good fit for:
The second form is network-attached block storage. Here, the volume lives on a separate storage system, and the server accesses it over the network.
From the server's point of view, it still looks like a normal local disk. The operating system formats it, mounts it, and reads and writes blocks just as it would with a physical drive.
The big advantage is that the storage is independent of the server. If a server fails, we can start a new server, attach the same volume, and continue where we left off. We can also resize volumes, change their performance tier, and take snapshots without touching the server.
The trade-off is latency. Every read and write now involves a network hop, so network block storage is usually slower than a local NVMe drive. A local NVMe read can take tens of microseconds, while a network volume read typically takes from a few hundred microseconds to around a millisecond.
| Local (e.g. NVMe instance store) | Network-attached (e.g. EBS, SAN) | |
|---|---|---|
| Latency | Lowest, well under a millisecond | Higher, includes a network hop |
| Survives server failure | No | Yes, attach the volume to a new server |
| Resize or change tier | No, fixed with the machine | Yes, often while in use |
| Snapshots | Not built in | Built in |
| Good for | Caches, temp files, self-replicating systems | Databases, boot disks, most stateful servers |
When choosing block storage, three numbers matter most.
IOPS (input/output operations per second) measures how many individual reads and writes the volume can handle each second. Databases with many small random operations care a lot about IOPS.
Throughput measures how many megabytes per second the volume can transfer. Workloads that read large files sequentially, like analytics scans or backups, care more about throughput.
The two are connected:
This is why a volume can hit one limit before the other. A database doing small 4 KB reads may run out of IOPS long before it uses its full throughput, while a backup job reading in 1 MB chunks may hit the throughput limit with only a few hundred IOPS.
Latency is how long a single read or write takes to complete. Even a volume with high IOPS can feel slow to a database if each individual operation takes too long. A transaction that does five reads one after another waits for all five, so per-operation latency adds up quickly.
Latency also rises when a volume is pushed close to its limits. Requests start to queue, and each one waits longer. It is good practice to watch queue length and latency, not just IOPS.
Cloud providers usually let you choose these numbers. For example, an Amazon EBS gp3 volume starts with a baseline of 3,000 IOPS and 125 MB/s, and you can pay to provision more. Volume types designed for heavy databases, like io2, offer much higher IOPS and more consistent latency, at a higher price.
| Metric | What It Measures | Who Cares Most |
|---|---|---|
| IOPS | Operations per second | Databases with small random reads and writes |
| Throughput | MB per second | Analytics scans, backups, log processing |
| Latency | Time per operation | Transactional databases, anything user-facing |
So what happens behind a cloud block storage volume?
When we create a volume, the provider does not dedicate one physical disk to us. Instead, the volume is split into pieces and stored across many storage servers. Each piece is replicated to multiple servers, so a single disk or server failure does not lose data.
When our server writes a block, the request travels over the network to the storage servers responsible for that block. The write is stored on the replicas before it is acknowledged.
This replication usually happens within a single availability zone. Keeping replicas close together keeps write latency low. But it also means:
Replication inside a zone also does not make a volume impossible to lose. AWS, for example, designs gp3 volumes for 99.8% to 99.9% durability per year, and io2 volumes for 99.999%. That is why snapshots, and database replicas in other zones, still matter.
Another important property of block storage is that a volume is usually attached to one server at a time. Why?
Remember that the file system on the server keeps track of which blocks belong to which file. It also caches some of that information in memory.
Now imagine two servers attach the same volume, and both write to it with a normal file system like ext4. Each server assumes it is the only one making changes.
Some cloud services do offer multi-attach for certain volume types, such as EBS io2. But to use it safely, the servers need a cluster-aware file system (such as GFS2 or OCFS2) or an application that coordinates access to the blocks itself.
When many servers need to share the same files, file storage is usually the simpler choice.
Block storage volumes often support snapshots. A snapshot captures the state of the volume at a specific point in time.
In Amazon EBS, snapshots are incremental:
This makes regular snapshots much cheaper than copying the full volume every time. Each snapshot can still restore the complete volume, because it refers back to unchanged blocks stored by earlier snapshots.
EBS stores snapshots in Amazon S3, which makes them far more durable than the volume itself. Snapshots are useful in several ways:
One thing to keep in mind is consistency. If a database is actively writing during a snapshot, some of its data may still be sitting in memory, not yet on disk. The snapshot then captures the disk as if the server had suddenly lost power. For a clean backup, databases usually need to flush writes first, or rely on their own crash recovery (such as replaying the write-ahead log) when restoring.
A volume restored from a snapshot can also be slow at first. Its blocks are loaded from S3 the first time they are read, so the first access to each block takes longer until the volume is fully warmed up.
Block storage mostly scales vertically.
On many cloud platforms, both changes can be made while the volume stays attached and in use. After resizing, we still need to extend the file system so it can use the new space:
But a single volume has limits. Each volume type has a maximum size, IOPS, and throughput. And because one volume is typically attached to one server, a single volume cannot absorb unlimited traffic from many machines.
When a system outgrows one volume, the usual next step is to add more servers, each with its own volume. For a database, this often means read replicas or sharding, where each node manages its own block storage.
Another option on a single server is to combine several volumes into one larger device with software RAID 0 (striping). This adds up their IOPS and throughput, but losing any one volume loses the whole array, so it only makes sense when the data is protected some other way.
In a system design, block storage usually sits underneath components that manage their own data:
In all of these cases, the application handles the higher-level logic, like replication across zones and sharding across nodes. Block storage provides a fast, reliable disk underneath each individual server.
This split is worth stating clearly in a design discussion. A single EBS volume lives in one zone, so it does not make your database highly available on its own. For that, you still need a replica in another zone, which the database (or a managed service like Amazon RDS) handles.
Block storage is the right choice when a single server needs fast, low-latency access to data that changes frequently. Databases, transaction logs, virtual machine disks, and search indexes are classic examples.
Block storage is not a good fit when many servers need to share the same files, or when you need to store huge amounts of media, backups, or logs cheaply. In those cases, file storage or object storage usually works better.
| Need | Best Fit |
|---|---|
| Fast random reads and writes for one server (databases, VM disks) | Network-attached block storage |
| Lowest latency, and the data is replicated elsewhere | Local block storage (NVMe) |
| Many servers sharing the same files and folders | File storage |
| Huge amounts of media, backups, or logs at low cost | Object storage |
11 quizzes