TechOps Examples
Hey — It's Govardhana MK 👋
Welcome to another technical edition.
Every Tuesday – You’ll receive a free edition with a byte-size use case, remote job opportunities, top news, tools, and articles.
Every Thursday and Saturday – You’ll receive a special edition with a deep dive use case, remote job opportunities and articles.
How much of your Kubernetes spend is going to waste? Enter a few details about your cloud environment in Zesty’s Cloud Cost Calculator to find out.
In one click, estimate:
→ Your cloud cost efficiency score.
→ How much of your spend may be unnecessary.
→ Where you could reduce costs.
IN TODAY'S EDITION
🧠 Use Case
How Amazon Aurora Handles High Availability and Scaling
👀 Remote Jobs
Flashbots is hiring a Senior DevOps Engineer
Remote Location: Worldwide
1001 is hiring a DevOps Engineer
Remote Location: Worldwide
📚 Resources
If you’re not a subscriber, here’s what you missed last week.
To receive all the full articles and support TechOps Examples, consider subscribing:
🛠 TOOL OF THE DAY
k8sgpt - A tool for scanning your Kubernetes clusters, diagnosing, and triaging issues in simple English.
🧠 USE CASE
How Amazon Aurora Handles High Availability and Scaling
A traditional relational database, whether self managed MySQL or Postgres on EC2, ties compute and storage together on the same instance. Losing that instance means losing both at once, and scaling read capacity usually means standing up read replicas that each maintain their own full copy of the data, replicated over the network from the primary. This model works, but it has a ceiling: storage grows only as fast as you can resize a single volume, replica lag grows with distance from the primary, and a failover means waiting for a new instance to come up and replay logs before it can serve traffic. Amazon Aurora was built specifically to break the coupling between compute and storage, and that single architectural decision is what makes its availability and scaling behavior fundamentally different from a standard managed RDS instance.
Amazon Aurora Architecture
The core idea behind Aurora starts with separating the database engine, which processes queries, from the storage layer, which persists data. In Aurora, that storage layer is not a single disk attached to one instance. It is a distributed, self healing volume shared across every instance in the cluster.

Inside an AWS Region, an Aurora cluster spans three Availability Zones, AZ 1, AZ 2, and AZ 3. AZ 1 holds the Writer (W) instance alongside one Reader (R) instance, while AZ 2 and AZ 3 each hold two Reader instances. All of these instances, writer and readers alike, connect to the same Shared Storage Volume, shown as one continuous band running beneath all three AZs. That volume is itself broken into small chunks and distributed across Storage nodes in each AZ, with each chunk replicated six ways, two copies per AZ, across the three zones. The entire storage layer also continuously backs up to Amazon S3.
The six-way replication across three AZs, two copies per zone, is also what gives Aurora its specific durability guarantee. Aurora storage is designed to tolerate the loss of an entire AZ plus one additional copy without losing data, and to tolerate the loss of an entire AZ without affecting write availability. A write is only acknowledged back to the client once a quorum of storage nodes, four out of six copies, confirms the write, which is why storage remains consistent and available even during a single AZ outage, without the writer instance itself needing to do anything special to handle that failure.
Amazon Aurora DB Cluster
Applications connecting to an Aurora cluster do not connect to individual instance IP addresses directly in a well architected setup. They connect through endpoints that abstract away exactly which physical instance is currently serving the request.

A client connects to two separate endpoints. Writes go to the Writer Endpoint, which always points to the master instance, shown flowing down to a single W database that writes into the Shared Storage Volume. Reads go to the Reader Endpoint, which performs Load Balancing across four separate reader instances, each marked R, and each reading from that same shared storage volume. The reader fleet itself sits under Auto Scaling, expanding or shrinking the number of reader instances based on load, and the underlying Shared Storage Volume grows through Auto Resize as data volume increases, all the way up to 128 TiB, without any manual intervention or downtime.
This split between a Writer Endpoint and a Reader Endpoint is what makes read scaling in Aurora operationally simple compared to manually managing individual replica connection strings in application code. An application team pointing all SELECT traffic at the Reader Endpoint gets automatic load distribution across however many reader instances currently exist in the cluster, and when Auto Scaling adds a fifth or sixth reader in response to increased read load, the application needs zero configuration changes, since the Reader Endpoint's DNS resolution simply starts including the new instance in its rotation.
The failover path benefits from this same architecture directly. If the writer instance fails, Aurora does not need to build a new writer from scratch or replay a backlog of replication logs. It promotes one of the existing reader instances, already attached to the identical shared storage volume, to become the new writer, and updates the Writer Endpoint's DNS to point at that promoted instance. This is why Aurora failover typically completes in under 30 seconds, versus the multi minute failover windows common with traditional replicated database setups where a standby has to catch up on replication lag before it can safely take writes.
Happy to bring you the most efficient Cloud Cost Calculator. In one click, estimate: Your cloud cost efficiency score, how much of your spend may be unnecessary, Where you could reduce costs.
🔴 Get my DevOps & Kubernetes ebooks! (free for Premium Club and Personal Tier newsletter subscribers)
Looking to promote your company, product, service, or event to 50,000+ DevOps and Cloud Professionals? Let's work together.



