Showing posts with label cloud. Show all posts
Showing posts with label cloud. Show all posts

Tuesday, October 30, 2012

How we got through the October AWS Service Degradation

After planning and practicing for Amazon Web Services outages at BuildLinks, we got to exercise our plans when AWS had a east coast service degradation this month.

The summary - AWS had a single zone failure in their US East region.  We initiated our plan to evacuate the failed zone and we were able to continue to deliver services to our clients from the non-affected AWS zones in the US East.  I am very proud of our infrastructure team and how they jumped on this and made sure we continued to deliver service.

We build, plan and practice for this stuff, it doesn't just magically happen. We just went through a drill for this very scenario a week before the AWS service degradation.  Also, as I have mentioned in a previous blog post, if you are too attached to an EC2 instance, you are probably doing something wrong.

The thing about product infrastructure and security is this:  no one really appreciates it until something goes wrong.  As a CTO, it's part of my job to make sure we can deliver product to our customers who run their businesses on our platform.  I have to champion for this all the time - "Yes, It is worth our time and effort to build a robust and secure platform, not just add features to the product".

We have some principles that govern how we deliver services on AWS, which helped us in this last service degradation.

  • Front end server EC2 instances are stateless.  If one goes down, there is another to take its place and customers don't get affected.
  • Front end server EC2 instances are templates.   They don't store configurations and application code. When a front end server starts up, it is told what it is, and it finds it's configuration and application code in a pre-configured S3 bucket, it then copies these locally and starts up.  This allows us to run many of these as we need, in multiple availability zones, and bring them up and down without issues.  We can also modify the config and code in once place and then have the front end servers can pick up the new versions easily. 
  • Front end services run in multiple availability zones fronted by Elastic Load Balancers (ELBs).  If a zone goes down, the ELB moves traffic to non-affected zones.  If ELBs have trouble, we can by-pass traffic to ELBs and shard traffic directly to the instances by reconfiguring DNS (we had do do this for the last service degradation).
  • Back end databases are state-full (of course),  but they are actively replicated to multiple availability zones. When one zone goes offline, we failover to the non-affected zone (again, we had do do this for the last service degradation).  For paranoia, we backup the databases to RackSpace. 

There is a lot of more stuff we do, but these principles helped us withstand these last set of issues.  We will continue to refine and work on product delivery.  For example, we are interested in having our services replicated across regions (East to West coast, for example).  There are tools to do this or we can roll our own.  We have toyed with this, but need to dedicate more time to it.

Security and service delivery is a process, not a destination.  We learned a lot and continue to refine and get better, but the moral is you have to think about this and dedicate time to this - and never really feel that you are "done".


Wednesday, June 27, 2012

Cloud-based Infrastructure: If you want to be available, be transient

So, there was outage a couple of weeks ago in the Amazon Web Services east coast region - a power problem affecting one of their zones. They have four (or now five) zones in the east coast, and apparently the others were unaffected.

What was amazing to me was not that there was an outage, but that so many users on the AWS forums had messages which essentially said "Help! I can't connect to my instance!".  Well, yes, the zone was offline for a few hours so instances running in that zone were offline.  But, if you rely so much on a single instance being available,  trouble is coming your way.

Cloud-based infrastructure needs to be thought of as transient - that is, like saying, "I like you but I am not committed to you".  If an instance or service is down, you should be able to shrug your shoulders and move on.

At BuildLinks, we are very heavy users of AWS.  We don't even own a single piece of hardware (except for laptops, a couple of hubs and wireless routers).  We are all "in" on AWS, but we continue to work hard and design our systems such that we are not committed to a particular zone or instance (heck, we even back up our databases to RackSpace outside of AWS, just to be paranoid).  If an instance goes offline, there is another to take its place.  AWS provides some services for doing this (like ELB), and some you have to build yourself.  For example, we use multi-zone web server instances fronted by ELB,  but we had to string together database replication ourselves -  our database is actively replicated to another zone in the east coast and also to the AWS west coast region. We back up our database to S3 and RackSpace every few minutes too.

Yes, we could experience an outage if an entire region goes offline, ELB fails or something of the like, but we would not be completely wiped out, given the replication and redundancy we took time to design and build, and continue to build and add with every product release. Cloud-based infrastructure reliability needs to be an on-going issue.

If you want to use cloud-based infrastructure, you need to take time to design and iterate your services for robustness - not just throw instances up with the assumption that they will be there forever.  It will also keep your blood pressure down and help you sleep. :)

If you have questions, drop me a line.  I am glad to share.

Monday, April 25, 2011

No need to Panic: Why we are sticking with Amazon Web Services.


Yes, the outage was very bad for Amazon Web Services.  Essentially, it was a multi-zone failure in the US East region for AWS that was not supposed to happen.  Amazon claimed that each zone in a region was independent from other zones, and that failures in one zone would not affect another.  The failure over last weekend showed that did not happen.  Failures spanned availability zones. That was bad.

Luckily, none of our production systems were affected.  Don’t know why, luck of the draw. Our stage systems went offline, and eventually we got tired of waiting for them to recover, so we copied the EBS volumes, and spun up new systems with the data of the old systems and everything worked fine.

Am I happy with the failure? No.  Especially since EC2 makes it impossible to migrate running instances to other zones or regions for that matter.  Also, the lack of transparency on how these availability zones work is very disturbing.  It’s hard to take their word for it until they prove how they work and why this won’t happen again.

Are we sticking with Amazon Web Services? Yes, but with some changes to our strategy going forward.  Cloud computing models are still the cheapest way to bring up services, platforms and infrastructure without a lot of upfront costs.  Our company could not be where it is now without AWS.  I figure that this failure will make Amazon Web Services better, and increase transparency.  This is not unlike some of the massive Internet service failures that were very common in the late 90’s, but as technology and platforms matured, and as we understood how to engineer better and stronger services, these outages have become very rarer.  Web sites and web-based services are more robust, and handle large volumes of traffic much better than they did 10 years ago.  Similarly, cloud computing will only get better and better.  The cost efficiencies alone are too great for it to go away.

What changes will we make to how we use Amazon Web Services?


  1. We will be duplicating our services to other AWS regions (US West and US East),  not relying on availability zones within a region to take care of us.
  2. We will be moving our system backups not only to S3, but now we will be moving them to an AWS competitor as well (RackSpace).  
  3. I am going to run our team through a fire-drill.  How do we recover our services in another AWS region in case of a disaster?  How do we recover our services somewhere else completely? 


At the end of the day, for our company to leave AWS would require a large capital commitment to hardware and infrastructure we are not prepared or really able to take on. Besides, I don’t want to manage another large set of IT infrastructure anymore.  It sucks to do that.  I got spoiled by letting AWS do that for me,  and frankly I like it that way.  With some refinement to the way we do things, I am just happy to keep paying AWS to do that for me, and that allows me to concentrate on delivering a cool and solid application.

Friday, February 18, 2011

To Infinity and Beyond!

Man, I really dig cloud-based computing infrastructure.  It was made my life as CTO of a SaaS lots easier.  The ability to grow new infrastructure with out having to order equipment, service and wait for install has made us more agile.

We are currently using Amazon's EC2.  One of the neatest things for me is how this really enables smooth upgrades of our production software.  When we are ready to do a release, we grow new "candidate" systems - front end serves and databases.  The candidate systems are populated with the new code and "smoke" tested by QA and development.  When cut over time comes around,  we simply shutdown previous version front end servers, move data over, move the dynamic IP addresses to point to the new systems and we are off and running.   We keep around the old systems just in case,  and decommission them about a week later.

Very cool.