At times we can be quick to bat off edge cases, this will only affect 1% of requests, or this has a one in a million chance of occuring! It’ll be fine!

Something often left under-considered is the relationship of this with scale.
That is to say, as you scale, and process more requests or messages or serve more pages or whatever it is you are doing more of - you will run into edge cases more and more often.

So one of the things that you need to do to make your program scale is not just make it performant and horizontally scalable and all of that jazz. But it is to also lift the reliability, and the percentiles you are thinking at.

Scale and percentiles

If you have only got 1,000 requests for something in a given minute, and you measure your p99, you are saying you’re okay for 10 requests to be kind of slow. Which might be okay, 10 isn’t a huge number.
But now you are servering 1,000,000 requests per minute, suddenly that’s 10,000 slow requests! in any given minute. 10,000 customers, or other systems trying to do something that are struggling to do this.
And on top of that, think about the amplification of problems you start to get when you are in a distributed system. The whole thing goes crazy.

So the remedy if that you must lift the bar, you add more 9’s. p99.9, p99.99 etc.
And the same for availability and reliability. You must lift your reliability targets as you scale, else you risk more and more people running into problems.
Of course, in a distributed system many of these problems are transient and might go away with a retry. But these retries must not be forgotten.

Scale and edge-cases

An Edge-case is something that we consider to be at the bounds of expected things to happen. We can picture it happening, and understand that it will eventually happen. But we tell ourselves it is too complex to deal with, and that it will never happen.

Scale again amplifies this. If you have a one in a million problem, and you a processing 1,000,000 items every minute. Well you will have a one in a million problem every single minute. Or 1,440 one in a million problems every single day.
I don’t know about you, but that’s more problems than I’d like.

In my experience, these edge cases have a great way of finding their ways to the people you’d least like them to. People like:

  • a journalist
  • your boss
  • your CEO
  • a really important strategic customer who is trialing your app and trying to decide if they want to pay your company lots of money

That all being said, we talk ourselves out of dealing with these for a reason.
If we caitered for every edge case imaginable, our program would likely be infinitely long, and impossibly difficult to operate.

Finding the balance

The technique I have generally seen is that you will push in one area for a while, until you start to drown a little in ops alerts and tech-debt and support tickets. Followed by a big tidy up and scaling project, potentially with a rewrite.

This approach is of course quite short-sited and deals with the problem as it comes up.
A better approach is that you need Engineering Leadership in your organisation which understands this problem, and the importance of continuous investment into these problems even while things look good.
Because if your business is growing, and you are scaling, holding the line requires continuous effort and investment.