As a member of the IBM Cloud IKS and ROKS offering management team, there are things that customers get tripped up on when they jump into the world of a cloud managed Kubernetes/OpenShift cluster. The following are 5 things I wish everyone would consider before they create their first cluster.
Your cluster size will change
As much as you would like to get your cluster size right, you won’t. I have yet to meet a customer that completely understands their workload size and throughput requirements before they create a cluster. After all, this is the cloud. You should expect to be able to grow and shrink as needed. Take advantage of the cluster autoscaler if you anticipate dynamic throughput. Add and remove worker nodes as you deploy more apps and you learn their true resource requirements. Also, explore solutions like Turbonomic or StormForge to help rightsize your cluster.
Plan to be without one of your worker nodes
You will need to live without a worker node at some point in the future for various reasons. Some of those are bad reasons. Nodes fail. Availability zones fail. Build a cluster that can handle that. Other reasons are not that bad. Upgrading worker nodes requires a reload and you need to be able to lose a node during the reload. This is good, this is working as intended. Plan for it. In general planning for redundancy and resilience will also help you when workloads grow (see #1 above).
Kubernetes/OpenShift versions change every 4 to 5 months
Please please please plan for and realize that you need to upgrade your Kubernetes or OpenShift versions regularly. The version progression comes on fast and you must pay attention to this. Understand when versions go out of maintenance. Practice your upgrades. We make it very easy to upgrade (IBM does it for you when you request), but you must initiate the upgrade. As a cloud provider, we have no choice but to follow the release cycle of Kubernetes and OpenShift. When versions go out of support you will lose critical security patches and you will quickly be running your workloads in a risky environment.
Understand storage solutions (and other add-ons) and their dependencies on the platform
Storage is very important. And you must understand what upgrading your cluster does to your storage choice. If you go down the path of an SDS (we support ODF and Portworx) you must understand the upgrade path and process for your SDS as well. Again, practice practice practice. What happens to the attached storage when a node goes down? What are the node requirements for your SDS. All if this is important to learn and understand before you get to a critical situation with a critical workload. If you add other things to your cluster, make sure you understand how it supports a new version of Kube or OCP.
Read the release notes
Read the release notes for each new version to understand when APIs change and update your applications accordingly. Many customers upgrade their clusters and then find that their workloads fail. It usually comes down to some CRD API change. Again, practice makes perfect. Understand your application and its dependencies on the platform you are running on.
A bit of planning ahead is all that is needed. But a failure to understand any of these will lead you to a bad experience. And I don't want that.