How to Write a Test Plan for AI Infrastructure
A practical guide to AI infrastructure testing, covering acceptance criteria, FAT, SAT, UAT, performance validation, and risk management for AI and HPC deployments.
A practical guide to AI infrastructure testing, covering acceptance criteria, FAT, SAT, UAT, performance validation, and risk management for AI and HPC deployments.
Processes That Matter At Red Oak I have had the privilege of working with numerous organisations to assist with expanding their HPC footprint into the cloud. With an ever-evolving set of resource types available from cloud providers, technical implementation is rarely the issue. Instead, defining the practical processes and governance through which the HPC service … Read more
A Practical Guide Artificial Intelligence (AI) and machine learning (ML) are no longer niche; they are driving innovation across all industries. Behind the scenes, GPUs handle the heavy processing that enables these workloads. However, GPUs are expensive, sometimes scarce, and not always easy to manage at scale. This is where Kubernetes comes in. Originally developed … Read more
After a successful proof-of-concept (PoC) phase, the natural next key milestone is a pilot deployment. But have you ever wondered why many pilot projects stall before becoming operational? In this post, we’ll walk through the essential stages, questions, and best practices for running a production-grade Cloud HPC Pilot following a successful PoC. A pilot bridges … Read more
Introduction A common topic during many of our projects is the sizing of a HPC cluster. More specifically, we often come across clusters which we can loosely characterise as having a low average utilisation. This may be an intentional design point or an artifact of over procurement or oversizing. Reasons why an oversized cluster may … Read more
Azure Batch, or Azure Kubernetes Service (AKS) Choosing the right platform for computational workloads is a business decision that impacts scalability, cost, and operational efficiency. Azure offers three powerful options for different use cases: Azure CycleCloud, Azure Batch, and Azure Kubernetes Service (AKS). Each tool has distinct strengths and suitability for specific types of workloads. This … Read more
Having assisted with the management of a number of cloud HPC systems over the last few years, I have noticed several cost components that are often overlooked. On the other hand, customers may have some difficulty keeping control of the more obvious costs. In this article, with a focus on cloud HPC deployments, I hope … Read more
Customising Your Cluster CycleCloud is a powerful tool for orchestrating and managing clusters of Virtual Machines (VMs) on Azure. Whether you are planning your own infrastructure or using the new Azure CycleCloud Workspace for Slurm, there are multiple ways to deploy and configure CycleCloud. After deploying your CycleCloud host and creating a cluster using one of … Read more
Introduction Cloud HPC can help speed up the procurement process and provide flexibility that on-premises HPC cannot offer. However, this does not mean that taking time to design, iterate, and engage with key members of your organisation can be skirted around. One key area that can often be overlooked is monitoring. Monitoring can provide a … Read more
Streamlining Your Cloud HPC Journey Over the last few years, I have had the privilege of being involved in several Cloud High-Performance Computer (HPC) cluster deployments. Despite the great variation in designs, which highlights the great flexibility on the Cloud, I have noticed some common roadblocks along the way. These roadblocks can in the best … Read more