Available 09:00-17:30
|

A career in HPC Part III

The Storage Era – From Clusters to Consultancy 

In Part II of my HPC career journey, I talked about how the small company I was working for had successfully moved into building affordable compute clusters, and how demand for our in-house designed systems grew rapidly. 

As the company evolved, so did the challenges, and once compute was becoming established, attention naturally began to turn to another equally important part of HPC infrastructure, storage. 

Time for a change 

After the heyday of the blade clusters, we had a look at a parallel storage product.

Following our work on blade-based clusters, this was a type of storage we thought there was a need for.  

At that time there were a few products out there offering parallel storage solutions, and many were quite expensive. 

We had a relationship with a Canadian company that had a software stack which could turn a group of storage servers, or “bricks” as they were known, into a performant parallel storage solution. 

This time, we bought in a 2U storage chassis that had a group of hot-pluggable disks in the front, and an Xeon-based system board at the back. 

To make a storage system performant, you need to understand what your requirements are. 

Do you need low latency with hundreds or thousands of transactions per second, or do you want high bandwidth performance because you have multi-gigabyte image files to move around? 

There is no magic bullet that fulfils all of these requirements unless you have a hybrid solution and an unlimited budget. 

We could provide a number of these storage bricks and form a parallel filesystem, but what we didn’t have was access to a large compute cluster that could throw masses of data at it so we could properly measure and tune the performance. 

So, we relied on providing demo systems to customers in order for them to try them out and give us some feedback on performance. 

We were targeting organisations where large,  parallel data transactions and high IOPS (input/output operations per second) were a critical requirement, so oil and gas, media organisations and our old friends bio-informatic companies. 

In many of the datacentres where I installed our demo systems, I would notice a stack of boxes from another well-known parallel storage company often in the corner. 

So we were competing with the big boys, and on many occasions, customers would want to entrust their data to a well-known parallel storage supplier. 

While this was going on, I had been thinking it may be time for a change of scenery, and I had a chat with a cluster solutions company in Warwick. 

They were a bigger operation than I was used to, and while I had a similar role, it was more focused on building, testing and maintaining their cluster products. 

All of the hardware was bought in, and they had a logistics department who racked, built and cabled all of the clusters, so all I had to do was install the HPC software stack and test everything, so all the fun but without the Meccano (google it). 

When I first arrived, the business used a subcontractor that they used that procured all the hardware, and it would then be built and tested at their premises, then shipped directly to the customer site. 

After a while, the company acquired a warehouse that we used as a build centre, which allowed us to procure, build and test all of the hardware ourselves. 

So, I was mostly based at this build centre, where the logistics team would put together a customer order, then I would install the cluster stack. 

The stack was written in BASH by one of our technical directors.

It was a simple but effective piece of software where we could install the head node, configure all of the compute nodes and have benchmarks running in a few hours. 

If the timing was right, we would get the cluster built towards the end of the week and leave it running the final benchmark, LINPACK , over the weekend. 

If all was well, the following Monday, the script would print out a test certificate with all of the benchmark results on it. 

If the benchmark results looked good, the system would then be torn down and packed by logistics, then reassembled on the customer’s site. 

I would then turn up, get it all powered on and rerun the benchmark test, but only for a few hours this time. 

I would print another test certificate, and hopefully it would have similar results to the build centre ones. 

As long as they were within a few percent of the build centre tests, we would present the results and the cluster to the customer. 

I think a lot of our customers were really impressed with this method of delivery. 

I really enjoyed this work. It was the best of both worlds; you got to play with some cool hardware, and you also got to visit the customer sites. 

I visited lots of different universities, aerospace companies, pharmaceuticals and some government sites. 

I did this work for five years, enjoying every minute. 

Towards the end, the company got into difficulties and was put into administration and subsequently closed down. 

I do not wish to pass any comment on the reasons why this happened, except to say I was bitterly, bitterly disappointed. 

Behind secure doors 

In the final weeks before the company closed, a colleague and I had begun a new project providing HPC consultancy for a government department under subcontract through Red Oak Consulting. 

We had only been in this role for a few weeks when our company went into administration, which left us in a bit of a quandary. We had just embarked on a new and exciting opportunity, but suddenly had no employer behind us. 

As it happened, Red Oak agreed to take us on as employees, and as we were required to hold a reasonably high security clearance, this had to be checked and cleared by the government agency in question before we could resume our roles on-site. 

Due to the nature of work this particular government department does, I am not able to go into the level of detail that I have done in my previous engagements. 

This was a unique, interesting and sometimes challenging role, and an experience that I will remember for the rest of my life. 

All I can really say is that in the six and a half years I was engaged there, I felt proud and privileged to work alongside the talented and dedicated staff that help keep this country safe. 

Although I was employed by Red Oak, I was officially engaged as a contractor, not a civil servant, carrying out work of a specific nature that the agency is not able to easily fill by using their own staff.

These types of engagement are not normally permanent, so eventually my work there was slowly coming to an end, and as the site was not easily commutable every day, I was living in a nearby flat during the week and only coming home at weekends.  

Naturally, this would put a strain on anyone’s personal life, so I started looking for something closer to home.

One day, I came across an opportunity that was not only closer to home, but it was literally a 10-minute commute, as it was in the town where I lived. 

So, once I was offered the role, although I was sad to leave behind all the new friends and colleagues I had made during my government work, and of course, my Red Oak colleagues, I was excited to take up my new position, partly because I could move back home. 

As you might expect, my government work was secure and very process-driven, so if you came across a new app or piece of code that solved a problem you were working on, you could not just download and install it.

It just does not work that way.

There were forms to fill in and authorisations to be requested before you could even access the app or piece of code. 

Not quite the same as working for a Formula One team 

 

Paul

Paul Ingram
Senior Principal Consultant
Red Oak Consulting

Recent Posts

Beyond SLAs

What Really Defines Excellent Support? A green SLA dashboard doesn’t necessarily mean you’re delivering excellent support. When people think about excellent IT support, they often

Read More »

It Costs How Much?!

Roco’s Beginner’s Guide to the Real Cost of HPC Hello Again, Humans! I’ve been having another educational adventure. This time, Jorge has been teaching me

Read More »

Discover how Red Oak Consulting can help your organisation get the very best from High-Performance Computing and Cloud Computing

Book a meeting

Get in touch with our team of HPC experts to find out how we can help you with your HPC, AI & Cloud Computing requirements across:

  • Strategy & Planning
  • Procurement
  • Implementation
  • Expansion & Optimisation
  • Maintenance & Support
  • Research Computing

Call us on
+44 (0)1242 806 188

Experts available:
9:00am – 5:30pm GMT

Name

Download Whitepaper

HPC and Formula One

"*" indicates required fields

Name*

Download Brochure

HPC AI – Deep Dive

"*" indicates required fields

Name*

Take something useful for when the time is right.

Download a short, shareable pack:

Because in a fast-moving market, staying connected matters — and timing is everything.

Download Brochure

HPC Procurement

"*" indicates required fields

Name*