Cloud infrastructure
10 GPU cluster operations engineers
These ten professionals have worked as GPU and HPC cluster operations engineers at GPU cloud providers, chip and cloud companies and research and energy organizations. A reader can learn how GPU cluster operations engineers approach their work from people who have done it.
Professionals to explore
01—10Aaron Hettinger
LinkedInExperience: Lead Staff Network Engineer - HPC Cluster Operations · NVIDIA
Lead Staff Network Engineer for HPC Cluster Operations at NVIDIA since 2024; the profile says the work includes designing, deploying and provisioning large-scale GPU clusters interconnected with NVLink, InfiniBand or Ethernet. Earlier roles include Lead Staff Network Engineer at Pure Storage (2018–2024) and data center operations engineering at Pure Storage and Sauce Labs.
Kristopher Whetham
LinkedInExperience: Fleet Engineering Manager · CoreWeave
Fleet Engineering Manager at CoreWeave since 2024; the profile describes an engineering and operations leader for AI/ML, big data storage and high-performance computing platforms. Before CoreWeave he was Director of DevOps (2023–2024), Managed Services Manager (2021–2024) and Managed Services Engineer (2019–2024) at Penguin Computing.
Kyle Hutson
LinkedInExperience: HPC Operations Engineer · Lambda
HPC Operations Engineer at Lambda since 2024; the profile describes Lambda as a GPU compute provider offering on-demand and reserved cloud NVIDIA GPUs. Earlier he was System Engineer at KanREN (2023–2024) and System Administrator at Kansas State University (2012–2023).
Lee Hobson
LinkedInExperience: Senior HPC Engineer · Fluidstack
Senior HPC Engineer at Fluidstack since 2024, where the profile says he develops and optimizes GPU cluster infrastructure; the headline describes site reliability work on GPU and CPU clusters and AI infrastructure. Earlier he was Lead HPC Systems Engineer (2020–2024) and HPC Systems Engineer (2017–2020) at OCF Limited.
Malik Chérif
LinkedInExperience: HPC/GPU Cluster Sytem Administrator · TotalEnergies
HPC/GPU cluster system administrator at TotalEnergies since 2019; the profile mentions the IBM Pangea III supercomputer and providing GPU HPC engineering to operate and support mission-critical infrastructure services. Earlier he was HPC Cluster System Administrator at Safran (2016–2018) and a system and network engineer at SF2i (2018–2019).
Prabu Sekar
LinkedInExperience: Senior Engineer - HPC Operations · Core42
Senior Engineer - HPC Operations at Core42 since 2026, after serving as Engineer - HPC Systems there (2023–2026). The profile headline lists large-scale GPU clusters of more than 20,000 GPUs, Slurm, InfiniBand and RDMA. Earlier roles include Research Engineer, HPC at Nanyang Technological University Singapore (2019–2022) and HPC Systems Engineer at Harrington HPC Microsystems.
Ryan Crawford
LinkedInExperience: Sr HPC Operations Engineer · Lambda
Sr HPC Operations Engineer at Lambda since 2024 and Sr HPC Systems Validation Engineer there since 2026. The profile describes deploying and optimizing mission-critical GPU clusters and delivering more than 30 cluster installations. Earlier roles include HPC Team Lead at Translucent Services (2023–2024) and HPC DevOps in the United States Air Force (2016–2022).
Venkatesh G
LinkedInExperience: Senior Cloud Operations Engineer - GPU Cluster Health & AI Infrastructure · Oracle
Senior Cloud Operations Engineer for GPU Cluster Health & AI Infrastructure at Oracle since 2026, after compute data plane and fleet management operations roles at Oracle from 2023. The profile describes work at Oracle Cloud Infrastructure on infrastructure reliability, fleet management, incident response and automation. Earlier roles include Linux System Administrator at Newt Global.
Wallace P.
LinkedInExperience: Senior Fleet Reliability Operations Engineer II · CoreWeave
Senior Fleet Reliability Operations Engineer II at CoreWeave since 2026, after Senior Fleet Reliability Operations Engineer I (2024–2026) and HPC Operations Engineer (2023–2024) there. The profile says the HPC operations role covered troubleshooting and debugging NVIDIA GPUs, delivering nodes and automating processes. Earlier: Linux System Administrator at the Whitehead Institute (2019–2023).
Zach Merendino
LinkedInExperience: HPC Operations Engineer · CoreWeave
Former HPC Operations Engineer at CoreWeave (2019–2025), where the rows show six years in the role, and Staff Operations Engineer there since 2025. The profile describes CoreWeave as an AI hyperscaler, and the rows also list a support engineering role at the company.
Choose the right perspective
Match the person's layer of the stack to your question: fleet reliability and node health engineers for hardware failures and repair loops, network operations engineers for InfiniBand and NVLink fabrics, and HPC systems engineers for schedulers, storage and user workloads. Also consider whether you need someone from a GPU cloud provider serving many customers or someone who runs a cluster for one organization.
Questions to take into the conversation
- 01How do you detect and drain unhealthy GPU nodes before they break a long training job?
- 02What causes most downtime in a large GPU cluster, and how do you track it?
- 03How do you plan maintenance and firmware updates on a cluster that is busy around the clock?
Reaching GPU cluster operations engineers
How can I contact one of these GPU cluster operations engineers?
Pick one of the 10 people on this page and choose "Book a paid call", or describe your project to find others. You offer a fee for a 15-minute call or a written answer, Instant Expert finds the person's work email and sends the invitation, and they decide whether to accept. The list covers 7 companies, including NVIDIA, CoreWeave and Lambda, based on profile data retrieved on October 9, 2026. Being listed here doesn't mean someone has agreed to take calls.
How much does it cost to reach GPU cluster operations engineers?
You choose the offer, starting at $5 per person. It includes Instant Expert's 20% fee, so an offer that pays the person $100 costs you $125. You're charged only when the person books the call or sends the answer.
What if they don't reply?
You pay nothing. Instant Expert sends follow-up reminders, and if the person hasn't booked or answered within 7 days, the request expires and any hold on your card is released. You can invite several of the 10 people on this list at once and cap your total spend, so you pay only for the ones who accept.
Can an AI agent ask one of these GPU cluster operations engineers a question?
Yes. An agent with a USDC wallet on Base can ask one question without an Instant Expert account. It names the person, for example by the LinkedIn URL on this page, and pays per ask over x402 (HTTP 402). The person answers in writing or by voice note, and if nobody answers within 7 days, the payment goes back to the wallet automatically. x402 docs
About this directory
This is a professional research starting point based on business profile data retrieved on . Titles and companies reflect that source snapshot and may describe past or present roles. Check the linked profiles for current details. Inclusion does not imply Instant Expert membership or availability.
Request a correction or removal