Head of GPU Cluster Engineering
Head of GPU Cluster Engineering
Nava12-14 Years
- Posted 22 days ago
- Be among the first 10 applicants
Job Description
About Nava
Nava is building next-generation AI infrastructure and inference platforms designed to power the future of AI. We are looking for a Head of GPU Cluster Engineering to lead the architecture, deployment, and operations of large-scale GPU clusters.
This is a highly strategic leadership role responsible for defining the engineering vision, driving execution across multiple infrastructure domains, and ensuring our GPU platforms deliver world-class performance, scalability, and reliability.
What You'll Do
Platform Leadership
Nava is building next-generation AI infrastructure and inference platforms designed to power the future of AI. We are looking for a Head of GPU Cluster Engineering to lead the architecture, deployment, and operations of large-scale GPU clusters.
This is a highly strategic leadership role responsible for defining the engineering vision, driving execution across multiple infrastructure domains, and ensuring our GPU platforms deliver world-class performance, scalability, and reliability.
What You'll Do
Platform Leadership
- Own end-to-end engineering outcomes, from reference architecture through production operations.
- Define the technical vision, architecture, and operating model for large-scale GPU infrastructure.
- Establish engineering standards, design principles, and operational excellence across the platform.
- Drive platform scalability, resiliency, and performance to support rapidly growing AI workloads.
- Lead the design and evolution of GPU cluster architectures across compute, networking, storage, orchestration, and observability.
- Evaluate and drive technology decisions around GPUs, interconnects, storage, Kubernetes, scheduling, and AI infrastructure software.
- Review and approve architecture decisions while balancing performance, reliability, cost, and operational complexity.
- Drive infrastructure standardization and automation across deployments.
- Lead execution across multiple engineering pillars:
- GPU Compute
- High-Speed Networking
- Storage
- Platform Engineering
- Site Reliability Engineering (SRE)
- Infrastructure Automation
- Partner closely with Product, Supply Chain, Data Centre Operations, and Customer Success teams to ensure successful platform delivery.
- Act as the technical escalation point for critical engineering decisions.
- Own engineering readiness gates from design through production deployment.
- Drive release planning, operational reviews, risk assessments, and post-incident analysis.
- Establish SLAs, SLOs, and operational metrics for platform health and reliability.
- Champion automation, observability, incident management, and continuous improvement across the engineering organization.
- Build, mentor, and scale a high-performing engineering organization.
- Develop technical leaders across infrastructure disciplines.
- Foster a culture of engineering excellence, ownership, collaboration, and innovation.
- Support hiring and talent development for critical infrastructure roles.
- 12+ years of experience in infrastructure engineering, distributed systems, cloud platforms, or AI infrastructure.
- Proven experience leading large-scale infrastructure or platform engineering teams.
- Deep understanding of GPU clusters, AI infrastructure, HPC, or large-scale distributed systems.
- Strong expertise across:
- GPU Compute (NVIDIA ecosystem preferred)
- High-performance networking (InfiniBand, RoCE, RDMA)
- Kubernetes and container orchestration
- Distributed storage systems
- Infrastructure automation and observability
- Linux systems and platform engineering
- Experience designing highly available, scalable production infrastructure.
- Strong architectural thinking with the ability to balance technical excellence and business priorities.
- Excellent stakeholder management and cross-functional leadership skills.
- Experience building AI factories, GPU cloud platforms, or large-scale inference infrastructure.
- Familiarity with CUDA, NCCL, Slurm, Ray, Kubeflow, or similar AI infrastructure technologies.
- Experience working with hyperscalers, cloud providers, or AI-first technology companies.
- Exposure to multi-region or global infrastructure deployments.
- Build one of the world's leading AI infrastructure platforms.
- Lead the engineering strategy behind next-generation GPU clusters and AI factories.
- Work alongside world-class engineers solving some of the most challenging infrastructure problems in AI.
- Shape the future of AI infrastructure from architecture through production at global scale.
More Info
Key Skills
RDMA
NVIDIA ecosystem
Distributed storage systems
RoCE
platform engineering
Linux systems
GPU Compute
container orchestration
observability
High-performance networking
