Data Centre Field Operations Engineer
Data Centre Field Operations Engineer
Nava5-7 Years
- Posted 15 hours ago
- Be among the first 10 applicants
Job Description
About The Role
We are seeking an Infrastructure Operations Lead to oversee operational performance across our high-density GPU data center footprint.
This is an on-site vendor governance and technical escalation role. You will act as our primary operational anchor—managing vendor performance and holding our Managed Service Provider (MSP), Co-Location Facility Provider, and Enterprise Customers accountable to their operational standards, service contracts, and SLAs.
Key Responsibilities
We are seeking an Infrastructure Operations Lead to oversee operational performance across our high-density GPU data center footprint.
This is an on-site vendor governance and technical escalation role. You will act as our primary operational anchor—managing vendor performance and holding our Managed Service Provider (MSP), Co-Location Facility Provider, and Enterprise Customers accountable to their operational standards, service contracts, and SLAs.
Key Responsibilities
- 360° Vendor & Service Governance: Oversee third-party service delivery to enforce hardware repair SLAs, ticket response times, and spare parts/RMA workflows. Ensure the co-location provider meets power and high-density cooling guarantees, while keeping enterprise customers within agreed operating boundaries.
- Crisis Management & Incident Command: Act as On-Site Incident Commander during high-severity outages or infrastructure degradation. Lead recovery efforts by coordinating across vendor engineering teams and internal stakeholders.
- Incident & Executive Communication: Provide concise, real-time updates to executive leadership and enterprise clients during major outages, translating complex technical failures into clear operational impact.
- Root Cause Analysis (RCA) & Post-Mortems: Lead technical investigations following major incidents. Audit and challenge technical RCAs provided by vendors (MSP/Colo) to identify systemic hardware, environmental, or workflow issues, ensuring permanent corrective actions are executed.
- Technical Escalation & Standards: Serve as the Subject Matter Expert (SME) for complex GPU/HPC hardware escalations that exceed standard vendor runbooks, and maintain ownership of operational standards.
- 5+ years in data center infrastructure operations, with a strong focus on third-party vendor management, SLA enforcement, and service delivery for high-density GPU/HPC platforms.
- Crisis management: Proven ability to drive incident recovery and direct third-party vendor teams during critical data center outages under high-pressure conditions.
- Stakeholder Communication: Outstanding verbal and written communication skills to bridge technical vendor teams, internal stakeholders, and enterprise client representatives.
- Root Cause Analysis (RCA) Expertise: Demonstrated capability in methodical troubleshooting, post-incident investigations, and auditing vendor-supplied RCAs (using frameworks like 5-Whys or Fishbone) to drive long-term infrastructure reliability.
- Technical & Facility Knowledge: Solid understanding of enterprise server platforms (NVIDIA Blackwell GPU architectures, high-speed networking RoCE,infiniband) and data center facility constraints (high-density power distribution, liquid cooling).
- Education: Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.
More Info
Key Skills
RoCE
incident recovery
data center facility constraints
high-density power distribution
5-Whys
SLA enforcement
GPU HPC hardware
liquid cooling
high-speed networking
NVIDIA Blackwell GPU architectures
