Search Jobs

Search by job, company or skills

Data Centre Field Operations Engineer

Data Centre Field Operations Engineer

Nava
5-7 Years
  • Posted 15 hours ago
  • Be among the first 10 applicants

Job Description

About The Role

We are seeking an Infrastructure Operations Lead to oversee operational performance across our high-density GPU data center footprint.

This is an on-site vendor governance and technical escalation role. You will act as our primary operational anchor—managing vendor performance and holding our Managed Service Provider (MSP), Co-Location Facility Provider, and Enterprise Customers accountable to their operational standards, service contracts, and SLAs.

Key Responsibilities

  • 360° Vendor & Service Governance: Oversee third-party service delivery to enforce hardware repair SLAs, ticket response times, and spare parts/RMA workflows. Ensure the co-location provider meets power and high-density cooling guarantees, while keeping enterprise customers within agreed operating boundaries.
  • Crisis Management & Incident Command: Act as On-Site Incident Commander during high-severity outages or infrastructure degradation. Lead recovery efforts by coordinating across vendor engineering teams and internal stakeholders.
  • Incident & Executive Communication: Provide concise, real-time updates to executive leadership and enterprise clients during major outages, translating complex technical failures into clear operational impact.
  • Root Cause Analysis (RCA) & Post-Mortems: Lead technical investigations following major incidents. Audit and challenge technical RCAs provided by vendors (MSP/Colo) to identify systemic hardware, environmental, or workflow issues, ensuring permanent corrective actions are executed.
  • Technical Escalation & Standards: Serve as the Subject Matter Expert (SME) for complex GPU/HPC hardware escalations that exceed standard vendor runbooks, and maintain ownership of operational standards.

Requirements

  • 5+ years in data center infrastructure operations, with a strong focus on third-party vendor management, SLA enforcement, and service delivery for high-density GPU/HPC platforms.
  • Crisis management: Proven ability to drive incident recovery and direct third-party vendor teams during critical data center outages under high-pressure conditions.
  • Stakeholder Communication: Outstanding verbal and written communication skills to bridge technical vendor teams, internal stakeholders, and enterprise client representatives.
  • Root Cause Analysis (RCA) Expertise: Demonstrated capability in methodical troubleshooting, post-incident investigations, and auditing vendor-supplied RCAs (using frameworks like 5-Whys or Fishbone) to drive long-term infrastructure reliability.
  • Technical & Facility Knowledge: Solid understanding of enterprise server platforms (NVIDIA Blackwell GPU architectures, high-speed networking RoCE,infiniband) and data center facility constraints (high-density power distribution, liquid cooling).
  • Education: Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.

Skills: oems,infrastructure,commissioning,data,operations,field execution,skills,critical infrastructure,maintenance,contractors

More Info

Job Type:
Industry:
Employment Type:

Key Skills

RoCE

incident recovery

data center facility constraints

high-density power distribution

5-Whys

SLA enforcement

GPU HPC hardware

liquid cooling

high-speed networking

NVIDIA Blackwell GPU architectures

About Company