Senior Network Engineer
Job Description
Our client is a global technology company operating large-scale AI/HPC and data center infrastructure across multiple international markets. As its AI cloud capabilities continue to scale, the company is looking for an experienced AI Cloud Network Operations Engineer to operate and optimize high-performance network infrastructure supporting large-scale GPU computing environments.
Job Responsibilities
- AI Network Operations — Monitor and operate large-scale data center and AI/HPC networks across switches, routers, optical links, bandwidth utilization and network health.
- Incident & Performance Troubleshooting — Own network incidents and troubleshoot complex performance issues including latency/jitter, packet loss, GPU-to-GPU communication and NCCL throughput degradation.
- Network Configuration & Changes — Execute and troubleshoot BGP, OSPF, VXLAN, EVPN and ECMP environments, including VLAN, routing, traffic isolation and bandwidth changes.
- AI/HPC Fabric Operations — Support high-performance GPU cluster networking using InfiniBand and/or RoCEv2, including lossless Ethernet technologies such as PFC and ECN.
- Monitoring & Automation — Improve network visibility, operational efficiency and reliability using monitoring platforms and Python/Go-based network automation.
Job Requirements
- 5+ years of Network Operations / Network Engineering experience within large-scale Cloud, Internet, Data Center, Carrier or HPC environments.
- Strong hands-on knowledge of BGP, OSPF, VXLAN, EVPN and ECMP, with independent troubleshooting capability on enterprise networking equipment.
- Practical exposure to InfiniBand and/or RoCEv2, with understanding of PFC, ECN and lossless networking for high-performance workloads.
- Experience operating large-scale infrastructure using vendors such as NVIDIA Spectrum/Quantum, Arista or Cisco, together with monitoring tools such as Prometheus, Grafana, Zabbix or equivalent.
- Exposure to large-scale GPU clusters / AI infrastructure is highly advantageous, particularly NVIDIA H100/GB200 environments, NCCL/MPI, NVIDIA UFM, network automation or optical/DCI networks.


