Search by job, company or skills

KMC Solutions

Senior Infrastructure Engineer (Linux / Automation) - Full Remote

5-7 Years
Save
  • Posted a day ago
  • Be among the first 10 applicants
Early Applicant

Job Description

About the Role

We are looking for an experienced Senior Site Reliability Engineer to help build, automate, and operate our large-scale Linux infrastructure. You will be responsible for maintaining highly available platforms, improving automation, managing infrastructure as code, and supporting production environments across multiple data centers.

This role is ideal for engineers who enjoy solving complex infrastructure problems, writing production-quality automation, and working closely with global engineering teams.

What You'll Do

  • Lead L3 platform operations and manage production incidents from start to resolution.
  • Troubleshoot Linux infrastructure, networking, storage, hardware, and system performance issues.
  • Build and maintain infrastructure automation using Ansible, Terraform/OpenTofu, and Python.
  • Provision and configure bare-metal Linux servers using automated deployment methods.
  • Perform software upgrades, configuration management, and infrastructure improvements.
  • Develop Infrastructure-as-Code modules following best engineering practices.
  • Create automation tools to improve operational efficiency and eliminate repetitive tasks.
  • Collaborate with global Infrastructure and Operations teams during shift handovers and major incidents.
  • Review code, improve documentation, and contribute to engineering standards and best practices.

What We're Looking For

  • 5+ years of experience in Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering.
  • Strong Linux administration experience (Debian or Ubuntu preferred).
  • Hands-on experience with bare-metal servers and data center environments.
  • Strong scripting skills using Bash.
  • Production-level Python development experience.
  • Advanced experience with Ansible automation.
  • Strong knowledge of Terraform or OpenTofu.
  • Experience with Git, CI/CD, and Infrastructure as Code.
  • Excellent English communication skills.

Nice to Have

  • PXE, iPXE, Preseed, Kickstart, MAAS, or fleet provisioning tools.
  • Kubernetes, OpenShift Virtualization, or Proxmox.
  • Prometheus, VictoriaMetrics, Harbor, or ArgoCD.
  • DCIM platforms such as NetBox or Nautobot.
  • HPC, AI infrastructure, GPU clusters, or InfiniBand.
  • Experience with SLURM, Portworx, Pure Storage, VAST, or NVIDIA BlueField.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 151350907