Search Jobs

Search by job, company or skills

Sr. Principal Infrastr Eng (HPC)

Sr. Principal Infrastr Eng (HPC)

Mphasis
Early Applicant
  • Posted 8 days ago
  • Be among the first 10 applicants

Job Description

Skill - HPC, Slurm, ClearML, Linux

Location - Bengaluru

Experience -7-10yrs

Early joiners are preferred

Job Summary:

We are seeking a highly skilled Senior Principal Infrastructure Engineer to administer and optimize our ClearML Server and associated infrastructure. The ideal candidate will have a strong background in MLOps platforms, HPC execution environments, and containerized solutions, with a focus on supporting GPU-based AI/ML workloads. This role requires a proactive approach to managing resources, ensuring security, and automating processes to enhance operational efficiency.

Responsibilities:

  • Administer ClearML Server, including management of agents, execution queues, projects, users, roles, experiment tracking, pipelines, datasets, artifacts, and model registry.
  • Configure ClearML Agents on CPU and GPU worker nodes, integrating with HPC execution platforms such as Slurm, PBS Professional, or Kubernetes.
  • Support GPU-based AI/ML workloads utilizing NVIDIA drivers, CUDA, NCCL, UCX, and containerized environments.
  • Maintain secure container execution using technologies like Apptainer/Singularity, Docker, Enroot, Pyxis, or Kubernetes.
  • Implement confidential-computing controls leveraging AMD SEV-SNP, Intel TDX, NVIDIA Confidential Computing, Secure Boot, TPM, and remote attestation.
  • Integrate authentication, RBAC, TLS certificates, secrets management, and audit controls for ClearML and confidential workloads.
  • Monitor ClearML services, agents, queues, GPU utilization, task failures, scheduler integration, and overall platform health.
  • Automate deployment, configuration, monitoring, and troubleshooting processes using Python, Bash, Ansible, and Git.

Mandatory Skills:

  • Strong Linux administration skills, particularly with RHEL, SLES, Rocky Linux, or Ubuntu.
  • Hands-on experience with ClearML administration or a comparable MLOps platform.
  • Familiarity with HPC schedulers such as Slurm, PBS Professional, or LSF.
  • Knowledge of GPU platforms, including NVIDIA drivers, CUDA, and distributed training basics.
  • Proficiency in container technologies: Apptainer/Singularity, Docker, Enroot, Pyxis, or Kubernetes.
  • Strong scripting and automation skills in Python, Bash, Ansible, and Git.
  • Understanding of security fundamentals, including RBAC, IAM, TLS, certificates, secrets management, secure boot, and audit logging.
  • Familiarity with confidential-computing concepts such as TEE, encrypted memory, TPM, remote attestation, AMD SEV-SNP, Intel TDX, or equivalent technologies.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

audit logging

ClearML

UCX

NVIDIA drivers

Intel TDX

Singularity

Apptainer

secure boot

Slurm

secrets management

Enroot

TLS certificates

AMD SEV-SNP

Pyxis

NCCL

About Company