Search by job, company or skills

Technical Lead ML Platform & Infrastructure

  • Posted 20 hours ago
  • Be among the first 10 applicants

Job Description

Role Overview

We are looking for an experienced Senior Software Engineer or Technical Lead to design and build scalable machine learning platforms for training, optimizing, evaluating, packaging, and deploying large-scale AI models.

The role requires strong expertise in distributed systems, GPU computing, ML infrastructure, and production-grade software engineering. You will work closely with applied scientists, platform engineers, hardware teams, compiler teams, and infrastructure specialists.

Key Responsibilities

  • Design and develop distributed ML platform services and reusable libraries.
  • Build scalable training capabilities for large language and multimodal models.
  • Support data, tensor, pipeline, and model parallelism across multi-node GPU clusters.
  • Improve training throughput, GPU utilization, memory efficiency, communication performance, and fault recovery.
  • Define stable APIs and platform interfaces for model onboarding, training, evaluation, and deployment.
  • Integrate model optimization techniques such as quantization, pruning, knowledge distillation, and compression.
  • Build evaluation and artifact-management workflows to measure model quality and system performance.
  • Develop automated validation, CI/CD, regression testing, observability, and release pipelines.
  • Profile GPU workloads and resolve end-to-end system performance bottlenecks.
  • Build production-operational mechanisms, including metrics, alarms, dashboards, runbooks, and root-cause corrective actions.
  • Collaborate with model, compiler, runtime, hardware, security, infrastructure, and product teams.
  • Prepare technical designs, evaluate architecture trade-offs, and drive engineering decisions.
  • Mentor engineers and improve code reviews, design reviews, testing, and development practices.

Required Qualifications

  • Bachelor's degree in Computer Science, Engineering, or a related technical field.
  • 5+ years of professional software development experience.
  • Strong programming skills in C++, Java, C#, Python, or a similar language.
  • Experience designing distributed systems, high-performance computing platforms, or scalable backend systems.
  • Strong understanding of system design, concurrency, reliability, scalability, and performance optimization.
  • Experience as a technical lead, architect, mentor, or engineering team lead.
  • Experience across the software development lifecycle, including coding, testing, source control, build systems, deployment, and production support.

Preferred Qualifications

  • Experience with PyTorch, TensorFlow, JAX, NeMo, or Megatron-LM.
  • Experience building distributed ML training, inference, evaluation, or data platforms.
  • Knowledge of CUDA, GPU kernels, GPU profiling, and performance optimization.
  • Experience with containers, Kubernetes, cloud infrastructure, CI/CD, and observability tools.
  • Knowledge of model compression, quantization, pruning, distillation, compilation, or edge AI deployment.
  • Experience with large language models, multimodal models, and distributed GPU training.
  • Experience designing extensible platform APIs and reusable software frameworks.
  • Experience working with applied science, hardware, compiler, runtime, or product engineering teams.

Key Skills

Distributed ML Systems | GPU Computing | ML Infrastructure | PyTorch | CUDA | Kubernetes | CI/CD | Model Optimization | Performance Engineering | System Architecture | Technical Leadership

 

More Info

Job ID: 152360273

Beware of Scammers

We don’t charge money for job offers