Search by job, company or skills

Senior DevOps Engineer

Senior DevOps Engineer

tarento group
3-5 Years
Not Disclosed
Early Applicant
  • Posted 19 days ago
  • Be among the first 50 applicants

Job Description

About Tarento

Tarento is a fast-growing technology consulting company headquartered in Stockholm, with a strong presence in India and clients across the globe. We specialize in digital transformation, product engineering, and enterprise solutions, working across diverse industries including retail, manufacturing, and healthcare. Our teams combine Nordic values with Indian expertise to deliver innovative, scalable, and high-impact solutions.

We're proud to be recognized as a Great Place to Work, a testament to our inclusive culture, strong leadership, and commitment to employee well-being and growth. At Tarento, you'll be part of a collaborative environment where ideas are valued, learning is continuous, and careers are built on passion and purpose.

Role Overview

We are looking for a Senior DevOps Engineer with 3+ years of experience in Multi-Cloud, Kubernetes, Jenkins, and automation. The role involves building and managing scalable infrastructure, driving automation, supporting migrations, and collaborating with clients.

Key Responsibilities

  • Design, build, and maintain cloud infrastructure (AWS/GCP/Azure) using IaC (Terraform/Pulumi/CloudFormation)
  • Architect and manage CI/CD pipelines for application and ML model deployment
  • Build and maintain Kubernetes clusters; manage containerized workloads for services and ML inference
  • Own observability stack — logging, metrics, tracing, alerting (Prometheus, Grafana, ELK/EFK, Datadog)
  • Support ML lifecycle infrastructure: feature stores, model registries, training pipelines, GPU cluster management
  • Manage GPU/accelerator provisioning and cost optimization (spot instances, autoscaling, reserved capacity)
  • Drive security best practices — IAM, secrets management, network policies, compliance (SOC2/ISO)
  • Lead incident response, on-call rotations, and postmortems; drive SRE culture (SLOs/SLIs)
  • Mentor junior engineers; collaborate with data science/ML teams on infra requirements
  • Manage multi-region/multi-cloud deployments and disaster recovery strategy
  • Own capacity planning, cost governance (FinOps) across cloud + ML compute

Required Skills & Experience

Cloud Platforms

  • Deep expertise in at least one hyperscaler (AWS/GCP/Azure), working knowledge of a second
  • Compute, networking (VPC, load balancers, CDN), storage, IAM

Infrastructure as Code & Automation

  • Terraform / Pulumi / CloudFormation
  • Ansible / Chef / Puppet
  • GitOps (ArgoCD/Flux)

Containers & Orchestration

  • Docker, Kubernetes (EKS/GKE/AKS or self-managed)
  • Helm, Kustomize
  • Service mesh (Istio/Linkerd) — good to have

CI/CD

  • Jenkins, GitLab CI, GitHub Actions, ArgoCD, CircleCI

Observability & Reliability

  • Prometheus, Grafana, Alertmanager
  • ELK/EFK stack, Loki
  • Datadog/New Relic
  • Distributed tracing (Jaeger/OpenTelemetry)

Security

  • Vault/Secrets Manager, IAM policy design
  • Network security, container security (Trivy, Falco)
  • Compliance frameworks

Scripting/Programming

  • Python (strong — for automation & ML tooling integration)
  • Bash/Go (good to have)

Databases & Messaging

  • SQL/NoSQL operations at scale (Postgres, MongoDB, Redis)
  • Kafka/RabbitMQ/Pub-Sub for streaming ML data pipelines

Preferred Qualifications

  • Certifications: AWS/GCP/Azure Solutions Architect, CKA/CKAD
  • Experience running LLM inference infrastructure at scale (vLLM, TGI, quantization-aware serving)
  • Experience with cost optimization for large-scale GPU training
  • Contributions to open-source infra/MLOps tooling

Education

BE/BTech, Information Technology, or related discipline.

More Info

Job Type:
Industry:
Employment Type:

About Company

Similar Jobs

5-8 yrs
Gurugram, India, Gurugram
Skills:
Sql, Rabbitmq, Grafana, RDS, Kafka, Nosql, Iam, Terraform, Git, S3, Vpc, Elk, Ansible, AWS, Prometheus, Kubernetes, Docker, Ec2, Cache, Jenkins, Linux Shell Scripting, AWS CLI scripting, message-queuing, ArgoCD, stream-processing, Amazon EKS, AI tools for Infrastructure as Code
6-8 yrs
Noida, India
Skills:
Python Scripting, Terraform, Docker, Prometheus, Grafana, Helm, Kubernetes Cluster Management, OpenShift Administration, CI CD Pipeline Design Implementation, DevSecOps Automation, Argo CD GitOps
5-7 yrs
Noida, India
Skills:
PowerShell, Bash, ARM templates, Terraform, Microsoft Azure, Python, Azure DevOps, Azure Migrate, GitHub Actions, Bicep, Azure SQL Managed Instance, Azure AI technologies
6-10 yrs
Gurugram, Gurugram, India
Skills:
Elk, Networking, Cloudformation, Prometheus, Bash, Grafana, Datadog, Docker, Terraform, Sonarqube, Azure, Kubernetes, Python, AWS, Prisma Cloud, Wiz, Go, Snyk, Falco, Trivy, Infrastructure-as-Code
7-9 yrs
Noida, India
Skills:
sentinel , VMware, Citrix, PowerShell, Terraform, Ansible, Oracle, Azure, Python, Azure DevOps, Infrastructure as Code, Key Vault, Active Directory, Microsoft Defender for Cloud