Senior DevOps Engineer
- Posted 19 days ago
- Be among the first 50 applicants
Job Description
About Tarento
Tarento is a fast-growing technology consulting company headquartered in Stockholm, with a strong presence in India and clients across the globe. We specialize in digital transformation, product engineering, and enterprise solutions, working across diverse industries including retail, manufacturing, and healthcare. Our teams combine Nordic values with Indian expertise to deliver innovative, scalable, and high-impact solutions.
We're proud to be recognized as a Great Place to Work, a testament to our inclusive culture, strong leadership, and commitment to employee well-being and growth. At Tarento, you'll be part of a collaborative environment where ideas are valued, learning is continuous, and careers are built on passion and purpose.
Role Overview
We are looking for a Senior DevOps Engineer with 3+ years of experience in Multi-Cloud, Kubernetes, Jenkins, and automation. The role involves building and managing scalable infrastructure, driving automation, supporting migrations, and collaborating with clients.
Key Responsibilities
Cloud Platforms
BE/BTech, Information Technology, or related discipline.
Tarento is a fast-growing technology consulting company headquartered in Stockholm, with a strong presence in India and clients across the globe. We specialize in digital transformation, product engineering, and enterprise solutions, working across diverse industries including retail, manufacturing, and healthcare. Our teams combine Nordic values with Indian expertise to deliver innovative, scalable, and high-impact solutions.
We're proud to be recognized as a Great Place to Work, a testament to our inclusive culture, strong leadership, and commitment to employee well-being and growth. At Tarento, you'll be part of a collaborative environment where ideas are valued, learning is continuous, and careers are built on passion and purpose.
Role Overview
We are looking for a Senior DevOps Engineer with 3+ years of experience in Multi-Cloud, Kubernetes, Jenkins, and automation. The role involves building and managing scalable infrastructure, driving automation, supporting migrations, and collaborating with clients.
Key Responsibilities
- Design, build, and maintain cloud infrastructure (AWS/GCP/Azure) using IaC (Terraform/Pulumi/CloudFormation)
- Architect and manage CI/CD pipelines for application and ML model deployment
- Build and maintain Kubernetes clusters; manage containerized workloads for services and ML inference
- Own observability stack — logging, metrics, tracing, alerting (Prometheus, Grafana, ELK/EFK, Datadog)
- Support ML lifecycle infrastructure: feature stores, model registries, training pipelines, GPU cluster management
- Manage GPU/accelerator provisioning and cost optimization (spot instances, autoscaling, reserved capacity)
- Drive security best practices — IAM, secrets management, network policies, compliance (SOC2/ISO)
- Lead incident response, on-call rotations, and postmortems; drive SRE culture (SLOs/SLIs)
- Mentor junior engineers; collaborate with data science/ML teams on infra requirements
- Manage multi-region/multi-cloud deployments and disaster recovery strategy
- Own capacity planning, cost governance (FinOps) across cloud + ML compute
Cloud Platforms
- Deep expertise in at least one hyperscaler (AWS/GCP/Azure), working knowledge of a second
- Compute, networking (VPC, load balancers, CDN), storage, IAM
- Terraform / Pulumi / CloudFormation
- Ansible / Chef / Puppet
- GitOps (ArgoCD/Flux)
- Docker, Kubernetes (EKS/GKE/AKS or self-managed)
- Helm, Kustomize
- Service mesh (Istio/Linkerd) — good to have
- Jenkins, GitLab CI, GitHub Actions, ArgoCD, CircleCI
- Prometheus, Grafana, Alertmanager
- ELK/EFK stack, Loki
- Datadog/New Relic
- Distributed tracing (Jaeger/OpenTelemetry)
- Vault/Secrets Manager, IAM policy design
- Network security, container security (Trivy, Falco)
- Compliance frameworks
- Python (strong — for automation & ML tooling integration)
- Bash/Go (good to have)
- SQL/NoSQL operations at scale (Postgres, MongoDB, Redis)
- Kafka/RabbitMQ/Pub-Sub for streaming ML data pipelines
- Certifications: AWS/GCP/Azure Solutions Architect, CKA/CKAD
- Experience running LLM inference infrastructure at scale (vLLM, TGI, quantization-aware serving)
- Experience with cost optimization for large-scale GPU training
- Contributions to open-source infra/MLOps tooling
BE/BTech, Information Technology, or related discipline.





