Search by job, company or skills

Senior Software Engineer (SRE)

Senior Software Engineer (SRE)

Maya
7-9 Years
Not Disclosed
  • Posted 3 hours ago
  • Be among the first 10 applicants

Job Description

Role Overview

We are looking for a Software Engineer specializing in Site Reliability Engineering (SRE) to build reliable, scalable, and highly available systems. This role combines software engineering and reliability engineering, focusing on developing automation, platform capabilities, observability solutions, and self-healing systems that improve the resilience and performance of our cloud-native applications.

The ideal candidate is a strong software engineer who can apply engineering principles to infrastructure, operations, and reliability challenges.

Key Responsibilities

  • Design and develop software, tools, and automation that improve platform reliability and operational efficiency.
  • Build and maintain cloud-native applications and services running on Kubernetes.
  • Develop self-service platforms, automation frameworks, and reliability tooling for engineering teams.
  • Implement observability solutions, including monitoring, logging, tracing, and alerting.
  • Define and manage SLOs, SLIs, error budgets, and reliability metrics.
  • Improve system resilience through automation, incident prevention, root cause analysis, and performance optimization.
  • Partner with product and engineering teams to ensure services are scalable, observable, and operationally ready.
  • Contribute to CI/CD, infrastructure automation, and GitOps practices.
  • Participate in incident response and drive long-term reliability improvements.

Required Qualifications

  • 7+ years of experience in Software Engineering, SRE, Platform Engineering, or related fields.
  • Strong programming experience in Go, Java, Python, or similar languages.
  • Experience designing and building distributed systems, APIs, and microservices.
  • Hands-on expertise with Kubernetes and containerized environments.
  • Experience with Terraform, Infrastructure as Code, and cloud platforms (AWS, Azure, or GCP).
  • Knowledge of CI/CD and GitOps tools such as GitLab CI, ArgoCD, or FluxCD.
  • Experience with observability platforms such as Dynatrace, Prometheus, Grafana, Datadog, or OpenTelemetry.
  • Strong understanding of high availability, resiliency, performance optimization, and incident management.
  • Familiarity with cloud security and governance best practices.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

About Company