Search by job, company or skills

Site Reliability Engineer (SRE) GCP Platform

  • Posted 10 hours ago
  • Be among the first 10 applicants

Job Description

Job Description

Role - Site Reliability Engineer (SRE) – GCP Platform

Location - Bangalore (O Shaughnessy Road)

Work type - Work from Office

Experience - 8+ Years

Key Responsibilities

Reliability & Operations

  • Own end-to-end production systems reliability, availability, scalability, cost and performance.
  • Drive measurable improvements in MTTR, MTTA, and incident response practices using automation and runbook additions and process enhancements.
  • Participate in 24x7 on-call rotations and handle high-severity incidents and document the learnings on ongoing basis.
  • Establish and manage SLI, SLO, SLA, Error Budgets, and operational metrics for mission critical services and partner with engineering teams with full accountability for upholding the SLOs.
  • Partner with the various engineering, operations and cloud management teams to deliver highly reliable service in a timely manner.

Cloud & Infrastructure

  • Design, deploy, and manage infrastructure on Google Cloud Platform (GCP).
  • Work extensively on:
  • GKE (Kubernetes Engine)
  • Compute, networking, IAM, Load Balancers, TLS Certs
  • BigQuery, Pub/Sub, cloud logging enhancement, metrics and logs analysis
  • Implement and manage infrastructure using Terraform (Infrastructure as Code).

Kubernetes & Containers

  • Deploy and manage containerized workloads using Kubernetes (GKE).
  • Troubleshoot issues related to:
  • Pods, nodes, networking, storage, services on an ongoing basis
  • Manage deployments using Helm, YAML, and rollout strategies (Canary/Blue-Green).

Automation & CI/CD

  • Build and maintain CI/CD pipelines using:
  • Jenkins (pipeline-based, Groovy / Shell / Python scripting)
  • Strong experience in using GitHub as a PowerUser
  • Develop automation using Python and Shell scripting.
  • Reduce operational toil through automation initiatives.

Observability & Monitoring

  • Implement and manage monitoring systems using:
  • Dynatrace, Grafana, logs and metrics explorer
  • Work with logs, metrics, and traces for deep observability to identify trends and arrest problems proactively.
  • Define alerting strategies based on system behaviour and SLOs and create runbooks.
  • Work alongside operations teams to identify, fix the production incidents and own the problem resolution.
  • Work with engineering teams to isolate infra and application issues and set up right tooling for debugging production incidents.

System & Application Troubleshooting

  • Perform deep troubleshooting for:
  • Distributed systems
  • Microservices-based architectures on containerised workloads
  • Java and Golang applications
  • Strong debugging of:
  • Application issues
  • Infrastructure issues
  • Network-related problems

Plan and execute continuous improvement

  • Identify and eliminate repetitive manual tasks.
  • Drive reliability engineering practices and culture. (DRY – Don't Repeat Yourself)
  • Collaborate with development teams to improve system design and resilience.

Technical Skills

Cloud & Platform

  • Strong expertise in Google Cloud Platform (GCP):
  • GKE, VPC, IAM, Load Balancing, LB, Certs, KMS, logs and metrics exploration
  • BigQuery, Pub/Sub
  • Good understanding of cloud architecture and landing zones

Infrastructure as Code

  • Strong hands-on experience with Terraform
  • Ability to write and debug Terraform code from scratch

Containers & Orchestration

  • Deep expertise in:
  • Kubernetes (GKE)
  • Docker
  • Strong troubleshooting experience in Kubernetes environments

CI/CD & Automation

  • Hands-on experience with:
  • Jenkins (pipeline-based CI/CD)
  • GitHub
  • Strong scripting skills:
  • Python (preferred)
  • Shell scripting
  • Experience with automation frameworks and tooling

Observability

  • Experience with:
  • Dynatrace / Grafana
  • Log, metrics, and trace-based monitoring

Programming & Debugging

  • Working knowledge of:
  • Java and/or Golang applications
  • Strong debugging skills across application and infrastructure layers

Linux & Networking

  • Strong Linux fundamentals
  • Deep understanding of TCP/IP networking
  • Ability to debug network issues in distributed systems

Reliability Engineering Skills

  • Solid understanding of:
  • SLI, SLO, SLA, Error Budgets
  • Demonstrable and Proven Experience improving:
  • MTTR, MTTA
  • Experience handling incident management lifecycle

Soft Skills

  • Strong analytical and troubleshooting mindset
  • Excellent communication and stakeholder management
  • Ability to work in high-pressure production environments
  • Ownership-driven and proactive approach

Preferred candidates with:

  • Experience in high-scale distributed systems
  • Exposure to banking/financial domain (optional but valuable)
  • Understanding of security and compliance practices
  • Experience with deployment strategies:
  • Canary, Blue-Green
  • Strong GCP + Kubernetes + Terraform core
  • Hands-on production troubleshooting expert
  • Good at automation + reducing toil
  • Deep understanding of SRE principles
  • Comfortable in 24x7 production environments

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 153307083

Beware of Scammers

We don’t charge money for job offers