Search by job, company or skills

VP, Site Reliability Engineer

VP, Site Reliability Engineer

Ambition Group Singapore Pte Ltd
Fresher
Not Disclosed
  • Posted 9 hours ago
  • Be among the first 10 applicants

Job Description

We are looking for an experienced and hands-on, Vice President of Site Reliability Engineering to lead the reliability, availability, and resilience strategy across critical platforms and services. This individual will establish SRE as a core engineering discipline, driving automation, observability, incident management, and reliability-by-design practices across the organization.

Key Responsibilities

  • Define and execute the enterprise SRE strategy, embedding reliability engineering across products, platforms, and services.
  • Own and govern SLIs, SLOs, and error budgets, ensuring reliability decisions are data-driven.
  • Drive improvements in service availability, resilience, recoverability, and performance.
  • Lead the strategy for observability, automation, self-healing capabilities, and resilience engineering platforms.
  • Oversee incident management, major incident response, post-incident reviews, and chaos engineering initiatives.
  • Drive operational excellence through automation, toil reduction, and platform standardization.
  • Partner with Engineering, Infrastructure, Security, and Product teams to embed reliability requirements early in the development lifecycle.
  • Build, mentor, and lead a high-performing team of SRE leaders and engineers.
  • Influence technical decisions across teams and stakeholders, driving reliability outcomes in a matrixed environment.

Required Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field (or equivalent practical experience).
  • At least 12 years of experience in Site Reliability Engineering, Software Engineering, Platform Engineering, Infrastructure, DevOps, or related disciplines.
  • Proven experience building, scaling, or transforming SRE functions within complex enterprise environments.
  • Strong expertise in:
  • SRE principles and automation-first operations
  • Observability and monitoring frameworks
  • Resilience engineering and incident management
  • Capacity planning, performance optimization, and disaster recovery
  • Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budget management
  • Solid software engineering background with experience in Java and developing or supporting large-scale distributed systems.
  • Experience designing and implementing observability, automation, self-healing, or reliability platforms.
  • Strong stakeholder management skills with the ability to influence engineering teams and senior leaders without direct authority.

More Info

Key Skills

reliability platforms

Error Budget management

automation-first operations

observability automation

observability and monitoring frameworks

SRE principles

resilience engineering

large-scale distributed systems

Similar Jobs

Singapore
Skills:
.NET, Python, Spring Boot, Splunk, Java, Cloud, Grafana, Dynatrace, Datadog, Gitlab, Prometheus, Networking, Artificial Intelligence, Android, Kubernetes, ECS, Terraform, Docker, Jenkins
1-3 yrs
Singapore
Skills:
System Administration, Java, Continuous Integration, Linux, Automation solutions, Python, UNIX, Build deployment processes
Singapore
Skills:
Java, Nginx, Golang, Rust, C, Http, Shell, Linux, MySQL, Python, K8S, Istio
Singapore
Skills:
Elasticsearch, Java, Golang, Unix, Linux, HBase, Python, Doris, RocksDB, ClickHouse, HDFS
Singapore
Skills:
powerdns , DHCP, Nat, Bind, Troubleshooting Skills, Bash, Dns, Ansible, Kerberos, Puppet, Python, APT repository management, NTP, Linux host management, blameless post-mortems, error budget management, active-passive architectures, Go, DevOps tooling, disaster recovery strategies, high availability design patterns, SLI, Salt, SLO, SRE principles