The Site Reliability Engineer responsible for maintaining the service-level agreement critical production platforms or products and providing automated operations to ensure the service to our clients is always of the best quality.
Responsibilities:
- Harden platforms before and after go-live by reviewing architecture, security, configurations, and implementing monitoring.
- Ensure reliability by monitoring availability, performance, and overall system health across infrastructure and applications.
- Lead incident response and recovery to meet SLAs, leveraging strong infrastructure expertise to restore services quickly.
- Conduct root cause analysis and post-mortems to drive continuous reliability improvements.
- Collaborate with development and product teams to enhance scalability, resilience, and operational readiness.
- Validate release readiness through CAB participation, automated testing, and infrastructure-aware risk assessment.
- Design and improve observability via dashboards, alerts, and monitoring tools (e.g., Azure Monitor, Prometheus, Grafana).
- Manage and optimize core infrastructure (Windows servers, AD, IIS, DNS, databases) supporting critical applications.
- Build and maintain CI/CD pipelines and standardized deployment processes (Azure DevOps, GitHub Actions).
- Support cloud and modernization initiatives (Azure, containers, Kubernetes, IaC) to enable scalable application reliability.
Qualifications:
- 2 years of experience in a Site Reliability Engineering, Systems Engineering, or Infrastructure Engineering role.
- Willing to work in a rotating schedule
- Strong experience administering Windows Server environments, including Active Directory, IIS, and file systems.
- Hands-on experience implementing CI/CD pipelines and automation using tools such as Azure DevOps.
- Foundational understanding of cloud architecture, preferably within Microsoft Azure environments.
- Familiarity with containerization and orchestration technologies (e.g., Docker, Kubernetes).
- Experience monitoring and optimizing SQL database performance.
- Knowledge of SQL database deployment strategies and administration practices.
- Exposure to monitoring, logging, and alerting frameworks, as well as participation in on-call operational support models.
Nice to Haves:
- Experience within payment services, card issuance platforms, or banking technology environments.
- Relevant certifications such as AZ-104, AZ-400, or other DevOps/SRE-related credentials.
- Familiarity with compliance and security frameworks such as PCI-DSS, PCI-CPP, or ISO 27001.
- Proficiency in PowerShell scripting and infrastructure automation.