Location: Manila, Phillipines
Experience: 1-3 years
Type: Full Time Consultant
Department: Engineering
About Us:
Qure.ai is a global healthtech company using artificial intelligence to improve disease detection and patient outcomes at scale. Our solutions are deployed across hospitals, diagnostic networks, and public health programs worldwide.
Today, our technology has impacted 40M+ lives across 105+ countries, deployed across 5,200+ sites, and powered by one of the world's largest real-world medical imaging datasets. With 26 FDA-cleared indications, we are helping shape the future of AI-driven healthcare.
The Opportunity
Qure.ai is seeking a talented and experienced Site Reliability Engineer to join our team. As a Site Reliability Engineer, you will play a crucial role in ensuring the reliability, scalability, and performance of our software systems. You will collaborate with cross-functional teams to build and maintain a highly available and efficient infrastructure that supports our artificial intelligence-based healthcare solutions. This role is based in the Philippines and requires frequent travel within the Philippines and across Southeast Asian countries based on business requirements. Additionally, you will regularly engage with client teams on-site, necessitating strong communication and soft skills.
What You'll Own &
- Drive, Deploy and distribute Qure.ai's AI and computer vision software across customer environments, supporting both Linux and Windows platforms following established deployment procedures.
- Manage and maintain on-premise server infrastructure at hospital and clinic sites covering provisioning, OS patching, configuration hardening, and capacity planning for GPU and compute-intensive AI workloads.
- Work directly with customers to diagnose and resolve recurring issues in production deployments and updates, applying structured debugging across OS, network, and application layers.
- Support the implementation and upkeep of highly reliable services using infrastructure-as-code tooling (Ansible, Terraform, or equivalent) to enable consistent and swift product delivery across customer sites.
- Monitor and analyse system performance across cloud and on-premise environments, identify bottlenecks, and implement solutions to optimise throughput and ensure high availability.
- Maintain monitoring, alerting, and logging systems (Prometheus, Grafana, ELK, or equivalent) to detect and respond to issues proactively, covering both cloud and on-premise deployments.
- Participate in incident response and post-mortem processes supporting root cause analysis and implementing fixes to prevent recurrence.
- Collaborate with security teams to maintain infrastructure hardening standards, manage access controls, and ensure compliance with healthcare data protection requirements.
- Translate technical findings into clear communication for non-technical hospital stakeholders, requiring strong interpersonal and written skills.
- Work within a culture that champions deep technical ownership, operating closely alongside Dev, Product, and Customer Success teams to ensure reliability is a shared responsibility across the organisation.
- Willingness to travel domestically and internationally as required to support customer deployments and on-site engagements.
Our Values
- Humble: Curious, open to feedback, and focused on team success over individual credit.
- Hungry: High ownership, relentless drive, and a constant push for better outcomes.
- Smart: Structured thinking with empathy, solving complex problems with clarity and strong collaboration.