Big Data Platform SRE Engineer
Job Description
What You'll Do
- Support the day-to-day operation and maintenance of Tencent's overseas big data platforms and clusters under the guidance of senior engineers; Help monitor platform health and escalate anomalies.
- Act as a first point of contact for user issues — platform usage, component errors, and basic performance problems — and work with senior engineers to drive them to resolution.
- Participate in reviewing existing operational workflows and standards; identify pain points and contribute improvement ideas.
- Assist in building and maintaining internal operational efficiency tools: capacity management, DevOps pipelines, coverage analytics, automation scripts, and health monitoring.
- Contribute to cluster cost-optimization initiatives, e.g. helping evaluate new technologies or architectures and measuring their impact on processing efficiency and cost.
- Document incidents, resolutions and runbooks to build up the team's shared knowledge base.
What We Look For
Must-have
- Some hands-on experience with big data technologies.
- Working knowledge of at least one of the following: Hadoop, YARN, Flink, Spark, HDFS, Presto, Hive — plus a genuine willingness to go deep on the rest.
- Basic proficiency in Shell and at least one of the following languages: Java / Python / Go.
- Solid Linux/Unix fundamentals: Comfortable on the command line, able to do basic performance observation and log-based troubleshooting.
- A troubleshooting mindset — you can break a problem down, search, ask, and iterate rather than getting stuck.
Nice to have
- Exposure to mainstream big data or cloud data platforms (e.g. Databricks, Snowflake, EMR).
- Coursework or projects in distributed systems, databases, or cloud infrastructure.
- Familiarity with observability tooling (Prometheus, Grafana, ELK) or orchestration/automation (Airflow, Ansible, Kubernetes).
- Any exposure to on-call or production incident handling.
