Search by job, company or skills

Application Reliability Engineer

  • Posted 2 hours ago
  • Be among the first 10 applicants

Job Description

As an Application Reliability Engineer, you will be responsible for the health and stability of a Java/Spring-based, microservices-driven eCommerce and logistics SaaS platform. You will triage and resolve incidents end-to-end, perform root cause analysis on production issues, troubleshoot API integrations between the Client platform and its surrounding systems, and work closely with Engineering and Product teams to drive permanent fixes and process improvements.

Responsibilities:

Production Support & Incident Management

  • Provide first-line (L1) triage of incoming production issues, in addition to owning L2/L3 escalations passed up from the front-line support team.
  • Investigate, diagnose, and resolve production incidents across APIs, integrations, and backend services within agreed timeframes.
  • Perform structured root cause analysis (RCA) for production incidents, document findings, and drive issues through to permanent resolution rather than workarounds.
  • Prioritize and manage multiple concurrent incidents in line with SLA and KPI expectations.
  • Escalate appropriately to Engineering where a fix requires code-level or architectural changes, while retaining ownership of the incident lifecycle.

Production Support & Incident Management

  • Provide first-line (L1) triage of incoming production issues, in addition to owning L2/L3 escalations passed up from the front-line support team.
  • Investigate, diagnose, and resolve production incidents across APIs, integrations, and backend services within agreed timeframes.
  • Perform structured root cause analysis (RCA) for production incidents, document findings, and drive issues through to permanent resolution rather than workarounds.
  • Prioritize and manage multiple concurrent incidents in line with SLA and KPI expectations.
  • Escalate appropriately to Engineering where a fix requires code-level or architectural changes, while retaining ownership of the incident lifecycle.
  • Validate and debug integrations between client platform and external carrier, marketplace, and partner systems.

System Monitoring & Observability

  • Proactively monitor system health, error rates, and performance using observability tools such as Splunk, ELK, or Datadog.
  • Interpret logs, dashboards, and alerts to detect emerging issues before they escalate into customer-impacting incidents.
  • Contribute to refining alert thresholds, dashboards, and monitoring coverage based on recurring incident patterns.

Collaboration & Cross-Functional Work

  • Work closely with Engineering and Product teams to resolve issues, validate fixes, and feed back recurring problems for permanent remediation.
  • Coordinate with the internal Customer-facing team (not external clients directly) to provide technical context on escalated issues.
  • Communicate clearly and precisely in writing for ticketing, incident documentation, and internal handovers.

Automation & Process Improvement

  • Contribute to the development and continuous improvement of runbooks and standard operating procedures for recurring incident types.
  • Identify opportunities for automation of repetitive diagnostic or remediation tasks.
  • Support knowledge-base development to reduce mean-time-to-resolution (MTTR) across the team.

Required Technical Skills & Experience

Candidates must demonstrate genuine, hands-on depth in the following areas — not surface-level familiarity:

  • Bachelor's degree in Computer Science, Engineering, or a related field, with a strong academic record.
  • 5–6 years of experience in technical support, production support, or backend engineering — ideally within SaaS, API-driven, or platform engineering environments.
  • Strong, hands-on experience with Java, Spring / Spring Boot, and microservices-based architectures, including how distributed services communicate and fail.
  • Solid, demonstrable understanding of REST and SOAP API integrations — able to read specifications, construct and debug calls, and diagnose integration failures independently.
  • Practical, day-to-day experience using Postman or equivalent API tooling for troubleshooting.
  • Hands-on experience monitoring system health using observability/monitoring tools such as Splunk, ELK, or Datadog.
  • Solid SQL skills with MySQL and/or PostgreSQL — comfortable writing and debugging queries to investigate data-related production issues.
  • Strong proficiency working in Linux/Unix environments, including command-line diagnostics and log analysis.

Strongly Preferred – More than Nice to Have

  • Hands-on exposure to cloud platforms (AWS, Azure, or GCP) — this is a significant differentiator for this role given the platform's cloud-hosted architecture.
  • Experience with messaging/event-driven systems such as Kafka, ActiveMQ, or RabbitMQ.
  • Basic front-end knowledge (Angular, HTML, CSS) to assist in isolating front-end vs. backend issues.
  • Experience with JIRA or similar ticketing/incident management tools.
  • eCommerce, logistics, or shipping domain experience, including familiarity with carrier APIs (e.g., FedEx, UPS, Purolator, Canada Post).

Core Competencies

  • Strong analytical thinking and structured problem-solving under pressure.
  • High technical ownership — sees issues through to genuine resolution rather than passing them along.
  • Intellectual curiosity and a strong, self-driven work ethic; comfortable learning unfamiliar parts of a large codebase independently.
  • Ability to manage multiple concurrent incidents against SLA/KPI expectations without losing accuracy or thoroughness.
  • Capable of working independently with minimal supervision, while collaborating effectively with internal Engineering and Product teams.
  • Clear, precise written communication for documentation and internal handovers (this role is not customer-facing).

More Info

Job Type:
Industry:
Employment Type:

Job ID: 151789323

Beware of Scammers

We don’t charge money for job offers