We are seeking a skilled and proactive Apollo L1 Support Engineer to join our AI Platform Operations team. In this role, you will serve as the first line of technical support for our AI/LLM platform, providing expert-level troubleshooting and incident resolution for complex technical issues involving APIs, networking, and Python-based applications.
This is a highly technical support role requiring strong developer skills and a deep understanding of IT infrastructure. You will work closely with development teams, platform engineers, and end-users to ensure the reliability and performance of our AI platform.
Key Responsibilities
Technical Support & Incident Management
- Incident Handling:Receive, triage, and resolve Level 1 technical incidents related to the Apollo AI platform, ensuring timely resolution and minimal service disruption.
- API Troubleshooting:Diagnose and resolve API connectivity, authentication, and performance issues using tools like Postman, curl, and logging platforms.
- Networking Support:Troubleshoot network-related issues including connectivity, latency, DNS, and firewall configurations affecting platform access.
- Python Application Support:Debug and resolve issues with Python-based automation scripts, data pipelines, and integration workflows.
System Monitoring & Operations
- Proactive Monitoring:Monitor platform health and performance using observability tools, identifying potential issues before they impact users.
- Alert Response:Respond to system alerts, perform initial diagnostics, and escalate complex issues to Level 2/3 engineers as needed.
- Runbook Execution:Follow documented runbooks and standard operating procedures for incident resolution and system maintenance tasks.
Documentation & Knowledge Management
- Knowledge Base:Create and maintain detailed documentation, knowledge articles, and troubleshooting guides to support end-users and internal teams.
- Incident Reports:Document incident root causes, resolution steps, and preventive measures to build a comprehensive knowledge repository.
- Continuous Improvement:Contribute to the improvement of support processes, runbooks, and automation scripts.
Collaboration & Communication
- Cross-functional Collaboration:Work closely with development, platform engineering, and product teams to resolve complex technical issues and communicate platform updates.
- Stakeholder Communication:Provide clear and professional updates to stakeholders on incident status, resolution timelines, and root cause analysis.
- Knowledge Sharing:Actively participate in knowledge transfer sessions, team stand-ups, and post-incident reviews.
Escalation & Incident Management
- Escalation:Escalate complex or unresolved issues to Level 2/3 engineers with clear documentation and diagnostic information.
- Triage & Prioritization:Prioritize incidents based on business impact and service level agreements (SLAs), ensuring critical issues are addressed immediately.
- Incident Documentation:Ensure accurate and detailed logging of all incidents and service requests in the ticketing system.