Search by job, company or skills

  • Posted a day ago
  • Be among the first 10 applicants

Job Description

Details:

Job Description

We are seeking a skilled and proactive Apollo L1 Support Engineer to join our AI Platform Operations team. In this role, you will serve as the first line of technical support for our AI/LLM platform, providing expert-level troubleshooting and incident resolution for complex technical issues involving APIs, networking, and Python-based applications.

This is a highly technical support role requiring strong developer skills and a deep understanding of IT infrastructure. You will work closely with development teams, platform engineers, and end-users to ensure the reliability and performance of our AI platform.

Key Responsibilities

Technical Support & Incident Management

  • Incident Handling: Receive, triage, and resolve Level 1 technical incidents related to the Apollo AI platform, ensuring timely resolution and minimal service disruption.
  • API Troubleshooting: Diagnose and resolve API connectivity, authentication, and performance issues using tools like Postman, curl, and logging platforms.
  • Networking Support: Troubleshoot network-related issues including connectivity, latency, DNS, and firewall configurations affecting platform access.
  • Python Application Support: Debug and resolve issues with Python-based automation scripts, data pipelines, and integration workflows.

System Monitoring & Operations

  • Proactive Monitoring: Monitor platform health and performance using observability tools, identifying potential issues before they impact users.
  • Alert Response: Respond to system alerts, perform initial diagnostics, and escalate complex issues to Level 2/3 engineers as needed.
  • Runbook Execution: Follow documented runbooks and standard operating procedures for incident resolution and system maintenance tasks.

Documentation & Knowledge Management

  • Knowledge Base: Create and maintain detailed documentation, knowledge articles, and troubleshooting guides to support end-users and internal teams.
  • Incident Reports: Document incident root causes, resolution steps, and preventive measures to build a comprehensive knowledge repository.
  • Continuous Improvement: Contribute to the improvement of support processes, runbooks, and automation scripts.

Collaboration & Communication

  • Cross-functional Collaboration: Work closely with development, platform engineering, and product teams to resolve complex technical issues and communicate platform updates.
  • Stakeholder Communication: Provide clear and professional updates to stakeholders on incident status, resolution timelines, and root cause analysis.
  • Knowledge Sharing: Actively participate in knowledge transfer sessions, team stand-ups, and post-incident reviews.

Escalation & Incident Management

  • Escalation: Escalate complex or unresolved issues to Level 2/3 engineers with clear documentation and diagnostic information.
  • Triage & Prioritization: Prioritize incidents based on business impact and service level agreements (SLAs), ensuring critical issues are addressed immediately.
  • Incident Documentation: Ensure accurate and detailed logging of all incidents and service requests in the ticketing system.

Job Requirements

Details:

What do you need to succeed

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
  • 3+ years of experience in a technical support, software development, or systems engineering role.
  • Proven experience troubleshooting complex technical issues in a production environment.
  • Knowledge of Japanese (spoken and written) will be a plus.

Technical Skills (Must-have):

  • Strong proficiency in Python - ability to read, debug, and write automation scripts
  • Deep understanding of RESTful APIs, authentication (OAuth, API keys), and troubleshooting tools (Postman, curl)
  • Solid understanding of networking fundamentals (TCP/IP, DNS, firewalls, load balancers, VPNs)
  • Familiarity with Linux command line for log analysis and system troubleshooting
  • Understanding of CI/CD pipelines and deployment processes
  • Experience with monitoring and observability tools
  • Experience with ServiceNow or similar ITSM platforms

Nice To Have:

  • Knowledge of Large Language Models, prompting, model inference, and AI platform operations
  • Experience with containerization and orchestration platforms
  • Familiarity with AWS, Azure, or GCP
  • Understanding of ITIL processes (incident, problem, change management)
  • Familiarity with authentication, authorization, and security best practices
  • Proficiency in Bash or Shell scripting for automation

Soft Skills:

  • Excellent Communication: Clear and professional verbal and written communication in English.
  • Structured Mindset: Highly organized with strong attention to detail and ability to prioritize effectively.
  • Problem-Solving: Strong analytical and troubleshooting skills with a proactive approach to issue resolution.
  • Knowledge Sharing: Willingness to share knowledge and contribute to team development.
  • Customer Focus: Strong customer service orientation and commitment to user satisfaction.
  • Flexibility: Willingness to work across morning, mid, and night shifts as required.
  • Attendance and schedule adherence are requirements of this position

,

More Info

Job Type:
Industry:
Employment Type:

Job ID: 153587351

Beware of Scammers

We don’t charge money for job offers