(Senior) Cloud Operations Manager - Roles and Responsibilities
The Senior Cloud Operations Manager is responsible for overseeing the daily operations and ensuring optimal performance of cloud environments. This role focuses on delivering exceptional service quality, managing critical incidents, and driving continuous improvement in cloud operations.
Key responsibilities include:
Ticket Validation and SLA Management
- Ensure that all incoming tickets are accurately validated, categorized, and assigned to the appropriate teams for resolution.
- Maintain oversight of Service Level Agreements (SLA) to ensure timely resolution of incidents and requests, adhering to the agreed-upon response and resolution times.
Leadership and Team Development
- Lead, mentor, and guide the Cloud Operations team towards the advancement of the Capability Maturity Model Integration (CMMI) framework, ensuring continuous improvement in processes and performance.
- Foster a culture of high performance, accountability, and professional growth within the team.
Incident Commander
- Act as the Incident Commander during major incidents or service outages, ensuring a coordinated and efficient response to restore services.
- Provide clear communication to all stakeholders, manage escalations, and ensure all incident details are captured for post-incident reviews.
Timezone Escalation Point
- Serve as the primary escalation point for operational issues and incidents within the designated timezone, ensuring swift and effective resolutions for any critical concerns.
- Provide guidance and oversight to on-call teams during off-hours and holidays as necessary.
Knowledgebase Maintenance
- Maintain and enhance the internal knowledgebase by documenting standard operating procedures (SOPs), troubleshooting guides, and incident resolution steps.
- Regularly review and update knowledgebase content to ensure it is accurate, comprehensive, and aligned with operational best practices.
Standard Operating Procedure (SOP) Management
- Develop, maintain, and periodically review SOPs for cloud operations, ensuring all operational processes are standardized and follow industry best practices.
- Ensure that all team members are properly trained on SOPs and are following them in day-to-day operations.
Regular Drills and Preparedness
- Conduct regular operational drills (e.g., disaster recovery, incident response) to ensure the team is well-prepared to handle various operational scenarios.
- Identify potential gaps or weaknesses during drills and implement corrective actions to improve team readiness.
Feedback Loop and Continuous Improvement
- Establish a robust feedback loop into the quarterly planning process, contributing at least three (3) improvement suggestions from operational experiences, incidents, and team feedback.
- Leverage data and metrics to recommend process optimizations and enhance operational efficiency.
Monitoring and Anomaly Detection
- Continuously monitor cloud environments and services to identify anomalies, trends, or potential risks that could impact service performance or availability.
- Collaborate with the engineering and development teams to resolve issues proactively and mitigate risks before they escalate.
Cross-Functional Collaboration
● Collaborate with other departments (e.g., Engineering, Security, Compliance) to ensure seamless operations, alignment with organizational goals, and adherence to security and compliance standards.
● Support cloud initiatives and contribute to projects that improve the overall cloud infrastructure.