Technical Lead The Cloud Operations Engineer L1L15 provides operational support for enterprise cloud environments with a primary focus on Google Cloud Platform GCP The role acts as the first point of engagement for cloud s incidents operational events and service requests while ensuring adherence to ITILaligned service management practices
Experience Requirements
Minimum 8 years of experience in Cloud Operations Infrastructure Operations NOC IT Operations SRE Operations or Cloud Support environments
Key Roles and Responsibilities
Monitor cloud infrastructure and respond to s from enterprise monitoring tools Perform L1 and L15 incident triage troubleshooting and remediation Manage incidents requests and operational tasks through ServiceNow Support GCP
Compute Engine IAM VPC DNS Storage Load Balancing VPN and Logging services Participate in major incident bridges and war rooms Execute runbookdriven operational activities and health checks Support change reviews risk assessments prepost implementation validations and approved operational changes
Maintain SOPs runbooks KEDB articles and knowledge of repositories Provide shift handovers and operational reporting Identify automation toil reduction and service improvement opportunities
Primary Skills
Google Cloud Platform GCP Linux Administration Basic Networking TCPIP DNS VPN Routing Firewalls Preferred Optional Skills AWS and Azure exposure Kubernetes GKE Docker Terraform Infrastructure as Code GitHub Azure DevOps Jenkins Ansible and workflow automation Cloud Security and Governance Certifications Preferred Google Associate Cloud Engineer Google Skilled Cloud Architect
Soft Skills
- Strong troubleshooting communication stakeholder management documentation customer focus and ability to work in 24x7 support environments
Technical Lead Responsibilities
- Lead and coordinate the Cloud Operations team across shifts to ensure 24x7 operational coverage and service continuity
- Act as the technical escalation point for complex incidents service disruptions and critical operational issues
- Drive Major Incident Management activities coordinate with crossfunctional teams and provide technical leadership during war rooms and bridge calls
- Mentor and guide L1L15 engineers on cloud operations troubleshooting methodologies and operational best practices
- Review incident trends recurring issues and problem records drive root cause analysis and permanent fixes
- Own operational governance including SLA adherence service quality shift readiness and operational maturity improvements
- Lead change reviews validate implementation plans and ensure operational readiness for production deployments
- Identify automation opportunities and drive implementation of operational efficiencies toil reduction and selfhealing capabilities
- Collaborate with Cloud Engineering Security Network and Application teams to improve platform reliability and performance
- Review and approve runbooks SOPs KEDB articles and operational documentation to ensure accuracy and effectiveness
- Provide leadership reporting on service health incident metrics risk areas and improvement initiatives
- Support capacity planning availability management operational risk assessments and cloud governance activities
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.