Role Overview: Technical Lead The Cloud Operations Engineer L1L15 provides operational support for enterprise cloud environments with a primary focus on Google Cloud Platform GCP The role acts as the first point of engagement for cloud s incidents operational events and service requests while ensuring adherence to ITILaligned service management practices
Experience Requirements
Minimum 8 years of experience in Cloud Operations Infrastructure Operations NOC IT Operations SRE Operations or Cloud Support environments
Key Roles and Responsibilities
Monitor cloud infrastructure and respond to s from enterprise monitoring tools
Perform L1 and L15 incident triage troubleshooting and remediation
Manage incidents requests and operational tasks through ServiceNow
Support GCP Compute Engine IAM VPC DNS Storage Load Balancing VPN and Logging services
Participate in major incident bridges and war rooms
Execute runbookdriven operational activities and health checks
Support change reviews risk assessments prepost implementation validations and approved operational changes
Maintain SOPs runbooks KEDB articles and knowledge of repositories
Provide shift handovers and operational reporting
Identify automation toil reduction and service improvement opportunities
Primary Skills
Google Cloud Platform GCP
Linux Administration
Basic Networking TCPIP DNS VPN Routing Firewalls
Preferred Optional Skills
AWS and Azure exposure
Kubernetes GKE
Docker
Terraform Infrastructure as Code
GitHub Azure DevOps Jenkins
Ansible and workflow automation
Cloud Security and Governance
Certifications Preferred
Google Associate Cloud Engineer Google Skilled Cloud Architect
Soft Skills
Strong troubleshooting communication stakeholder management documentation customer focus and ability to work in 24x7 support environments
Technical Lead Responsibilities
Lead and coordinate the Cloud Operations team across shifts to ensure 24x7 operational coverage and service continuity
Act as the technical escalation point for complex incidents service disruptions and critical operational issues
Drive Major Incident Management activities coordinate with crossfunctional teams and provide technical leadership during war rooms and bridge calls
Mentor and guide L1L15 engineers on cloud operations troubleshooting methodologies and operational best practices
Review incident trends recurring issues and problem records drive root cause analysis and permanent fixes
Own operational governance including SLA adherence service quality shift readiness and operational maturity improvements
Lead change reviews validate implementation plans and ensure operational readiness for production deployments
Identify automation opportunities and drive implementation of operational efficiencies toil reduction and selfhealing capabilities
Collaborate with Cloud Engineering Security Network and Application teams to improve platform reliability and performance
Review and approve runbooks SOPs KEDB articles and operational documentation to ensure accuracy and effectiveness
Provide leadership reporting on service health incident metrics risk areas and improvement initiatives
Support capacity planning availability management operational risk assessments and cloud governance activities
Shifts will be rotated on below timings
06:30 AM 4 PM IST
2:30 PM 12:00 AM IST
Ready to take oncall on weekends based on oncall scheduling