Key Roles and Responsibilities • Monitor cloud infrastructure and respond to alerts from enterprise monitoring tools.
Perform L1 and L1.5 incident triage, troubleshooting, and remediation.
Manage incidents, requests, and operational tasks through ServiceNow.
Support GCP Compute Engine, IAM, VPC, DNS, Storage, Load Balancing, VPN and Logging services.
Participate in major incident bridges and war rooms.
Execute runbook-driven operational activities and health checks.
Support change reviews, risk assessments, pre/post implementation validations, and approved operational changes.
Maintain SOPs, runbooks, KEDB articles, and knowledge of repositories.
Provide shift handovers and operational reporting.
Identify automation, toil reduction, and service improvement prospects.
Primary
Skills • Google Cloud Platform (GCP) • Linux Administration • Basic Networking: TCP/IP, DNS, VPN, Routing, Firewalls Preferred / Optional Skills • AWS and Azure exposure • Kubernetes (GKE) • Docker • Terraform / Infrastructure as Code • GitHub, Azure DevOps, Jenkins • Ansible and workflow automation • Cloud Security and Governance