21 Aug
|
VikingCloud India
|
Vadodara
21 Aug
VikingCloud India
Vadodara
Work location : Vadodara (Work from office)
Shift : US/Night Shift
Role Overview
We are hiring an Operations Engineer to own L3 operations across VMware and Google Cloud Platform (GCP). The role is responsible for designing, building, and operating production-grade cloud environments, including Kubernetes (GKE), Kafka, and infrastructure-as-code with Terraform and GitHub. You will ensure platform reliability, secure change delivery, and high-quality L3 incident/problem ownership.
Key Responsibilities
VMware - L3 Administration
- Act as L3 VMware administrator for vSphere / ESXi / vCenter environments.
- Troubleshoot complex compute, storage, networking, and HA/DRS cluster issues.
- Lead root-cause analysis for major incidents related to virtualization and on-prem compute.
- Manage capacity planning, performance tuning, patching windows, and lifecycle upgrades.
- Support VM provisioning standards, templates, and integration with backup/DR tooling.
- Mentor L1/L2 teams and define escalation playbooks for VMware operations.
Google Cloud Platform (GCP) - Build & Operate
- Create, configure, and manage GCP environments (projects, folders, org policies, IAM).
- Design and operate core GCP services: Compute Engine, VPC, Cloud DNS, Load Balancing, Cloud Storage, Cloud SQL / managed data services as required.
- Implement secure networking (Shared VPC, firewall rules, Private Google Access, Cloud NAT, VPN/Interconnect patterns).
- Own environment lifecycle: non-prod to prod promotion, landing zone hygiene, and cost/FinOps tagging.
- Configure monitoring, logging, and alerting using Cloud Monitoring / Logging (and complementary tools such as Prometheus/Grafana where used).
- Participate in incident response, post-incident reviews, and reliability improvements.
Kubernetes
- Deploy and manage GKE (or Kubernetes) clusters: node pools, upgrades, workloads, namespaces, RBAC.
- Support CI/CD deployments to Kubernetes; troubleshoot pods,
services, ingress, HPA, and networking issues.
- Implement cluster security baselines (network policies, workload identity, secret management practices).
- Maintain Helm charts / manifests and operational runbooks for platform services.
Kafka
- Install, configure, and operate Apache Kafka (self-managed and/or managed offerings such as Confluent / Managed Service for Apache Kafka as applicable).
- Manage topics, partitions, ACLs, consumer groups, retention, and performance tuning.
- Troubleshoot lag, broker health, replication, and connectivity issues across applications.
- Partner with application teams on event-driven design and production readiness.
Terraform GitHub (IaC & Delivery)
- Build and maintain Terraform modules/stacks for GCP (and related platform components).
- Enforce IaC standards: remote state, workspaces/environments, code review, plan/apply governance.
- Use GitHub for source control, pull requests, branch protection, Actions/workflows, and release process.
- Automate environment provisioning, drift detection, and repeatable infrastructure delivery.
- Collaborate with security/DevOps on policy-as-code and least-privilege service accounts.
L3 Operations & Collaboration
- Own L3 escalations across VMware GCP Kubernetes Kafka stack.
- Drive problem management, permanent fixes, and toil reduction through automation.
- Maintain accurate runbooks, architecture diagrams, and operational documentation.
- Work closely with application, security, network, and SRE/DevOps stakeholders.
- Support on-call / shift handovers as per operational model.
Required Skills & Experience
- 6-7 years overall experience in infrastructure / cloud / platform operations.
- L3-level VMware administration experience (vSphere/ESXi/vCenter) with proven major-incident ownership.
- 4-5 years hands-on experience creating and managing GCP production environments.
- Strong practical experience with Kubernetes / GKE.
- Hands-on experience operating Kafka in production (or equivalent event-streaming platforms at scale).
- Strong Terraform skills for GCP infrastructure-as-code.
- Practical experience with GitHub (repos, PRs, Actions/CI, workplace promotion).
- Solid Linux administration, networking fundamentals, and IAM/security awareness.
- Experience with monitoring/observability and incident management tools.
- Strong troubleshooting, RCA writing, and stakeholder communication skills.
Preferred / Nice to Have
- Google Cloud certifications (Associate Cloud Engineer, Professional Cloud Architect, Professional DevOps Engineer).
- VMware certification (VCP or equivalent).
- Experience with Anthos, Cloud Run, Pub/Sub, BigQuery, or service mesh (Istio/Anthos Service Mesh).
- Ansible / Python / Bash automation experience.
- Exposure to GitOps (Argo CD / Flux), Helm, and policy tools (OPA/Gatekeeper).
- Experience in captive centers / product companies / US-client delivery models.
- Familiarity with cost optimization and SRE practices (SLIs/SLOs, error budgets).
Education
Bachelor's degree in Computer Science, IT, Engineering, or equivalent practical experience.
Soft Skills
- Ownership mindset for production reliability and customer impact.
- Clear English communication for L3 updates, RCA, and cross-team coordination.
- Ability to work under incident pressure and prioritize effectively.
- Documentation discipline and continuous improvement attitude.
📌 Operations Engineer (Vadodara)
🏢 VikingCloud India
📍 Vadodara