What you’ll do
• Own the cloud infrastructure and Kubernetes platform that hosts services end-to-end — provision, upgrade, and operate multi-cloud (AWS + Azure) clusters.
• Build and maintain infrastructure-as-code — Helm charts, ARM/CloudFormation templates, cloud-init, and Python automation for cluster creation, CNI, node scaling, machine health checks, and workload identity.
• Bake and harden machine images using Packer + Ansible for Linux and Windows, validate them with automated image tests (goss), and manage the container image supply chain.
• Run observability and alerting — Prometheus/Grafana, Splunk (OTel collectors, saved searches/alerts)
• Lead production support and on-call — triage incidents, perform RCA, remediate issues and manage node/pod health, capacity, and autoscaling.
• Design and operate CI/CD and release pipelines using Jenkins and deployment pipelines.
• Deploy, monitor, and troubleshoot platform services — Spring Boot services, node agents, and containerised engine workloads — across Kubernetes and clouds.
• Optimize scalability, cost,
and security — right-size clusters, tune autoscaling/workload placement, enforce network policies and admission control.
• Python/Shell scripting to develop and maintain platform code.
• Participate in code/infra reviews, cross-functional collaboration, and mentor junior engineers.
What you need to succeed
• 6–8 years in software/platform engineering with strong DevOps / SRE / infrastructure focus and end-to-end production ownership.
• Deep hands-on Kubernetes experience — Cluster API / kubeadm, Helm, networking/CNI (Calico), autoscaling, and Linux/Windows node/container workloads.
• Solid hands-on AWS and/or Azure experience across compute, storage, networking, identity/IAM, and managed services; multi-cloud is a strong plus.
• Infrastructure-as-code and automation proficiency — Helm, ARM/CloudFormation/Terraform, Packer + Ansible + Python
• Proven production-operations experience — monitoring, logging, ale
📌 Computer Scientist 2 (Noida)
🏢 Adobe
📍 Noida