Key Responsibilities:
Senior DevOps Engineer / Lead – Infrastructure & Kubernetes Operations - Apple
About the Role
We are looking for a Senior DevOps Engineer / Lead to join our Infrastructure team, with a robust focus on Kubernetes operations, infrastructure automation, reliability, scalability, and solution architecture.
The role will own the health and reliability of Kubernetes environments, drive operational improvements, and lead the development of automation and internal tooling used by platform and service teams.
The ideal candidate should be a hands-on solution owner who can understand requirements, design scalable and reliable infrastructure solutions, evaluate trade-offs, implement solutions, and create operational runbooks — not simply follow existing procedures.
You will work closely with service owners, platform engineers, tooling teams, architects, and leadership while mentoring junior and mid-level engineers.
Key Responsibilities
- Own and drive the Kubernetes cluster operations roadmap, including cluster upgrades, node group management, scaling, draining, replacement, and overall cluster health.
- Troubleshoot complex production and non-production infrastructure issues using kubectl, logs, metrics, system diagnostics, and cloud-native tools.
- Diagnose and resolve issues related to pods, nodes, scheduling, networking, resource constraints, deployments, configuration, and cluster health.
- Understand infrastructure requirements and design appropriate technical solutions considering reliability, scalability, security, performance, and cost.
- Architect, develop, and maintain internal automation and infrastructure tooling, including scripts, CLIs, controllers, and operators.
- Drive automation of operational workflows such as cluster upgrades, remediation, scaling, maintenance, and recovery.
- Identify manual or repetitive operational processes and convert them into reliable, reusable automation with appropriate safeguards.
- Develop and improve observability strategies across logging, metrics, monitoring, alerting, and SLOs.
- Define effective incident management, issue tracking, triage, and root-cause analysis practices.
- Lead on-call rotations and incident response, including post-incident reviews and continuous improvement initiatives.
- Partner with service owners, platform teams, and leadership on capacity planning, infrastructure requirements, upgrades, and maintenance schedules.
- Create and maintain technical designs, operational runbooks, troubleshooting guides, and recovery procedures.
- Establish engineering standards through technical documentation, design/code reviews, and operational best practices.
- Mentor and guide junior and mid-level DevOps/SRE engineers and provide technical leadership on infrastructure initiatives.
- Evaluate architecture and infrastructure trade-offs across availability, scalability, security, performance, operational complexity, and cost.
Required Qualifications
- 10–15 years of experience in DevOps, SRE, Infrastructure Engineering, or a related field.
- Deep, hands-on experience with Kubernetes administration and operations in large-scale environments.
- Strong experience with:
- Kubernetes cluster upgrades
- Node group management
- Scaling, draining, and node replacement
- Pod scheduling and resource management
- Kubernetes networking and troubleshooting
- Production incident troubleshooting
- Strong programming/scripting skills in Python, Go, Bash, or similar languages.
- Proven experience building and maintaining automation and internal infrastructure tooling used across teams.
- Strong ability to analyze logs, metrics,
system state, and distributed-system behavior to identify root cause.
- Extensive experience with at least one major cloud platform: AWS, GCP, or Azure.
- Strong understanding of cloud infrastructure architecture, including IAM, networking, security, availability, scalability, and cost optimization.
- Strong understanding of distributed systems, infrastructure reliability, scalability, and performance.
- Ability to translate business/technical requirements into infrastructure architecture and implementation plans.
- Experience making infrastructure architecture and cost/performance trade-off decisions.
- Ability to independently define solutions and operational approaches rather than relying solely on predefined runbooks.
- Strong written and verbal communication skills, including technical documentation, incident reports, design documents, and status updates.
- Demonstrated experience mentoring engineers and leading technical initiatives.
Preferred Qualifications
- Strong experience with Infrastructure as Code, particularly:
- Terraform
- Helm
- Ansible
- Hands-on experience with CI/CD platforms such as Jenkins, GitHub Actions, Spinnaker, or similar.
- Experience implementing GitOps using ArgoCD, Flux, or similar tools.
- Strong experience with observability platforms such as:
- Prometheus
- Grafana
- Datadog
- CloudWatch
- Experience operating multiple Kubernetes clusters across regions/environments.
- Experience with capacity planning, disaster recovery, and high availability.
- Experience contributing to or leading infrastructure/platform architecture decisions.
- Experience designing AWS solutions with appropriate security, access control, scalability, availability, and cost considerations.
- Experience creating reusable infrastructure patterns, modules, tools, or operational frameworks used by multiple teams.
- Experience with OpenShift or other enterprise Kubernetes platforms is a plus .
📌 Senior DevOps Engineer/lead (India)
🏢 Artech
📍 India