19 Sep
|
Rackspace US
|
India
19 Sep
Rackspace US
India
Role Overview
The L3 Dev Ops / Cloud Engineer will act as a senior technical resource responsible for leading complex infrastructure and Dev Ops activities, driving automation, defining technical standards, and improving platform reliability.
The role requires solid hands-on expertise in AWS, Terraform, AWX/Ansible, ArgoCD, Git Ops, Kubernetes, CI/CD, and Disaster Recovery.
Key Responsibilities
- Act as the L3 escalation point for complex Cloud, Dev Ops, Kubernetes, infrastructure, and deployment issues.
- Lead major technical activities across production and non-production environments.
- Take ownership of complex infrastructure changes, upgrades, migrations, and platform improvements.
- Lead and coordinate Disaster Recovery (DR) activities, including:
- DR planning
- Restore testing
- Failover/failback activities
- Application and infrastructure recovery
- DR validation
- Runbook preparation and improvement
- Design and implement infrastructure using Terraform.
- Develop and maintain reusable Terraform modules.
- Review Terraform code and infrastructure changes implemented by L1/L2 engineers.
- Define Terraform coding standards, repository structure, and implementation best practices.
- Strong hands-on experience with AWX / Ansible.
- Develop reusable Ansible roles, playbooks, and automation workflows.
- Use AWX to automate:
- Server configuration
- Patching
- Application deployment
- Infrastructure operations
- Repetitive BAU activities
- Identify manual operational activities and convert them into automated workflows.
- Design and maintain ArgoCD-based deployments.
- Implement and support Git Ops practices for Kubernetes and application deployments.
- Define Git Ops repository structures and deployment standards.
- Manage and troubleshoot ArgoCD:
- Applications
- Sync issues
- Configuration drift
- Deployment failures
- Environment promotion
- Rollback activities
- Design and maintain Kubernetes workloads running on Amazon EKS.
- Develop and maintain reusable Helm Charts.
- Define standards for Helm values, templates, and environment-specific configurations.
- Design and improve CI/CD deployment processes.
- Support and improve Jenkins and other CI/CD automation.
- Define deployment strategies and operational standards.
- Lead automation initiatives across Cloud and Dev Ops platforms.
- Develop automation using:
- Terraform
- Ansible / AWX
- Bash
- Python
- CI/CD pipelines
- Define and enforce technical standards for:
- Infrastructure as Code
- Terraform
- Git Ops
- Kubernetes
- Helm
- CI/CD
- Automation
- Cloud operations
- Patching
- Deployment processes
- Establish reusable templates, modules, pipelines, and automation frameworks.
- Perform technical reviews for changes implemented by L1 and L2 engineers.
- Lead complex production incidents and perform Root Cause Analysis.
- Identify recurring issues and implement permanent automated solutions.
- Improve monitoring, logging, alerting, and operational reliability.
- Participate in architecture and technical design discussions.
- Lead production deployments, infrastructure upgrades, patching, and maintenance activities.
- Mentor and provide technical guidance to L1 and L2 engineers.
- Create and maintain:
- SOPs
- Runbooks
- Technical standards
- Architecture documentation
- DR documentation
- Operational procedures
Experience
8–12 years of relevant experience in cloud engineering, infrastructure, Dev Ops, or a related field is required.
Core Technical SkillsStrong hands-on expertise in:
- AWS
- Amazon EKS
- Kubernetes
- Terraform
- Terragrunt
- AWX
- Ansible
- ArgoCD
- Git Ops
- Helm
- Docker
- Jenkins
- CI/CD
- Git
- Linux
- Windows
- Bash / Python
- SSL/TLS
- Infrastructure Automation
- Disaster Recovery
- Monitoring & Logging
- Production Troubleshooting
- Root Cause Analysis
L3 ExpectationsAn L3 Engineer should be capable of:
- Leading technical activities independently
- Owning complex Cloud and Dev Ops changes
- Designing and implementing automation
- Leading Disaster Recovery activities
- Implementing Git Ops using ArgoCD
- Building automation using AWX/Ansible
- Developing and reviewing Terraform code
- Defining technical standards and best practices
- Driving operational improvements
- Mentoring L1/L2 engineers
- Handling complex production incidents and RCA
The L3 engineer should move beyond routine operational support and focus on engineering, automation, standardization, reliability, and technical leadership.
About Rackspace Technology
We are the multicloud solutions experts. We combine our expertise with the world’s leading technologies — across applications, data and security — to deliver end-to-end solutions. We have a proven record of advising customers based on their business challenges, designing solutions that scale, building and managing those solutions, and optimizing returns into the future. Named a best place to work, year after year according to Fortune, Forbes and Glassdoor, we attract and develop world-class talent. Join us on our mission to embrace technology, empower customers and deliver the future.
More on Rackspace Technology
Though we’re all different, Rackers thrive through our connection to a central goal: to be a valued member of a winning team on an inspiring mission. We bring our whole selves to work every day. And we embrace the notion that unique perspectives fuel innovation and enable us to best serve our customers and communities around the globe. We welcome you to apply today and want you to know that we are committed to offering equal employment opportunity without regard to age, color, disability, gender reassignment or identity or expression, genetic information, marital or civil partner status, pregnancy or maternity status, military or veteran status, nationality, ethnic or national origin, race, religion or belief, sexual orientation, or any legally protected characteristic. If you have a disability or special need that requires accommodation, please let us know.
Originally posted on Himalayas
📌 AWS Cloud Engineer III (India)
🏢 Rackspace US
📍 India