- Key Responsibilities
- Manage and maintain highly available and scalable AWS cloud infrastructure
- Monitor applications servers and cloud resources to ensure system reliability and performance
- Automate operational tasks using Python Shell scripting and Infrastructure as Code IaC
- Troubleshoot production issues and perform root cause analysis RCA
- Implement and support CI CD pipelines for seamless application deployments
- Manage incident problem and change management activities
- Configure and maintain monitoring and alerting tools
- Work closely with development and DevOps teams to improve system stability
- Ensure security compliance backup and disaster recovery standards are met
- Participate in on call support and production support activities
- Minimum Qualifications
- 5 to 9 years of experience in AWS Cloud and SRE Production Support
- Robust knowledge of AWS services such as EC2 S3 RDS IAM VPC and Route 53
- Experience with Linux Unix administration
- Proficiency in Python or Shell Scripting
- Hands on experience with monitoring tools such as CloudWatch Grafana or Prometheus
- Experience with incident management and troubleshooting production environments
- Knowledge of networking concepts including DNS Load Balancers TCP IP and VPN
- Experience with Git and CI CD tools
- Preferred Qualifications
- Experience with Docker and Kubernetes
- Experience with Terraform or CloudFormation
- Knowledge of Jenkins and DevOps practices
- Experience with ELK Splunk Datadog or New Relic
- Familiarity with ServiceNow and ITIL processes
- Experience in High Availability HA and Disaster Recovery DR environments
- AWS Certification is an added advantage
- Experience working in Agile Scrum environments