- Key Responsibilities
- Build and operate reliable scalable services on AWS with a strong focus on uptime latency and efficiency
- Define and track SLI SLOs error budgets and reliability metrics drive improvements based on data and trends
- Implement automation to reduce manual toil across deployments provisioning and operational workflows
- Develop and maintain CI CD pipelines and DevOps practices to enable secure repeatable releases
- Set up and enhance observability monitoring logging alerting dashboards for proactive issue detection
- Lead participate in incident response perform root cause analysis and deliver post incident action plans
- Improve system resilience through capacity planning performance tuning and fault tolerant architecture patterns
- Collaborate with engineering teams to embed reliability into design reviews releases and operational readiness
- Minimum Qualifications
- 3 5 years of experience in AWS SRE DevOps or production engineering roles supporting cloud based systems
- Bachelor s degree in BTECH BCA MCA MTECH or equivalent practical experience
- Strong hands on experience with AWS services and operating workloads in production environments
- Solid understanding of SRE principles including incident management on call practices and reliability engineering
- Experience implementing DevOps practices such as CI CD infrastructure automation and release reliability
- Minimum Qualifications
- 3 5 years of experience in AWS SRE DevOps or production engineering roles supporting cloud based systems
- Bachelor s degree in BTECH BCA MCA MTECH or equivalent practical experience
- Strong hands on experience with AWS services and operating workloads in production environments
- Solid understanding of SRE principles including incident management on call practices and reliability engineering
- Experience implementing DevOps practices such as CI CD infrastructure automation and release reliability
- Preferred Qualifications
- Experience with Infrastructure as Code and automated provisioning for AWS environments e
- g
- Terraform CloudFormation
- Strong containerization and orchestration exposure e
- g
- Docker Kubernetes EKS for scalable deployments
- Proven ability to build effective observability using tools like CloudWatch Prometheus Grafana ELK OpenSearch
- Experience hardening systems with security best practices IAM least privilege secrets management patching
- Track record of reducing MTTR and improving reliability through runbooks automation and operational playbooks
- Familiarity with microservices reliability patterns load balancing autoscaling and high availability design on AWS