27 Aug
|
Prowess Publishing
|
Hyderabad
27 Aug
Prowess Publishing
Hyderabad
Site Reliability Engineer (SRE) AWS DevOPS
Experience: 34 Years
Location: Hyderabad
Work Mode: Work from Office
Job Summary
We are looking for a hands-on Site Reliability Engineer (SRE) / AWS Cloud Operations Engineer with 3–4 years of experience to take ownership of production infrastructure, cloud operations, incident management and reliability.
The ideal candidate should be comfortable working in a fast-paced environment, troubleshooting production issues independently, monitoring critical systems and implementing automation to improve system availability and operational efficiency.
This is not a generic DevOps role. Solid hands-on experience in AWS, Linux, production support, troubleshooting and incident management is essential.
Key Responsibilities
- Own and support production AWS infrastructure and cloud environments.
- Monitor applications, infrastructure and services and proactively identify performance or availability issues.
- Troubleshoot and resolve production incidents within defined SLAs.
- Perform root cause analysis (RCA) for recurring production issues and implement permanent fixes.
- Manage AWS services including EC2, ECS/EKS, S3, IAM, VPC, ELB/ALB, CloudWatch and related services.
- Work with Linux environments, system processes, logs, networking and performance troubleshooting.
- Support and troubleshoot Kubernetes/Docker based workloads.
- Implement infrastructure automation using Terraform / Infrastructure as Code.
- Build and maintain monitoring, alerting and observability using CloudWatch, Prometheus, Grafana, Datadog, Splunk or equivalent tools.
- Automate repetitive operational activities using Python, Bash or Shell scripting.
- Support and maintain CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI/CD or similar tools.
- Participate in incident response,
problem management and post-incident reviews.
- Collaborate with Development, QA and other engineering teams to improve application reliability and deployment processes.
- Identify opportunities for automation, capacity improvement, performance optimization and reduction of operational toil.
- Participate in on-call / production support activities as required.
Mandatory Skills
- 3–4 years of relevant experience in SRE, AWS DevOps, Cloud Operations or Production Engineering.
- Strong hands-on AWS Cloud experience.
- Strong Production Support / Production Operations experience.
- Excellent Linux troubleshooting skills.
- Hands-on experience with Kubernetes / EKS and Docker.
- Good understanding of AWS networking – VPC, subnets, security groups, load balancers, routing, DNS, etc.
- Experience with monitoring and alerting tools.
- Strong troubleshooting, incident management and RCA skills.
- Experience with Bash/Shell or Python scripting.
- Good understanding of CI/CD concepts and tools.
- Ability to independently investigate and resolve production issues.
Good to Have
- Terraform / CloudFormation
- Prometheus / Grafana / Datadog
- AWS ECS / Fargate
- Jenkins / GitHub Actions / GitLab
- ELK / OpenSearch / Splunk
- Infrastructure as Code
- Auto Scaling and high-availability architecture
- Performance and capacity monitoring
- ITIL / Incident & Problem Management
- Experience supporting microservices environments
What We Are Looking For We are particularly interested in candidates who demonstrate:
Production Ownership | Strong AWS | Linux Troubleshooting | Incident Management | RCA | Kubernetes | Monitoring | Automation
Candidates who have primarily worked on CI/CD pipeline development without significant production-support ownership may not be suitable for this position.
📌 SRE - AWS DevOPS Engineer (Hyderabad)
🏢 Prowess Publishing
📍 Hyderabad