14 Aug
|
Zensar
|
Hyderabad
Zensar Technologies is looking for an experienced AWS Site Reliability Engineer (SRE) to join our Cloud Engineering team. The ideal candidate will be responsible for ensuring the reliability, availability, scalability, and performance of mission-critical enterprise applications running on AWS cloud platforms.
The role demands deep expertise in production operations, incident management, cloud-native architectures, observability, automation, and Kubernetes-based microservices environments. The candidate will work closely with development, infrastructure, and platform teams to improve system reliability, automate operations, and drive operational excellence.
Key Responsibilities
- Manage and support large-scale production environments hosted on AWS Cloud.
- Ensure high availability, reliability, scalability, and performance of business-critical applications.
- Drive Site Reliability Engineering practices across cloud-native and microservices-based applications.
- Perform incident management, troubleshooting, root cause analysis (RCA), and postmortem documentation.
- Participate in on-call support rotations and provide timely resolution for production issues.
- Build and enhance observability solutions using monitoring, logging, and APM tools.
- Automate infrastructure provisioning and operational processes using Infrastructure as Code (IaC).
- Design and implement CI/CD pipelines for application deployment and platform automation.
- Collaborate with development teams to improve resiliency, reliability, and operational efficiency.
- Support cloud migration initiatives and disaster recovery strategies.
- Build monitoring and alerting solutions for Kubernetes infrastructure and backend microservices.
- Contribute towards performance tuning, capacity planning, and reliability improvements.
- Implement security best practices for cloud-native applications and infrastructure.
Must Have Skills Cloud & Infrastructure
- Strong hands-on experience with AWS Cloud services.
- Experience in AWS Networking, Compute, Storage, IAM, Security, and Cloud Operations.
- Experience supporting enterprise-scale production environments.
Site Reliability Engineering
- Minimum 2.5+ years of dedicated SRE experience.
- Strong experience in Incident Management, Change Management, Problem Management, and Production Support.
- Expertise in RCA (Root Cause Analysis) and Postmortem Analysis.
- Experience working in on-call and shift-based support environments.
Containerization & Microservices
- Strong experience with Kubernetes and Docker.
- Experience managing and supporting Microservices-based architectures.
- Experience with Kubernetes platform monitoring and troubleshooting.
Observability & Monitoring
- Experience with APM tools such as:
- Splunk APM (Preferred)
- Datadog
- Dynatrace
- New Relic
- AppDynamics
- Experience with Monitoring & Visualization tools:
- Grafana
- Prometheus
- Experience with Logging platforms
- Splunk
- ELK Stack
- CloudWatch
- Sumo Logic
DevOps & Automation
- CI/CD pipeline implementation using Jenkins and AWS CodePipeline.
- Experience with GitHub and GitLab.
- Strong automation mindset and operational excellence practices.
Scripting & Programming
- Strong scripting skills in Bash/Shell and Python.
- Programming exposure to Java.
- Ability to develop automation utilities and operational tools.
Operating Systems
- Strong Linux Administration and troubleshooting experience.
Good to Have
- Terraform and Infrastructure as Code implementation.
- Helm Charts and Kubernetes deployment automation.
- Service Mesh implementation experience.
- Enterprise cloud migration projects.
- Disaster Recovery and Business Continuity planning.
- OAuth, Authentication and Authorization concepts.
- Java-based Microservices application support.
- Experience in eCommerce or large-scale customer-facing platforms.
- Experience with OKTA, Keycloak and enterprise security frameworks.
Desired Candidate Profile
- Solid communication and stakeholder management skills.
- Ability to work independently with minimal supervision.
- Excellent analytical and troubleshooting capabilities.
- Strong understanding of cloud-native technologies and distributed systems.
- Experience working in Agile and DevOps environments.
- Flexible to work in shifts and participate in production on-call rotations.
- Available to join at short notice will be preferred.
📌 AWS Site Reliability Engineer (SRE) (Hyderabad)
🏢 Zensar
📍 Hyderabad