Roles and Responsibilities :
- Design, implement, and maintain scalable cloud infrastructure using Kubernetes on AWS/Azure.
- Collaborate with development teams to ensure seamless deployment of applications through CI/CD pipelines.
- Monitor application performance using tools like Grafana, Splunk, and Open Telemetry to identify bottlenecks and optimize system resources.
- Troubleshoot issues related to infrastructure failures or service disruptions in a timely manner.
Job Requirements :
- 5.1-8 years of experience in Site Reliability Engineering (SRE) role with expertise in DevOps practices.
- Robust understanding of containerization using Docker and orchestration using Kubernetes.
- Experience with monitoring tools such as Grafana, Splunk, and Open Telemetry for log analysis.