Ensure high availability, reliability, and performance of production applications and infrastructure. Implement and manage Infrastructure as Code (IaC) using Terraform and Ansible. Automate infrastructure provisioning, configuration management, and deployment processes.
Monitor system health, performance, and availability using observability and monitoring tools. Manage incident response, troubleshooting, root cause analysis (RCA), and service restoration. Define and maintain SLAs, SLOs, and SLIs to improve system reliability.
Support CI/CD pipelines and automate release and deployment activities. Collaborate with development, cloud, security, and operations teams to improve system stability. Perform capacity planning, performance tuning, and scalability assessments.
Manage and support cloud infrastructure across AWS, Azure, or GCP environments. Implement disaster recovery, backup, and business continuity solutions. Develop automation scripts and self-healing mechanisms to reduce manual intervention.
Create and maintain operational documentation, runbooks, and knowledge articles. Participate in 24x7 on-call support and production issue resolution.
📌 Site Reliability Engineer (SRE) (Hyderabad)
🏢 Infosys
📍 Hyderabad
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.