09 Aug
|
Edge Executive Search
|
Gurugram
09 Aug
Edge Executive Search
Gurugram
This is a Site Reliability Engineer (SRE) role with 2 to 6 years of experience in responsible for building, supporting, protecting, and managing highly available and resilient production platforms. The engineer will work closely with software development, infrastructure, database, network, security, and operations teams to improve system reliability and operational efficiency.
Key Responsibilities
- Build and maintain scalable cloud infrastructure using Infrastructure as Code (IaC).
- Develop automation solutions using Python.
- Monitor system health and improve platform reliability using observability tools.
- Define and support SLIs, SLOs, and participate in on-call rotations.
- Perform capacity planning and Disaster Recovery (DR) design and execution.
- Ensure database availability, performance, and security.
- Work with Kubernetes, Docker, and cloud platforms.
- Use AI and Machine Learning tools to automate operations, detect anomalies, and improve incident response.
- Collaborate with engineering and product teams to improve reliability throughout the software development lifecycle.
Required Skills
- Solid Python programming.
- Infrastructure as Code using Terraform, Ansible, CloudFormation, or Pulumi.
- Cloud platforms (AWS, Azure, or GCP).
- Kubernetes and Docker.
- Monitoring tools such as Prometheus and Grafana.
- SQL and database troubleshooting.
- Understanding of networking, SSL/TLS, messaging technologies (IBM MQ/AMQ, XML/XSLT), and production incident management.
📌 Site Reliability Engineer (Gurugram)
🏢 Edge Executive Search
📍 Gurugram