Job Title: Site Reliability Engineer (SRE)
Experience: 3–15 Years
Location: Pan India
Work Model: Hybrid
About the Role
We are looking for a skilled Site Reliability Engineer (SRE) to build, automate, and maintain highly available, scalable, and secure infrastructure. The ideal candidate will have hands-on experience in cloud platforms, infrastructure automation, containerization, monitoring, CI/CD, and Linux administration while driving reliability and operational excellence.
Key Responsibilities
- Design, deploy, and manage highly available cloud infrastructure.
- Build and maintain CI/CD pipelines to automate application deployment.
- Automate infrastructure provisioning using Infrastructure as Code (IaC) tools.
- Monitor system health, performance, and availability using enterprise monitoring tools.
- Troubleshoot production incidents and perform root cause analysis (RCA).
- Develop automation scripts to eliminate repetitive operational tasks.
- Manage containerized workloads using Docker and Kubernetes.
- Collaborate with Development, DevOps, Security, and Infrastructure teams to improve system reliability.
- Optimize application performance, scalability, and infrastructure costs.
- Ensure system security, compliance, backup, and disaster recovery best practices.
Required Skills
- Hands-on experience with any Cloud Platform: AWS,
Azure, or GCP.
- Experience in CI/CD tools and deployment automation.
- Strong knowledge of Infrastructure as Code using Terraform, Ansible, Puppet, or Chef .
- Experience with monitoring and observability tools such as Prometheus, Grafana, Datadog, or Recent Relic .
- Proficiency in Python or Shell Scripting .
- Strong understanding of Git version control.
- Good knowledge of SQL .
- Hands-on experience with Docker and Kubernetes .
- Strong Linux/Unix administration skills.
Preferred Skills
- Experience in production support and incident management.
- Knowledge of microservices architecture.
- Understanding of networking fundamentals, DNS, load balancing, and security best practices.
- Experience with logging and observability platforms.
- Exposure to Agile/Scrum environments.
Qualifications
- Bachelor's or Master's degree in Computer Science, Information Technology, or a related field.
- Relevant cloud or Kubernetes certifications are an added advantage.
What We're Looking For
- Strong analytical and troubleshooting skills.
- Excellent communication and collaboration abilities.
- Ability to work in fast-paced, production-critical environments.
- Passion for automation, reliability, and continuous improvement.
📌 Site Reliability Engineer (India)
🏢 Harp
📍 India