29 Aug
|
Infobell IT Solutions
|
Hyderabad
29 Aug
Infobell IT Solutions
Hyderabad
Site Reliability Engineer (SRE)
Role Summary
We are looking for a Site Reliability Engineer (SRE) with 6-8 years of experience to ensure the reliability, availability, scalability, and performance of mission-critical applications and infrastructure. The ideal candidate should have solid expertise in GCP, Kubernetes, observability, automation, incident management, and DevOps practices.
Key Responsibilities
- Design, maintain, and optimize highly available and scalable production systems.
- Define and manage SLIs, SLOs, and Error Budgets to improve service reliability.
- Build and enhance observability solutions using tools such as Prometheus, Grafana, Datadog, Splunk, New Relic, Checkly, Anodot, and Coralogix.
- Lead incident response, troubleshooting, RCA, and post-mortem activities.
- Manage cloud-native applications on GCP, including GKE/Kubernetes and Docker environments.
- Develop automation and Infrastructure as Code (IaC) solutions using Terraform, Ansible, and scripting languages.
- Build and support CI/CD pipelines using GitHub Actions, GitLab CI/CD, Jenkins, or similar tools.
- Collaborate with development and platform teams to improve reliability, operational efficiency, and deployment practices.
Required Skills
- 4-8 years of experience in SRE,
DevOps, Platform Engineering, or Production Operations.
- Strong hands-on experience with Google Cloud Platform (GCP) and Kubernetes (GKE preferred).
- Proficiency in Python, Go, or Java and scripting using Bash/Shell.
- Strong Linux administration and networking fundamentals (TCP/IP, DNS, HTTP/S, Load Balancing).
- Experience with observability, monitoring, distributed tracing, and OpenTelemetry.
- Expertise in Terraform, Ansible, Infrastructure as Code, and automation.
- Experience in production incident management, reliability engineering, and performance optimization.
Preferred Qualifications
- Certifications such as Google Professional Cloud Engineer, CKA, CKAD, or Terraform Associate.
- Experience supporting large-scale distributed systems and implementing SRE best practices.
- Knowledge of cloud security, capacity planning, and operational excellence.
Key Success Measures
- Improved uptime and system reliability.
- Reduced MTTR and operational toil through automation.
- Enhanced observability and proactive incident detection.
- Consistent achievement of defined SLOs and service performance targets.
📌 Site Reliability Engineer (SRE) (Hyderabad)
🏢 Infobell IT Solutions
📍 Hyderabad