19 Sep
|
Infobell IT Solutions
|
Hyderabad
19 Sep
Infobell IT Solutions
Hyderabad
Site Reliability Engineer (SRE)
Role Summary
We are looking for a Site Reliability Engineer (SRE) with 10+ years of experience to ensure the reliability, availability, scalability, and performance of mission-critical applications and infrastructure. The ideal candidate should have robust expertise in GCP, Kubernetes, observability, automation, incident management, and DevOps practices.
Key Responsibilities
Design, maintain, and optimize highly available and scalable production systems.
Define and manage SLIs, SLOs, and Error Budgets to improve service reliability.
Build and enhance observability solutions using tools such as Prometheus, Grafana, Datadog, Splunk, Current Relic, Checkly, Anodot, and Coralogix.
Lead incident response, troubleshooting, RCA, and post-mortem activities.
Manage cloud-native applications on GCP, including GKE/Kubernetes and Docker environments.
Develop automation and Infrastructure as Code (IaC) solutions using Terraform, Ansible, and scripting languages.
Build and support CI/CD pipelines using GitHub Actions, GitLab CI/CD, Jenkins, or similar tools.
Collaborate with development and platform teams to improve reliability, operational efficiency, and deployment practices.
Required Skills
10 years of experience in SRE, DevOps, Platform Engineering, or Production Operations.
Solid hands-on experience with Google Cloud Platform (GCP) and Kubernetes (GKE preferred).
Proficiency in Python, Go, or Java and scripting using Bash/Shell.
Strong Linux administration and networking fundamentals (TCP/IP, DNS, HTTP/S, Load Balancing).
Experience with observability, monitoring, distributed tracing, and OpenTelemetry.
Expertise in Terraform, Ansible, Infrastructure as Code, and automation.
Experience in production incident management, reliability engineering, and performance optimization.
Preferred Qualifications
Certifications such as Google Professional Cloud Engineer, CKA, CKAD, or Terraform Associate.
Experience supporting large-scale distributed systems and implementing SRE best practices.
Knowledge of cloud security, capacity planning, and operational excellence.
Key Success Measures
Improved uptime and system reliability.
Reduced MTTR and operational toil through automation.
Enhanced observability and proactive incident detection.
Consistent achievement of defined SLOs and service performance targets.
📌 Site Reliability Engineer Sre Hyderabad
🏢 Infobell IT Solutions
📍 Hyderabad