Senior Site Reliability Engineer (Senior level)
Description
Site Reliability Engineering (SRE) is a up-to-date way of delivering IT Solutions by imbibing Software engineering principles in Service Delivery to reduce IT Risk to business, improve business resilience, attain predictability & reliability, optimize cost of IT Infra and Ops
A Site Reliability Engineer typically has deep software engineering experience encompassing design, build, deploy and manage / maintain an IT solution ensuring resilience, reliability, and performance.
An SRE is a bridge between development and operations by applying a software engineering mindset to the development, deployment, and maintenance of applications to maximize system reliability & automation, while improving efficiencies by optimizing resources
Responsibilities
Defining SLA/SLO/SLI for a product / service
Engineering in resilient design and implementation practices into solutions as they go through the product life cycle
Engineering out manual effort (Toil) through the development of automated processes and services (e.g.,
Automated Management of Systems, CI/CD improvements)
Developing Observability Solutions to track, report, and measure SLA adherence
Help Optimize Cost of IT Infra & Operations - FinOps
Critical Situation management
SOP / Runbook automation, Toil reduction
Data Analytics & System trend analysis
Typical Skills and Background
4+ years of experience in software product engineering principles, processes and systems
Hands-on experience in Java / J2EE, one of web server (Apache Tomcat or IBM HTTP Server), one of the application servers (Tomcat/WebSphere), and any major RDBMS like Oracle
Hands-on experience in at least one CI-CD (Azure DevOps, GitLab CI/CD, Jenkins) and IaC tools (Terraform, AWS CloudFormation, Ansible etc.)
Experience in at least one cloud technology (AWS/Azure/GCP etc. and Docker, Pivotal, Kubernetes, OpenShift etc.) and its reliability tools (Azure AppInsight, CloudWa
📌 Resilience And Reliability Engineer Bengaluru
🏢 EY
📍 Bengaluru