Site Reliability Engineer (Mumbai)

Site Reliability Engineer (Mumbai)

09 Oct
|
Envision Technology Solutions
|
Mumbai

09 Oct

Envision Technology Solutions

Mumbai

We are seeking a highly experienced Lead Site Reliability Engineer SRE with a strong background in Cloud Kubernetes Automation Observability and Production Operationscombined with a keen architectural perspective In this role you will not only drive reliability improvements and operational excellence initiatives but also play a key part in designing and evolving robust scalable and modern platforms

You will interact daily with clients and stakeholders to understand technical requirements architect solutions and guide platform modernization efforts Your engineering mindset passion for automation and ability to leverage AIpowered tools will be critical in shaping the reliability and efficiency of our infrastructure and applications You should be adept at both handson engineering and highlevel architectural planning working independently and in close collaboration with customer teams to deliver impactful results and regular updates

Responsibilities

Define architect and drive SRE best practicesincluding SLIs SLOs error budgets and operational excellence initiativesensuring they are embedded in system and platform designs

Work closely with crossfunctional teams to architect design and continuously improve system reliability scalability and performance

Participate in PI planning resolve dependencies provide architectural guidance and communicate reliability requirements to teams

Lead incident management root cause analysis RCA problem management and postmortem reviews

Collaborate with Product Owners Engineering Leads Architects and Platform teams to design and build resilient scalable architectures and systems

Architect and drive automation and selfhealing capabilities as foundational elements of the platform to reduce operational toil and improve availability





Mentor engineering teams on observability reliability engineering cloudnative technologies and production readiness

Define and implement endtoend observability strategy across infrastructure platform and applications

Build proactive monitoring for availability latency error rate saturation and businesscritical transactions

Design and maintain rolebased operational dashboards for SRE engineering leadership and customer stakeholders

Configure actionable s with proper severity routing escalation policies and noisereduction tuning

Establish and track SLISLO measurements and align s to error budget and customer impact

Integrate observability signals into incident response RCA postmortems and problem management

Technical and Process Skills

Must Have

Strong knowledge of cloudnative ecosystems including Docker Kubernetes Helm and microservices architectures

Strong experience in Linux networking troubleshooting and scripting using Shell Bash or Python

Experience supporting and operating largescale production environments on AWS Cloud

Good understanding of Infrastructure as Code IaC practices using Terraform andor CloudFormation

Excellent understanding of Git concepts release strategy branching strategy and CICD pipelines

Experience with GitOps solutions such as ArgoCD and CICD tools like GitLab CI Jenkins or GitHub Actions





Strong observability engineering experience across metrics logs traces events and service health telemetry

Strong experience implementing monitoring and observability solutions using Dynatrace Splunk Prometheus Grafana ELKOpenSearch and OpenTelemetry

Experience creating dashboards s SLISLO measurements automated remediation workflows and operational runbooks

Strong understanding of incident management availability management capacity planning disaster recovery and production support processes

Proven experience architecting and implementing scalable resilient secure and highly available cloud infrastructure following industry best practices

Ability to provide architectural guidance and technical leadership to engineering teams and stakeholders

Experience implementing scalable resilient secure and highly available cloud infrastructure following industry best practices

Experience reducing fatigue through threshold tuning deduplication suppression correlation and runbookdriven response

Experience integrating s with incident and collaboration workflows for example PagerDuty ServiceNow Jira Teams Slack Opsgenie

Experience leveraging AIpowered engineering tools such as GitHub Copilot Claude Codex ChatGPT or similar platforms for automation troubleshooting code generation operational analysis and productivity improvements

Valuable to Have

Experience building selfhealing systems using strong SRE principles and automation frameworks

Handson experience managing productiongrade Kuberne

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

📌 Site Reliability Engineer (Mumbai)
🏢 Envision Technology Solutions
📍 Mumbai

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (mumbai) / mumbai