09 Oct
|
Envision Technology Solutions
|
Mumbai
09 Oct
Envision Technology Solutions
Mumbai
We are seeking a highly experienced Lead Site Reliability Engineer SRE with a strong background in Cloud Kubernetes Automation Observability and Production Operationscombined with a keen architectural perspective In this role you will not only drive reliability improvements and operational excellence initiatives but also play a key part in designing and evolving robust scalable and modern platforms
You will interact daily with clients and stakeholders to understand technical requirements architect solutions and guide platform modernization efforts Your engineering mindset passion for automation and ability to leverage AIpowered tools will be critical in shaping the reliability and efficiency of our infrastructure and applications You should be adept at both handson engineering and highlevel architectural planning working independently and in close collaboration with customer teams to deliver impactful results and regular updates
Responsibilities
Define architect and drive SRE best practicesincluding SLIs SLOs error budgets and operational excellence initiativesensuring they are embedded in system and platform designs
Work closely with crossfunctional teams to architect design and continuously improve system reliability scalability and performance
Participate in PI planning resolve dependencies provide architectural guidance and communicate reliability requirements to teams
Lead incident management root cause analysis RCA problem management and postmortem reviews
Collaborate with Product Owners Engineering Leads Architects and Platform teams to design and build resilient scalable architectures and systems
Architect and drive automation and selfhealing capabilities as foundational elements of the platform to reduce operational toil and improve availability
Mentor engineering teams on observability reliability engineering cloudnative technologies and production readiness
Define and implement endtoend observability strategy across infrastructure platform and applications
Build proactive monitoring for availability latency error rate saturation and businesscritical transactions
Design and maintain rolebased operational dashboards for SRE engineering leadership and customer stakeholders
Configure actionable s with proper severity routing escalation policies and noisereduction tuning
Establish and track SLISLO measurements and align s to error budget and customer impact
Integrate observability signals into incident response RCA postmortems and problem management
Technical and Process Skills
Must Have
Strong knowledge of cloudnative ecosystems including Docker Kubernetes Helm and microservices architectures
Strong experience in Linux networking troubleshooting and scripting using Shell Bash or Python
Experience supporting and operating largescale production environments on AWS Cloud
Good understanding of Infrastructure as Code IaC practices using Terraform andor CloudFormation
Excellent understanding of Git concepts release strategy branching strategy and CICD pipelines
Experience with GitOps solutions such as ArgoCD and CICD tools like GitLab CI Jenkins or GitHub Actions
Strong observability engineering experience across metrics logs traces events and service health telemetry
Strong experience implementing monitoring and observability solutions using Dynatrace Splunk Prometheus Grafana ELKOpenSearch and OpenTelemetry
Experience creating dashboards s SLISLO measurements automated remediation workflows and operational runbooks
Strong understanding of incident management availability management capacity planning disaster recovery and production support processes
Proven experience architecting and implementing scalable resilient secure and highly available cloud infrastructure following industry best practices
Ability to provide architectural guidance and technical leadership to engineering teams and stakeholders
Experience implementing scalable resilient secure and highly available cloud infrastructure following industry best practices
Experience reducing fatigue through threshold tuning deduplication suppression correlation and runbookdriven response
Experience integrating s with incident and collaboration workflows for example PagerDuty ServiceNow Jira Teams Slack Opsgenie
Experience leveraging AIpowered engineering tools such as GitHub Copilot Claude Codex ChatGPT or similar platforms for automation troubleshooting code generation operational analysis and productivity improvements
Valuable to Have
Experience building selfhealing systems using strong SRE principles and automation frameworks
Handson experience managing productiongrade Kuberne
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Site Reliability Engineer (Mumbai)
🏢 Envision Technology Solutions
📍 Mumbai