01 Sep
|
Pmgs India
|
Mumbai
Count on us. Our 'we-care' culture is more than just a motto; its a promise. From day one, we prioritize your growth, well-being, and success. You can count on us to support your career journey and help you achieve your professional goals. Join us.
Make your mark. The Plante Moran Technology Services team is recognized as a Computerworld Top 100 Places to Work in IT. We're also previous recipients of the InformationWeek IT Excellence award, the CIO 100 award, and the InformationWeek IT Excellence award. If you're seeking professional growth, like being innovative and challenged, and desire to work on impactful business technology projects, we want to hear from you!
We are excited to expand our Application Development team by building out a Site Reliability Engineering capability that will enable us to deliver more reliable, scalable, and observable solutions for our clients and our staff. We are seeking candidates who are passionate about software engineering, operational excellence, and solving business problems with state-of-the-art technology solutions. Join our award-winning technology team and grow your career!
Your role.
Your work will include, but not be limited to:
- Serve as the Site Reliability Engineer for Application Development, championing reliability, observability, performance, and operational excellence across .NET and React application.
- Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets; partner with development teams to drive improvements when budgets are at risk.
- Design, implement, and maintain monitoring, alerting, logging, and observability solutions (e.g., Azure Monitor, Application Insights, Log Analytics, Grafana) to ensure early detection and rapid resolution of incidents.
- Lead incident response, root cause analysis (RCA), and post-incident reviews; drive blameless postmortems and follow-through on corrective actions.
- Leverage AI tools (GitHub Copilot, Copilot in Azure, and similar)
to accelerate incident triaging, and root cause analysis.
- Champion the adoption of AI-assisted SRE practices across the Application Development team, including AI-powered observability, intelligent alerting, and automated RCA drafting
- Contribute to building internal AI-driven tooling - such as chatops bots, runbook copilots, or log-analysis assistants - that reduce toil for engineers.
- Partner with the Application Development team on the design and support of resilient, scalable, and secure full stack solutions built on .NET (C#) and React.
- Write SQL queries for investigation, analysis, and performance tuning during development and incident response.
- Collaborate with Infrastructure, Security, and Networking teams to ensure applications meet enterprise standards for availability, security, and compliance.
- Drive system and performance testing, and chaos/resilience testing to proactively identify weaknesses.
- Mentor developers on reliability best practices, observability, and production-readiness reviews.
- Create and maintain documentation of SRE processes.
- Participate in continuous improvement initiatives. The qualifications.
- Bachelor s degree in computer science, Information Technology, or a related field.
- Minimum of 3 years of experience troubleshooting and debugging applications built with C# .NET, React, MS SQL and APIs; hands-on development experience with these technologies will be an added advantage.
- Proven experience as a full stack developer with progressive responsibility moving into SRE and DevOps responsibilities.
- Hands-on experience with Azure platform services (App Service, Functions, Storage, Queues, Key Vault, API Management, Application Insights, Log Analytics) is required.
- Experience with git source control, GitHub and Azure DevOps including CI/CD pipelines feature-flag releases and GitHub Actions is required.
- Solid understanding of deployment processes including blue/green, canary, and feature-flag-based release patterns is required.
- Experience with AIOps platforms or AI-driven observability (e.g., Azure Monitor with AI insights, or similar) will be an added advantage.
- Experience with review and analyzing Infrastructure as Code (Bicep, ARM, or Terraform) is required.
- Strong understanding of SRE principles: SLOs/SLIs/error budgets, toil reduction, incident management, and blameless postmortems.
- Strong coding, debugging, and troubleshooting skills with a solid understanding of relational databases, data models, and distributed systems.
- Strong understanding and experience using Postman or a similar tool to test APIs during development.
- Experience in performing code reviews via the pull request process in GitHub and/or Azure DevOps.
- Excellent leadership and mentoring skills.
- Strong problem-solving and analytical abilities, particularly under pressure during production incidents.
- Knowledge of service level agreements (SLAs), SLOs, and performance metrics.
- Excellent communication and interpersonal skills, with the ability to translate technical issues for non-technical audiences.
- Ability to work in a fast-paced, dynamic environment and participate in an on-call rotation.
Disclaimer: This job posting and Location has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Site Reliability Engineer (Mumbai)
🏢 Pmgs India
📍 Mumbai