07 Aug
|
NEWAGE TECH SOLUTION PRIVATE
|
Mumbai
07 Aug
NEWAGE TECH SOLUTION PRIVATE
Mumbai
Key Responsibilities Define measure and enforce service level objectives service level indicators and error budgets for production systems in collaboration with engineering and product teams Build maintain and improve monitoring alerting and observability frameworks across production infrastructure using tools such as Prometheus Grafana Datadog ELK Stack or similar Respond to and lead incident management activities including detection triage escalation resolution and postincident review for highseverity production incidents Conduct thorough postmortem and root cause analysis for production incidents and drive implementation of corrective and preventive actions to eliminate recurring failures Design and implement automation frameworks to eliminate toil improve operational efficiency and reduce manual intervention in repetitive infrastructure and deployment tasks Collaborate with software engineering teams to embed reliability practices including chaos engineering fault injection and resilience testing into the software development lifecycle Manage and optimize cloud infrastructure on AWS Azure or GCP for high availability fault tolerance and cost efficiency Oversee capacity planning performance benchmarking and traffic forecasting to ensure production systems can scale to meet business demand Drive adoption of DevOps and SRE best practices including CI CD infrastructure as code configuration management and change management processes Maintain comprehensive runbooks operational documentation and oncall playbooks to ensure consistent and effective incident response across the SRE team
📌 Site Reliability Engineer (Mumbai)
🏢 NEWAGE TECH SOLUTION PRIVATE
📍 Mumbai