What you’ll do:
Responsible for ensuring reliability, performance, and stability of critical enterprise platforms through a balance of software engineering and operational support. Applies SRE principles to monitor, maintain, and improve system health, with an initial focus on Supplier Lifecycle Platform (SLP) and related integrations. Works with moderate guidance, independently resolving issues and contributing to continuous improvement across systems.
Support reliability, availability, and performance of enterprise platforms, with initial focus on Supplier Lifecycle Platform (SLP) and associated systems
Monitor system health using key service metrics (latency, error rates, traffic, and saturation) and respond to issues impacting production stability
Execute incident, problem, and change management processes to restore service and prevent recurrence
Develop, enhance, and maintain automation, workflows, and monitoring solutions to reduce manual effort and improve system reliability
Participate in design, testing,
and deployment activities to ensure solutions meet reliability and performance expectations across the lifecycle
Collaborate with development, integration, and business teams to support system enhancements, onboarding workflows, and data flows across platforms
Create and maintain system documentation, runbooks, and operational playbooks to support consistent execution and knowledge transfer
Perform root cause analysis for incidents and implement corrective and preventive actions to improve long-term system stability
Contribute to backlog management for reliability improvements, defect remediation, and continuous system optimization
Support secure development and operational practices aligned with enterprise standards across the engineering lifecycle
Support supplier onboarding and qualification processes by ensuring stability and effectiveness of onboarding workflows and associated system interactions
Participate in test strategy exe
📌 Senior Analyst It Site Reliability Engineer Pune
🏢 Eaton
📍 Pune