Duration: Permanent Role The Role
Plan, manage, and oversee all aspects of a Production Setting
Define strategies for Application Performance Monitoring, Optimization in Production setting
Respond to Incidents and improvise platform based on feedback and measure the reduction of incidents over time.
Support deployment of code into multiple lower environments. Supporting current processes with an emphasis on automating everything as soon as possible.
Design, develop and standardize Monitoring and Alerting mechanism for the supported applications.
Take a holistic approach to problem solving, by connecting the dots during a production event through the various technology stack that makes up the platform, to optimize meantime to recover.
Engage in and improve the whole lifecycle of services—from inception and design, through deployment, operation and refinement.
Analyse ITSM activities of the platform and provide feedback loop to development teams on operational gaps or resiliency concerns.
Support services before they go live through activities such as system design consulting, capacity planning and launch reviews.
Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead in DevOps automation and best practices.
Maintain services once they are live by measuring and monitoring availability, latency, and overall system health.
Scale systems sustainably through mechanisms like automation and evolving systems by pushing for changes that improve reliability and velocity.
Work with a global team spread across tech hubs in multiple geographies and time zones.
Ability to share knowledge and explain processes and procedures to others.
Able to perform on-call duties on a rotational basis.
Occasional off hours work required.
Must Have
Kubernetes / MKS - Must
Java / Spring Boot - Knowledge is required
Linux/Unix
Kafka (Axon)
Splunk / Dynatrace
CI/CD Platforms
GitHub / Automation Tooling
Cloud-native observability and SRE practices
Remedy / servicenow
Operations Skills:
Production support leadership
Runbook and support model creation
Monitoring and alert strategy definition
Disaster recovery and resiliency planning
Deployment readiness validation
Operational process design
Service ownership mindset
Continuous improvement and toil reduction
Root cause analysis and problem management
📌 Site Reliability Engineer Pune
🏢 SRM Digital
📍 Pune
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.