04 Oct
|
Persistent
|
Pune
Job Description
About Persistent
We are an AI-led, platform-driven Digital Engineering and Enterprise Modernization partner, combining deep technical expertise and industry experience to help our clients anticipate what's next. Our offerings and proven solutions create a unique competitive advantage for our clients by giving them the power to see beyond and rise above. We work with many industry-leading organizations across the world, including 20 Fortune 50 companies and 4 of the 5 top banks in both the US and India, and numerous innovators across the healthcare ecosystem.
Our disruptor's mindset, commitment to client success, and agility to thrive in the dynamic environment have enabled us to sustain our growth momentum. Persistent has been recognized across top industry platforms for innovation, leadership, and inclusion. We reported $1,654.4M FY26 revenue with 17.4% Y-o-Y growth. We have delivered 24 sequential quarters of growth with $436.0M in Q4 FY26 revenue, up 3.2% Q-o-Q and 16.2% Y-o-Y growth. Our 27,500+ global team members, located in 18 countries, have been instrumental in helping the market leaders transform their industries. We have been recognized as the Fastest Growing IT Services Brand Globally in the 2026 Brand Finance IT Services 25 Report. We named a Leader in the Everest Group Private Equity (PE) Services PEAK Matrix Assessment 2026 and Software Product Engineering PEAK Matrix Assessment 2026.
About Position:
We are seeking an experienced Site Reliability Engineer (SRE) to improve the reliability, scalability, observability, and operational excellence of Government Digital Portals (GPD). This role combines software engineering, automation, cloud technologies, observability practices, and operational excellence to build highly available and resilient digital platforms. The ideal candidate will focus on reducing operational toil through automation, enhancing system reliability, and developing tools and dashboards that enable proactive monitoring and incident management.
-
Role: Site Reliability Engineer
- Location: All Persistent Locations
- Experience: 5 to 12 Years
- Job Type: Full Time Employment
What You'll Do:
- Own the availability, reliability, performance, scalability, and resiliency of GPD applications and platforms.
- Design, develop, and maintain automation and reliability tools using Python, Java, or other programming languages.
- Automate release validations, deployment checks, system health monitoring, certificate validations,
and dependency verification processes.
- Build automated workflows for incident detection, triage, remediation, and operational support.
- Implement and enhance end-to-end observability across application, infrastructure, platform, and network layers.
- Design and maintain dashboards using Dynatrace, Splunk, Elastic, Azure Monitor, and custom telemetry solutions.
- Develop UI-based operational tools and visualizations for release tracking, operational reporting, and incident management.
- Monitor system health, performance metrics, application logs, and infrastructure telemetry.
- Collaborate with development teams to ensure production readiness and operational excellence.
- Partner with platform engineering, DevOps, cloud engineering, and frontend teams to improve system reliability.
- Support incident response activities, on-call processes, and WAR room operations during critical production incidents.
- Drive continuous improvement initiatives to reduce manual effort and operational overhead.
- Identify reliability risks and implement proactive solutions to improve system stability.
- Participate in architecture reviews and contribute to platform modernization efforts.
- Establish and maintain operational best practices and service reliability standards.
Expertise You'll Bring:
- 5 to 12 years of experience in Site Reliability Engineering, DevOps, Platform Operations, or Production Support environments.
- Strong proficiency in at least one programming language such as Python, Java, or UI development technologies.
- Hands-on experience developing automation tools, scripts, and operational workflows.
- Experience working with cloud platforms including Azure, AWS, or Google Cloud Platform (GCP).
- Strong expertise in observability, monitoring, logging, and telemetry platforms.
- Hands-on experience with Dynatrace, Splunk, Kubernetes, Elastic Stack, or Azure Monitor.
- Experience analyzing logs, identifying root causes, and troubleshooting complex production issues.
- Strong understanding of application performance monitoring and infrastructure monitoring.
- Experience building operational dashboards,
visualizations, and reporting solutions.
- Knowledge of incident management processes, production support, and service reliability practices.
- Experience designing automation for health checks, deployment validation, dependency management, and incident remediation.
- Understanding of CI/CD pipelines and DevOps methodologies.
- Hands-on experience with Git, Jenkins, GitHub, Docker, and up-to-date deployment practices.
- Familiarity with containerized environments and Kubernetes platforms.
- Strong analytical, debugging, and problem-solving abilities.
- Experience working in Agile and collaborative engineering environments.
- Excellent communication and stakeholder management skills.
- Ability to work effectively during critical production incidents and high-pressure situations.
- Bachelor's degree in Computer Science, Engineering, Information Technology, or equivalent relevant experience.
Benefits:
- Competitive salary and benefits package
- Culture focused on talent development with quarterly growth opportunities and company-sponsored higher education and certifications
- Opportunity to work with cutting-edge technologies
- Employee engagement initiatives such as project parties, flexible work hours, and Long Service awards
- Annual health check-ups
- Insurance coverage: group term life, personal accident, and Mediclaim hospitalization for self, spouse, two children, and parents
Values-Driven, People-Centric & Inclusive Work Environment:
Persistent is dedicated to fostering diversity and inclusion in the workplace. We invite applications from all qualified individuals, including those with disabilities, and regardless of gender or gender preference. We welcome diverse candidates from all backgrounds.
- We support hybrid work and flexible hours to fit diverse lifestyles.
- Our office is accessibility-friendly, with ergonomic setups and assistive technologies to support employees with physical disabilities.
- If you are a person with disabilities and have specific requirements, please inform us during the application process or at any time during your employment
Let's unleash your full potential at Persistent - persistent.com/careers
“Persistent is an Equal Opportunity Employer and prohibits discrimination and harassment of any kind.”
1
Open Positions
Linux,Kubernetes,Python
Skills Required
PUNE
Location
Linux,Kubernetes,Docker,Python
Desirable Skills
185506
Job Code
📌 Site Reliability Engineer (SRE) (Pune)
🏢 Persistent
📍 Pune