07 Aug
|
Synechron
|
Hyderabad
07 Aug
Synechron
Hyderabad
- Synechron is seeking a Lead Systems Operations Engineer
- Lead Site Reliability Engineer (App SRE) is responsible for driving reliability, automation, observability, and performance for missioncritical applications and platforms.
- This role blends software engineering excellence with operational expertise to deliver stable, scalable, and resilient services, while reducing toil and shifting operations left across the application lifecycle.
- The Lead SRE acts as a technical authority and mentor, partnering with application, platform, and DevOps teams to embed reliability into design, delivery, and operations.
Required Qualifications:
- 6+ years of Systems Engineering, Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education
Job Expectations:
- Partner with application, platform, and business stakeholders to define, implement, and govern SLIs, SLOs, and error budgets, balancing reliability with delivery velocity.
- Lead the design and continuous improvement of observability, telemetry, monitoring, and alerting, ensuring actionable insights and reduced alert fatigue.
- Identify, prioritize, and implement automation and selfhealing solutions to eliminate operational toil and improve service resilience.
- Own and lead production readiness and golive activities, including NFR validation, Permit to Operate (PTO), and operational risk assessments.
- Provide engineeringled application production support, acting as an escalation point for complex application and platform issues.
- Lead and troubleshoot major incident response (P1/P2/P3), drive indepth root cause analysis (RCA), and ensure preventative actions are implemented to achieve longterm stability.
- Influence and guide teams to shift reliability left by embedding SRE practices into design, CI/CD pipelines, and release processes.
- Mentor junior engineers and contribute to SRE standards, best practices, and operating models.
- Collaborate and consult with technical peers, colleagues, and mid to more experienced level managers to resolve systems support issues and achieve goals
Additional Required Qualifications:
- 6+ years of handson experience in production application support engineering, with a robust focus on reliability, availability, and operational excellence.
- 5+ years of experience leading and operating production systems in a Site Reliability Engineering, DevOps, or Reliability Engineering role.
- 6+ years of experience working with enterprise schedulers and databases, such as Autosys, Oracle, and MS SQL Server.
- 4+ years of experience supporting applications on Kubernetes / OpenShift platforms.
- Strong understanding of observability concepts (metrics, logs, traces, APM) using tools such as AppDynamics, ThousandEyes, Prometheus, Grafana, Splunk, and Aternity.
- Solid experience with webbased applications and application servers .
- Proven experience providing technical leadership and handson execution in complex enterprise environments.
- Excellent communication and documentation skills, with the ability to influence both technical and nontechnical stakeholders.
📌 Site Reliability Engineer Lead (Hyderabad)
🏢 Synechron
📍 Hyderabad