Role description
Senior SRE Lead (AWS & Monitoring Tools)
Job Summary
We are seeking an experienced Senior SRE Lead to drive production service reliability, monitoring excellence, incident management, and operational improvements. This role will provide technical leadership, oversee production monitoring activities, coordinate with cross-functional teams, and mentor engineers while ensuring platform stability and continuous service optimization. Strong hands-on experience in AWS and Monitoring tools (Grafana/ Datadog/ prometheus etc) is essential.
Key Responsibilities
- Oversee and personally review production monitoring s and incidents to ensure timely investigation and resolution.
- Drive best practices in production service management and reliability engineering through leadership and mentoring.
- Conduct daily operational stand-ups with SRE and supporting engineering teams to review incidents, priorities, and operational health.
- Manage and prioritize SRE backlog activities, follow-up actions, and sprint planning in collaboration with relevant stakeholders.
- Identify recurring operational patterns, reliability risks, and opportunities for continuous improvement.
- Act as the primary escalation point for complex production issues and systemic reliability challenges.
- Coordinate closely with engineering, architecture, and operational teams to improve service stability and observability.
- Provide technical mentorship and guidance to engineers working on SRE-related initiatives.
- Participate in release planning discussions to proactively assess operational impact, monitoring requirements, and production risks.
- Design, implement,
and enhance monitoring solutions using Monitoring tools (Grafana/ Datadog/ prometheus etc) to improve visibility and service reliability.
- Support cloud infrastructure operations and service management activities on AWS.
- Contribute to operational automation and deployment processes using Docker OR CICD OR Jenkins, depending on project requirements.
Required Skills
Proficient
- AWS
- Monitoring tool (Grafana/ Datadog/ prometheus etc)
Intermediate
- Docker OR CICD OR Jenkins
Experience
- Senior-level experience in Site Reliability Engineering, Production Operations, or related disciplines.
Why Join UST
- Be part of a global technology organization focused on innovation and digital transformation.
- Work on large-scale enterprise platforms and reliability engineering initiatives.
- Collaborate with highly skilled global teams across architecture, engineering, and operations.
- Benefit from continuous learning, career growth, and opportunities to work with up-to-date cloud technologies.
Skills
AWS, Monitoring tool, CICD
About UST
UST is a global digital transformation solutions provider. For more than 20 years, UST has worked side by side with the world’s best companies to make a real impact through transformation. Powered by technology, inspired by people and led by purpose, UST partners with their clients from design to operation. With deep domain expertise and a future-proof philosophy, UST embeds innovation and agility into their clients’ organizations. With over 30,000 employees in 30 countries, UST builds for boundless impact—touching billions of lives in the process.
📌 Senior SRE Engineer with AWS (Karnataka)
🏢 UST
📍 Karnataka