30 Jul
|
Michael Page
|
Chennai
30 Jul
Michael Page
Chennai
Role purpose The Senior Site Reliability Engineer strengthens the reliability of business-critical services by converting observability signals into practical operational action.
This is a hands-on role for someone who can read telemetry, isolate faults, guide recovery, improve runbooks and work with product, platform, cloud, observability and Event Management teams to make services more reliable over time.
Role boundaries The SRE does not replace product, platform, cloud or Incident Management ownership. The role improves how those teams detect, isolate, restore and prevent issues.
The role is not a generic L3 support position. It exists to improve reliability practices, event quality, technical triage, automation and production readiness.
Experience required
12+ years overall experience across software engineering, production operations, platform engineering, cloud engineering or reliability engineering.
Solid hands-on experience in SRE, production reliability, high-availability operations or platform reliability.
Proven exposure to 24x7 production environments, on-call support, high-severity troubleshooting and incident response.
Transparent evidence of reducing incidents through engineering improvements, not only through firefighting.
Experience working with product, platform,
cloud, service management, observability and operations teams.
Must have :
Linux troubleshooting, performance analysis and network
fundamentals such as DNS, TCP, latency, packet loss and connectivity checks. bservability across metrics, logs, traces, synthetics, dashboards, alert tuning and service maps. • Controlled resilience testing or chaos engineering.
Hands-on use of Datadog, Grafana, Prometheus, Dynatrace,
Splunk, ELK, OpenTelemetry or similar tools. • High-throughput, API-heavy, customer-facing or distributed platforms.
Cloud experience in AWS, Azure or GCP, including compute,
storage, load balancing, autoscaling and network paths. • Linking technical service health to business journeys or customer impact.
Kubernetes, Docker, containers, microservices, APIs, databases,
queues and distributed service ecosystems.
Automation or scripting using Python, Shell, Go, Ansible,
Terraform, CloudFormation or similar tools.
CI/CD, rollback, feature flags, blue-green releases, canary releases
and release safety.
Working awareness of ServiceNow, Jira, Incident, Problem,
Change, Knowledge and service reporting practices.
📌 Site Reliability Engineer Lead Chennai
🏢 Michael Page
📍 Chennai