Site Reliability Engineer Lead (Mumbai)

Site Reliability Engineer Lead (Mumbai)

23 Sep
|
ETP INTERNATIONAL
|
Mumbai

23 Sep

ETP INTERNATIONAL

Mumbai

Job Title: Lead Site Reliability Engineer (SRE) MACH SaaS Platform

About Us
We are a product company delivering a multi-tenant SaaS platform for the Retail and e-commerce industry, built on MACH architecture (Microservices, API-First, Cloud-Native, Headless). Our platform includes 400+ microservices, API gateways, load balancers, service registries, distributed caching, and cloud-native infrastructure.
We are looking for a Lead Site Reliability Engineer (SRE) who will ensure high availability, optimal performance, and operational excellence across our setting.

Key Responsibilities
Ensure uptime SLAs and overall reliability of production, staging, and test environments.

Continuously assess all platform components for correct configuration including instance sizes, memory allocation, thread pools, JVM tuning, and log levels.

Review and optimize API gateway, service registry, load balancer, and cache service configurations.

Implement and maintain observability stack (metrics, logs, traces, dashboards, alerts).

Plan and execute capacity planning and autoscaling strategies.

Conduct performance, load, and stress testing to validate scalability and resilience.

Run chaos engineering experiments to ensure fault tolerance.

Collaborate with engineering teams to resolve performance bottlenecks and improve deployment practices.

Document configuration standards and operational best practices.





Skills & Qualifications
Must-Have:
5+ years in SRE, Platform Engineering, or Performance Engineering roles.

Solid experience with Kubernetes / container orchestration in production.

Proficiency in cloud platforms (AWS / GCP / Azure) and autoscaling mechanisms.

Expertise in JVM-based service tuning (heap sizing, GC tuning, thread pool config).

Hands-on experience with API gateway technologies (e.g., Kong, Apigee, NGINX, Envoy).

Proficient in observability tools (Prometheus, Grafana, ELK, Jaeger, OpenTelemetry).

Experience in load testing tools (k6, Gatling, JMeter) and chaos engineering (Gremlin, LitmusChaos).

Strong understanding of microservices performance patterns and distributed systems.

Nice-to-Have:
Familiarity with MACH architecture principles.

Experience in e-commerce SaaS or other high-scale transactional platforms.

Knowledge of service mesh (Istio, Linkerd).

Experience with infrastructure-as-code (Terraform, Helm, Ansible).

What We Offer
Opportunity to work on a large-scale MACH-based platform with modern tech stack.

Collaborative, engineering-driven culture with focus on innovation and reliability.

Competitive salary and benefits package.

Skilled growth opportunities in cloud-native, microservices, and SRE best practices.

Base Location: Mumbai

📌 Site Reliability Engineer Lead (Mumbai)
🏢 ETP INTERNATIONAL
📍 Mumbai

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer lead (mumbai) / mumbai