Site Reliability / Platform Engineer (India)

Site Reliability / Platform Engineer (India)

06 Aug
|
Qualtrix Consulting
|
India

06 Aug

Qualtrix Consulting

India

Location: INDIA (Remote)

Employment Type: Full time / Contract

Experience Level: 7+ Years

About the Role

We are seeking a hands-on engineer with deep expertise in chaos engineering, performance/load testing with K6, and AWS compute optimization across EKS (Kubernetes) and ECS. This role is central to improving the resilience, scalability, and cost-efficiency of our production infrastructure by proactively identifying failure modes, validating system behavior under load, and tuning compute resources for performance and cost.

Key Responsibilities

- Design, implement, and run chaos engineering experiments (e.g., pod/node failure, network latency injection, resource starvation, AZ/region failure simulation) using tools such as Chaos Mesh, LitmusChaos, Gremlin, or AWS Fault Injection Simulator (FIS).
- Build and maintain K6 performance/load testing scripts to validate throughput, latency, and stability of services under varying load conditions; integrate K6 into CI/CD pipelines for continuous performance validation.
- Analyze test and chaos experiment results to identify bottlenecks, single points of failure, and degradation patterns; produce actionable remediation recommendations.
- Optimize AWS EKS cluster configurations — node groups, Karpenter/Cluster Autoscaler, pod resource requests/limits, HPA/VPA tuning, and workload right-sizing.
- Optimize AWS ECS (Fargate and EC2 launch types) task definitions, service scaling policies, and cluster capacity providers for cost and performance efficiency.
- Drive compute cost optimization initiatives — Spot/Reserved/Savings Plans strategy, right-sizing, bin-packing, and resource utilization analysis across EKS/ECS workloads.
- Collaborate with SRE, DevOps, and application teams to define resilience SLOs/SLIs and embed chaos/performance testing into the software delivery lifecycle.
- Build observability and reporting around chaos experiments and load tests (dashboards, alerts, post-experiment reports)



using tools such as CloudWatch, Prometheus/Grafana, or Datadog.
- Document runbooks, failure scenarios, and lessons learned; contribute to a culture of resilience engineering and continuous improvement.

Required Skills & Experience
- Proven hands-on experience with chaos engineering practices and tooling (Chaos Mesh, Gremlin, LitmusChaos, AWS FIS, or equivalent).
- Strong practical experience writing and executing K6 test scripts (load, stress, spike, soak testing), including scripting in JavaScript and integrating K6 with CI/CD.
- Deep working knowledge of AWS EKS — cluster architecture, node/pod scaling, autoscaling strategies, networking (VPC CNI, security groups), and workload optimization.
- Deep working knowledge of AWS ECS — Fargate/EC2 launch types, task/service definitions, capacity providers, and auto-scaling configuration.
- Experience with AWS compute cost optimization techniques (Spot instances, Savings Plans, right-sizing, Compute Optimizer).
- Solid understanding of Kubernetes fundamentals (Deployments, Services, HPA/VPA, resource quotas, node affinity/taints).
- Experience with Infrastructure-as-Code (Terraform, CloudFormation, or CDK) for provisioning and managing EKS/ECS environments.
- Familiarity with observability stacks (Prometheus, Grafana, CloudWatch, Datadog, or similar) for monitoring resilience and performance metrics.
- Strong scripting/automation skills (Python, Bash, or Go).
- Excellent troubleshooting skills with the ability to root-cause distributed system failures under load.

Preferred / Nice-to-Have
- AWS Certifications (Solutions Architect, DevOps Engineer, or SysOps Administrator).
- Certified Kubernetes Administrator (CKA) or equivalent.
- Experience with GitOps tools (ArgoCD, Flux) and CI/CD platforms (Jenkins, GitLab CI, GitHub Actions).
- Exposure to service mesh technologies (Istio, App Mesh) for fault injection testing.
- Prior experience running Game Days or resilience/DR exercises in production environments.

📌 Site Reliability / Platform Engineer (India)
🏢 Qualtrix Consulting
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability / platform engineer (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: site reliability / platform engineer (india) / india