Site Reliability Engineer (India)

Site Reliability Engineer (India)

27 Sep
|
Qualtrix Consulting
|
India

27 Sep

Qualtrix Consulting

India

JOB TITLE: Site Reliability Engineer (SRE) – Performance & Chaos Engineering

Location: India (Remote)

Experience: 3–5 Years

:

Scope of work:

Strong experience in Site Reliability Engineering or production operations role supporting large-scale, customer-facing systems.

- Own service reliability end to end: define and maintain SLIs, SLOs and error budgets, establish steady-state behaviour for critical user journeys, and drive the reliability backlog that comes out of it.
- Performance and load engineering with k6 – author JavaScript/TypeScript test scripts, build scenarios and executors (constant-VUs, ramping-VUs, ramping-arrival rate), define custom metrics, checks and thresholds, parameterise test data, and cover HTTP, WebSocket and gRPC protocols.
- Integrate k6 into CI/CD pipelines (Harness, GitHub Actions, Jenkins) as automated performance regression gates; run distributed and containerised tests via the k6 Operator on Kubernetes or Grafana Cloud k6; stream results to Prometheus, InfluxDB, Grafana or Dynatrace for trend analysis.
- Chaos engineering – design and run hypothesis-driven experiments with an explicit steady state, controlled blast radius, abort conditions and rollback plan; plan and facilitate GameDays with application, infrastructure and client teams; convert every finding into a tracked remediation item.
- Fault injection using AWS Fault Injection Service (FIS), Gremlin, Chaos Mesh, LitmusChaos or Chaos Toolkit – CPU and memory pressure, pod and node termination, network latency and packet loss, dependency and third-party API failure, AZ evacuation, and database failover.
- Dynatrace administration and engineering – OneAgent deployment and lifecycle, management zones, entity model and tagging strategy, alerting profiles, SLO definitions, dashboards and notebooks, and day-to-day administration of tooling in the APM space.
- Dynatrace Davis AI and the Davis / Dynatrace API v2 – programmatic access to the problems, metrics, events, entities and SLO endpoints; ingest custom metrics and deployment / chaos events; tune Davis anomaly detection, alerting sensitivity and root-cause behavior; correlate load tests and chaos experiments with Davis-detected problems to validate detection and MTTR.




- Automate release validation and quality gates using Dynatrace Site Reliability Guardian / Cloud Automation (or equivalent) so that performance and resilience evidence is evaluated automatically as part of every deployment.
- Run production workloads on AWS – EKS (node groups, clusters, autoscaling), EC2, ALB/NLB, Route 53, RDS, Lambda and other AWS native services, with a strong grasp of container monitoring best practices.
- Infrastructure and observability as code with Terraform; scripting in Python, Go, Bash and JavaScript/TypeScript for automation, API integration and internal tooling.
- Participate in the on-call rotation; drive incident triage, root cause analysis and corrective actions, and work with cross-functional teams and Problem Management on escalations.
- Manage uptime and availability reporting, capacity planning and performance tuning, using both load-test evidence and production telemetry.

REQUIRED QUALIFICATIONS - KNOWLEDGE/SKILLS

- Demonstrable hands-on experience building and maintaining a k6 performance testing suite, not just running someone else’s scripts.
- Practical chaos engineering experience in a production or production-like environment, with evidence of the reliability defects it surfaced.
- Working knowledge of the Dynatrace API v2 and Davis AI, including authentication and token scopes, rate limits, and consuming problem and metric data from automation.
- Ability to define alert standards for production environments and implement them, including tuning to reduce false positives and alert fatigue.
- Strong understanding of distributed systems, networking and troubleshooting techniques – latency, timeouts, retries, connection pooling, DNS, TLS and load balancing.
- Experience with automated build pipelines and continuous integration. Source control, branching and merging:



git/svn/etc (Repository Management).
- Familiarity with configuration management software and observability standards such as OpenTelemetry.
- Provide support to teams for alarms and outages on an as needed basis, and work with development teams and management to ensure high availability.
- Communication Skills- The ability to communicate verbally and in writing with all levels of employees and management, speaks and writes clearly and understandably at the right level.
- Integrity and Trust- Involves being widely trusted, being seen as a direct, truthful individual, can present the unvarnished truth in an appropriate and helpful manner, keeps confidences, admits mistakes, and doesn’t misrepresent him/herself for personal gain.
- Teamwork- Works well in a collaborative setting, volunteering for and completing assignments, acting as a positive team member by contributing to discussions.

Skills Required:

Essential Skills:

SRE practices (SLI/SLO/error budgets, incident and problem management); k6 performance and load engineering; Chaos Engineering; Dynatrace including Davis AI and the Davis / Dynatrace API v2. Supporting stack: AWS (EKS, EC2, native services), Kubernetes, Terraform and CI/CD pipeline automation. Refer to the scope of work for the detail.

Desired Skills:

- Grafana, Prometheus and OpenTelemetry; experience with other load testing tools (JMeter, Gatling, Locust) in addition to k6.
- Harness for CI/CD pipeline automation, deployment orchestration and release management; progressive delivery patterns (canary, blue-green, feature flags).
- Certifications such as Dynatrace Associate/Professional, Certified Kubernetes Administrator (CKA), or AWS Solutions Architect / DevOps Engineer.
- Hands-on experience and familiarity with AI-assisted “vibe coding” using tools such as GitHub Copilot, Claude Code, Cursor, or similar AI development platforms is preferred. Given the evolving nature of these technologies, practical exposure and a transparent understanding of AI-assisted development concepts are acceptable.

Education Qualification:

B.E. / B.Tech / MCA / M.Sc. in Computer Science, Information Technology or an equivalent discipline

📌 Site Reliability Engineer (India)
🏢 Qualtrix Consulting
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (india) / india