27 Sep
|
Qualtrix Consulting
|
India
27 Sep
Qualtrix Consulting
India
JOB TITLE: Site Reliability Engineer (SRE) – Performance & Chaos Engineering
Location: India (Remote)
Experience: 3–5 Years
:
Scope of work:
Strong experience in Site Reliability Engineering or production operations role supporting large-scale, customer-facing systems.
- Own service reliability end to end: define and maintain SLIs, SLOs and error budgets, establish steady-state behaviour for critical user journeys, and drive the reliability backlog that comes out of it.
- Performance and load engineering with k6 – author JavaScript/TypeScript test scripts, build scenarios and executors (constant-VUs, ramping-VUs, ramping-arrival rate), define custom metrics, checks and thresholds, parameterise test data, and cover HTTP, WebSocket and gRPC protocols.
- Integrate k6 into CI/CD pipelines (Harness, GitHub Actions, Jenkins) as automated performance regression gates; run distributed and containerised tests via the k6 Operator on Kubernetes or Grafana Cloud k6; stream results to Prometheus, InfluxDB, Grafana or Dynatrace for trend analysis.
- Chaos engineering – design and run hypothesis-driven experiments with an explicit steady state, controlled blast radius, abort conditions and rollback plan; plan and facilitate GameDays with application, infrastructure and client teams; convert every finding into a tracked remediation item.
- Fault injection using AWS Fault Injection Service (FIS), Gremlin, Chaos Mesh, LitmusChaos or Chaos Toolkit – CPU and memory pressure, pod and node termination, network latency and packet loss, dependency and third-party API failure, AZ evacuation, and database failover.
- Dynatrace administration and engineering – OneAgent deployment and lifecycle, management zones, entity model and tagging strategy, alerting profiles, SLO definitions, dashboards and notebooks, and day-to-day administration of tooling in the APM space.
- Dynatrace Davis AI and the Davis / Dynatrace API v2 – programmatic access to the problems, metrics, events, entities and SLO endpoints; ingest custom metrics and deployment / chaos events; tune Davis anomaly detection, alerting sensitivity and root-cause behavior; correlate load tests and chaos experiments with Davis-detected problems to validate detection and MTTR.
- Automate release validation and quality gates using Dynatrace Site Reliability Guardian / Cloud Automation (or equivalent) so that performance and resilience evidence is evaluated automatically as part of every deployment.
- Run production workloads on AWS – EKS (node groups, clusters, autoscaling), EC2, ALB/NLB, Route 53, RDS, Lambda and other AWS native services, with a strong grasp of container monitoring best practices.
- Infrastructure and observability as code with Terraform; scripting in Python, Go, Bash and JavaScript/TypeScript for automation, API integration and internal tooling.
- Participate in the on-call rotation; drive incident triage, root cause analysis and corrective actions, and work with cross-functional teams and Problem Management on escalations.
- Manage uptime and availability reporting, capacity planning and performance tuning, using both load-test evidence and production telemetry.
REQUIRED QUALIFICATIONS - KNOWLEDGE/SKILLS
- Demonstrable hands-on experience building and maintaining a k6 performance testing suite, not just running someone else’s scripts.
- Practical chaos engineering experience in a production or production-like environment, with evidence of the reliability defects it surfaced.
- Working knowledge of the Dynatrace API v2 and Davis AI, including authentication and token scopes, rate limits, and consuming problem and metric data from automation.
- Ability to define alert standards for production environments and implement them, including tuning to reduce false positives and alert fatigue.
- Strong understanding of distributed systems, networking and troubleshooting techniques – latency, timeouts, retries, connection pooling, DNS, TLS and load balancing.
- Experience with automated build pipelines and continuous integration. Source control, branching and merging:
git/svn/etc (Repository Management).
- Familiarity with configuration management software and observability standards such as OpenTelemetry.
- Provide support to teams for alarms and outages on an as needed basis, and work with development teams and management to ensure high availability.
- Communication Skills- The ability to communicate verbally and in writing with all levels of employees and management, speaks and writes clearly and understandably at the right level.
- Integrity and Trust- Involves being widely trusted, being seen as a direct, truthful individual, can present the unvarnished truth in an appropriate and helpful manner, keeps confidences, admits mistakes, and doesn’t misrepresent him/herself for personal gain.
- Teamwork- Works well in a collaborative setting, volunteering for and completing assignments, acting as a positive team member by contributing to discussions.
Skills Required:
Essential Skills:
SRE practices (SLI/SLO/error budgets, incident and problem management); k6 performance and load engineering; Chaos Engineering; Dynatrace including Davis AI and the Davis / Dynatrace API v2. Supporting stack: AWS (EKS, EC2, native services), Kubernetes, Terraform and CI/CD pipeline automation. Refer to the scope of work for the detail.
Desired Skills:
- Grafana, Prometheus and OpenTelemetry; experience with other load testing tools (JMeter, Gatling, Locust) in addition to k6.
- Harness for CI/CD pipeline automation, deployment orchestration and release management; progressive delivery patterns (canary, blue-green, feature flags).
- Certifications such as Dynatrace Associate/Professional, Certified Kubernetes Administrator (CKA), or AWS Solutions Architect / DevOps Engineer.
- Hands-on experience and familiarity with AI-assisted “vibe coding” using tools such as GitHub Copilot, Claude Code, Cursor, or similar AI development platforms is preferred. Given the evolving nature of these technologies, practical exposure and a transparent understanding of AI-assisted development concepts are acceptable.
Education Qualification:
B.E. / B.Tech / MCA / M.Sc. in Computer Science, Information Technology or an equivalent discipline
📌 Site Reliability Engineer (India)
🏢 Qualtrix Consulting
📍 India