06 Aug
|
TESCRA
|
Bengaluru
Key Responsibilities
- Design, provision, and manage AWS infrastructure using Infrastructure as Code (Terraform), ensuring scalability, security, and cost efficiency.
- Build and maintain robust CI/CD pipelines to support fast, reliable, and automated deployments across development, staging, and production environments.
- Develop and execute comprehensive performance and load testing strategies using K6, integrating test scripts into CI/CD pipelines to proactively identify and address performance regressions.
- Design and implement sophisticated chaos engineering experiments using Harness Chaos Engineering and Terraform, simulating real-world failure scenarios such as node failures, network latency, and pod/service disruptions to validate system resilience.
- Configure and manage Dynatrace with its advanced Davis AI engine for full-stack observability, enabling proactive anomaly detection, efficient root-cause analysis, and reduction of Mean Time to Resolution (MTTR).
- Set up intelligent alerting and automated remediation workflows based on Davis AI-driven insights to minimize service disruptions.
- Collaborate closely with development, QA, and SRE teams to embed reliability, performance, and observability best practices throughout the software development lifecycle.
- Continuously monitor AWS environments for opportunities in cost optimization, security compliance, and performance tuning.
- Document runbooks, chaos experiment results, and performance benchmarks to foster a culture of continuous improvement and knowledge sharing.
- Participate in on-call rotations and incident response efforts, driving blameless postmortems and implementing long-term solutions to prevent recurrence.
Educational Qualifications
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
Must-Have Skills
- K6 (Performance & Load Testing):
Hands-on experience writing and executing K6 test scripts (JavaScript-based) for load, stress, spike, and soak testing. Proven ability to integrate K6 into CI/CD pipelines (e.g., Jenkins, GitLab CI, GitHub Actions) for automated performance gating. Demonstrated skill in analyzing K6 test results and translating them into actionable performance improvements.
- Chaos Testing with Harness & Terraform: Practical experience designing and running chaos engineering experiments using Harness Chaos Engineering (CE).
Strong
Terraform skills for provisioning and managing the infrastructure used in chaos experiments, including fault injection targets and isolated environments.
Experience defining steady-state hypotheses, blast radius controls, and rollback strategies for chaos tests. Familiarity with common failure modes in distributed/cloud-native systems.
- Dynatrace Davis AI: Deep working knowledge of Dynatrace, with a particular focus on the Davis AI engine for automated anomaly detection and root-cause analysis.
Experience configuring Dynatrace monitoring (infrastructure, application, and synthetic) across AWS workloads. Ability to build custom dashboards, define Service Level Objectives (SLOs), and create AI-driven alerting policies.
Experience leveraging Davis AI insights for proactive incident prevention.
Good-to-Have Skills
- Robust AWS expertise, including EC2, ECS/EKS, Lambda, VPC, IAM, S3, RDS, CloudWatch, and ALB/NLB.
- Solid Terraform/IaC experience beyond chaos testing, including module design, state management, and workspaces.
- Proficiency with containerization and orchestration technologies such as Docker and Kubernetes.
- Experience with CI/CD tooling like Jenkins, GitLab CI, GitHub Actions, or similar platforms.
- Scripting skills in Python, Bash, or JavaScript.
- Understanding of SRE principles, including SLIs/SLOs/SLAs, error budgets, and incident management frameworks.
- AWS certifications (e.g., AWS Certified DevOps Engineer, AWS Certified Solutions Architect) are considered a significant advantage.
📌 Devops Engineer (Bengaluru)
🏢 TESCRA
📍 Bengaluru