Manage and maintain infrastructure reliability, automate deployment processes using Terraform and Ansible, monitor systems with Splunk and Dynatrace, troubleshoot production issues, implement disaster recovery and resiliency strategies, optimize Linux environments, develop Python automation scripts, collaborate with cross-functional teams, and ensure high availability, performance, and operational excellence.
Preferred candidate profile
- 5+ years of experience in Infrastructure/SRE or Platform Engineering.
- Strong knowledge of Terraform, Ansible, Python, and Linux.
- Experience with Splunk and Dynatrace monitoring tools.
- Understanding of Infrastructure as Code (IaC) and automation.
- Familiarity with resiliency, disaster recovery (DR), and incident management.
- Valuable analytical, troubleshooting, and communication skills.