- Design, build, and maintain robust CI/CD pipelines for automated deployment and management of data infrastructure and applications.
- Implement and manage containerization and orchestration solutions using Docker and Kubernetes.
- Develop and maintain infrastructure-as-code IaC with tools such as Terraform, ensuring version-controlled and reproducible infrastructure.
- Establish and manage observability solutions monitoring, logging, alerting using tools like Prometheus, Grafana, or Datadog to ensure system reliability and performance.
- Ensure cloud infrastructure security, compliance, and cost-efficiency across the platform.
- Collaborate closely with data engineers to troubleshoot and resolve infrastructure-related issues.
Required Skills
Must-Have
- 4+ years of experience in DevOps, Site Reliability Engineering SRE, or Data Platform Engineering roles.
- Strong hands-on experience with CI-CD tools such as Jenkins, GitLab CI-CD.
- Proven expertise in containerization Docker and orchestration Kubernetes.
- Proficiency in scripting languages such as Python or Bash.
- Experience with infrastructure-as-code tools preferably Terraform.
- Practical experience with one or more major cloud providers AWS, GCP, Azure.
- Familiarity with monitoring and observability tools such as Prometheus, Grafana, Datadog.
Valuable-to-Have
- Familiarity with data platforms and customer analytics tools such as RudderStack, Amplitude, Kochava, and Braze.
- Experience managing infrastructure for data technologies like Apache Spark, Apache Airflow, or Kafka.
- Knowledge of network security, IAM policies, and cloud security best practices.
- Relevant certifications such as Certified Kubernetes Administrator CKA or AWS Certified DevOps Engineer.