Role : Data Platform Infrastructure Engineer with Databricks
Location : Scottsdale AZ (onsite)
Preferred Skill Combination (Must-Have Exposure to One or Both):
Cloudera + Databricks + Terraform + AWS Role Overview:
We are seeking a highly skilled Data Platform Infrastructure Engineer to design, build, and manage scalable data platforms across on-premise and cloud settings. The role involves working with cluster technologies, infrastructure automation, and up-to-date data ecosystems to enable reliable and high-performing data platforms. Key Responsibilities:
Design, deploy, and manage data platform infrastructure across on-prem (Cloudera) and cloud (AWS, Databricks) environments
Build and maintain distributed data clusters ensuring high availability, scalability, and performance
Automate infrastructure provisioning using Terraform and Ansible
Manage and optimize Cloudera Hadoop ecosystems (HDFS, Hive, Spark, YARN, etc.)
Deploy and manage Databricks workspaces, clusters, and integrations on AWS
Implement infrastructure-as-code (IaC) and configuration management best practices
Monitor cluster performance, troubleshoot issues, and ensure system reliability
Collaborate with data engineers, architects, and DevOps teams to support data pipelines and analytics workloads
Ensure security, compliance, and governance across data platforms
Support migration from on-prem to cloud-based data platforms Technical Skills Required:
Core Technologies
Strong experience in Cloudera (CDH/CDP) cluster setup and administration
Hands-on experience with Databricks (cluster management, jobs, notebooks)
Solid exposure to AWS (EC2, S3, IAM, VPC, EMR, networking concepts)
Infrastructure & Automation:
Expertise in Terraform (mandatory) for infrastructure provisioning
Proficiency in Ansible for configuration management and automation
Experience with CI/CD pipelines for infrastructure deployments
Cluster & Data Technologies:
Experience managing distributed systems / cluster technologies
Strong understanding of
Hadoop ecosystem (HDFS, Hive, Spark, Kafka, etc.)
Spark performance tuning and cluster optimization
Knowledge of containerization (Docker/Kubernetes) is a plus
📌 Data Platform Infrastructure Engineer With Databricks Dehradun
🏢 Git
📍 Dehradun