23 Aug
|
WebSenor
|
Noida
Job Description – Cloud Data EngineerPosition
Cloud Data Engineer – Databricks / Snowflake / Azure
Experience
3+ Years
Role Overview
We are looking for a highly skilled Cloud Data Engineer to join our team for a Cloud Data Modernization initiative. The ideal candidate will have strong hands-on expertise in Databricks as the primary data platform, with Snowflake as a secondary skill, along with experience in Azure and/or AWS cloud infrastructure.
The role involves modernizing on-premises ETL workloads and migrating them to cloud-based data platforms while building scalable, secure, and high-performing data solutions.
Key Responsibilities
- Migrate and modernize on-premises ETL workloads to Azure Cloud, Databricks, and Snowflake.
- Design and implement data solutions using Azure Databricks, ADF, SHIR, Logic Apps, ADLS Gen2, Blob Storage, and Snowflake.
- Develop scalable data pipelines using Apache Spark, PySpark, Python, and SQL.
- Analyze existing on-premises ETL processes and identify opportunities for cloud modernization.
- Implement Bronze/Silver/Gold (Medallion Architecture) and Lakehouse solutions.
- Work with Delta Lake and optimize Databricks workloads for performance and scalability.
- Develop and maintain Snowflake data engineering and analytics workloads.
- Implement CI/CD and DevOps practices using GitHub Actions.
- Manage GitHub branching, pull requests, code reviews, and engineering best practices.
- Implement data quality checks, validation, reconciliation, monitoring, and observability.
- Collaborate with cross-functional teams to ensure reliable data integration and flow.
- Ensure data solutions meet required security, governance, compliance,
and quality standards.
- Troubleshoot and optimize data pipelines, Databricks jobs, Spark workloads, and Snowflake processes.
- Leverage approved AI-assisted development tools such as GitHub Copilot, Databricks Assistant, ChatGPT, or Claude to improve engineering productivity.
- Work within Agile delivery methodologies and effectively communicate risks and technical issues.
Required Qualifications
- 3+ years of experience as a Cloud Data Engineer.
- Strong hands-on experience with Azure Databricks – primary skill.
- Experience with Snowflake – secondary skill.
- Experience with Azure and/or AWS cloud platforms.
- Strong knowledge of Apache Spark, PySpark, Python, and SQL.
- Hands-on experience with Azure Data Factory, SHIR, Logic Apps, ADLS Gen2, and Blob Storage.
- Strong ETL development experience with on-premises databases and ETL technologies.
- Strong SQL skills, including complex queries, stored procedures, views, and transformations.
- Experience with GitHub, GitHub Actions, branching, pull requests, and CI/CD.
- Experience with automated testing and data quality processes.
- Understanding of Agile methodologies.
- Strong analytical, problem-solving, communication, and collaboration skills.
- Ability to work independently, manage ambiguity,
and deliver within tight deadlines.
- Ability to quickly learn and adopt new technologies.
Preferred Qualifications
- Experience with Databricks Delta Lake and Lakehouse architecture.
- Experience implementing Medallion Architecture – Bronze, Silver, and Gold.
- Knowledge of data modeling and database design.
- Experience with Databricks and Snowflake performance optimization.
- Experience with Airflow or other orchestration frameworks.
- Knowledge of data governance and data quality best practices.
- Experience with other cloud platforms and contemporary data technologies.
- Experience leveraging AI/ML for data engineering automation and workflows.
- Experience with healthcare payer data domains such as Claims, Membership, Enrollment, Provider, Clinical, and Financial data.
- Azure or Databricks certification is a plus.
Technical Skills
Primary: Azure Databricks, PySpark, Apache Spark, Python
Secondary: Snowflake, SQL
Cloud: Azure and/or AWS
Azure Data Services: ADF, SHIR, Logic Apps, ADLS Gen2, Blob Storage
Data Architecture: Delta Lake, Lakehouse, Medallion Architecture
DevOps: GitHub, GitHub Actions, CI/CD
Orchestration: Airflow / Databricks Workflows
AI Productivity: GitHub Copilot, Databricks Assistant, ChatGPT, Claude
Preferred Domain
Experience with Healthcare Payer Data is highly preferred.
Recruiter / Vendor Note
Databricks is the primary data skill and Snowflake is secondary. Candidates should have strong Azure and/or AWS cloud platform experience.
Work Location: Hybrid remote in Noida, Uttar Pradesh (Noida)
📌 Cloud Data Engineer (Noida)
🏢 WebSenor
📍 Noida