Data Engineer (India)

Data Engineer (India)

15 Sep
|
WebSenor InfoTech
|
India

15 Sep

WebSenor InfoTech

India

– Cloud Data Engineer

Cloud Data Engineer – Databricks / Snowflake

Experience: 3–8 Years

Employment Type: Full-Time

Technology: Data Engineering / Cloud Data / Data Modernization

Primary Skill: Databricks

Secondary Skill: Snowflake

Cloud: Azure and/or AWS

About the Role

We are seeking an experienced Cloud Data Engineer to join our Cloud Data Modernization initiative. The ideal candidate will have strong hands-on expertise in Databricks, Apache Spark, PySpark, Azure/AWS cloud data services, and Snowflake, with experience modernizing and migrating enterprise ETL workloads from on-premises environments to scalable cloud-based data platforms.

The candidate will contribute to the design, development, migration, optimization, and production support of modern data platforms using Databricks as the primary data platform, supported by Snowflake and cloud services. The role requires strong experience in data engineering, ETL development, data quality, CI/CD, performance optimization, and Agile delivery.

Key Responsibilities

- Design, develop, and maintain scalable data engineering solutions using Databricks, Apache Spark, and PySpark.
- Contribute to the migration and modernization of on-premises ETL workloads to Azure/AWS cloud platforms.
- Analyze existing ETL processes, identify modernization opportunities, and convert legacy workloads into cloud-native data pipelines.
- Develop and maintain data pipelines using Databricks, Azure Data Factory (ADF), ADLS Gen2, Blob Storage, Logic Apps, and related cloud services.
- Develop and optimize data processing workloads using PySpark, Spark SQL, Python, and SQL.
- Implement Delta Lake and Lakehouse architecture, including Bronze, Silver, and Gold/Medallion data layers.
- Develop and support Snowflake-based data engineering and analytics workloads as a secondary data platform.
- Design efficient data models, transformations, ingestion frameworks, and data integration solutions.
- Implement automated data validation, reconciliation, profiling, quality checks, and monitoring.
- Develop solutions for data observability, pipeline monitoring, failure handling, and operational support.
- Optimize Databricks/Spark and Snowflake workloads for performance, scalability, reliability, and cost efficiency.
- Implement CI/CD and DevOps practices using Git, GitHub, GitHub Actions, and related engineering tools.




- Participate in GitHub branching strategies, pull requests, peer reviews, code reviews, and release processes.
- Develop automated testing frameworks and validation processes for data pipelines.
- Collaborate with data architects, cloud engineers, application teams, QA, DevOps, security, and business stakeholders.
- Ensure data solutions meet enterprise security, compliance, governance, and engineering standards.
- Troubleshoot production data pipeline issues and perform root-cause analysis.
- Contribute to technical design discussions, documentation, estimation, and delivery planning.
- Leverage AI-assisted development tools such as GitHub Copilot, Databricks Assistant, ChatGPT, or Claude to improve engineering productivity.
- Support continuous improvement and adoption of modern cloud data engineering practices.

Required Qualifications
- 3–8 years of experience in Data Engineering, Cloud Data Engineering, or related roles.
- Strong hands-on experience with Databricks as a primary data engineering platform.
- Strong experience with Apache Spark and PySpark for large-scale data processing.
- Robust experience with Python and SQL.
- Hands-on experience with Azure and/or AWS cloud data services.
- Experience with Azure services such as:
- Azure Data Factory (ADF)
- Self-hosted Integration Runtime (SHIR)
- Azure Data Lake Storage Gen2 (ADLS Gen2)
- Azure Blob Storage
- Logic Apps
- Hands-on experience with Snowflake for data engineering and analytics workloads.
- Experience developing complex SQL queries, stored procedures, views, transformations, and data pipelines.
- Experience with ETL/ELT development and on-premises database technologies.
- Experience with data migration and cloud modernization projects.
- Experience with Git/GitHub, branching strategies, pull requests, and CI/CD pipelines.
- Experience with GitHub Actions or equivalent CI/CD technologies.
- Experience implementing data quality, validation, reconciliation, and automated testing.




- Strong understanding of data engineering best practices and software development lifecycle.
- Experience working in Agile/Scrum delivery environments.
- Strong analytical, troubleshooting, communication, and collaboration skills.
- Ability to work independently, manage ambiguity, and deliver solutions within deadlines.

Preferred Qualifications
- Experience implementing Bronze/Silver/Gold (Medallion) architecture.
- Strong knowledge of Delta Lake and Lakehouse architectures.
- Experience with Databricks optimization techniques including:
- Spark performance tuning
- Partitioning
- Caching
- Cluster optimization
- Job optimization
- Delta optimization
- Experience with Snowflake performance optimization, warehouse sizing, query optimization, and data loading strategies.
- Experience with data modeling and database design.
- Experience with orchestration tools such as Apache Airflow or equivalent.
- Knowledge of data governance, metadata management, lineage, and data quality frameworks.
- Experience with other cloud platforms and data technologies.
- Experience with AWS data services such as S3, Glue, EMR, Lambda, or Step Functions.
- Experience with Azure DevOps or additional DevOps tooling.
- Experience leveraging AI/ML and Generative AI solutions for data engineering automation.
- Experience working with enterprise-scale data platforms.
- Azure or Databricks certification is preferred.
- Experience in healthcare payer data domains, including Claims, Membership, Enrollment, Provider, Clinical, or Financial data, is a plus.

Technical SkillsPrimary
- Databricks
- Apache Spark
- PySpark
- Python
- Delta Lake
- Lakehouse Architecture

Secondary
- Snowflake
- SQL
- Data Modeling
- ETL/ELT

Azure Cloud
- Azure Data Factory
- ADLS Gen2
- Blob Storage
- SHIR
- Logic Apps
- Azure DevOps

AWS – Preferred
- S3
- AWS Glue
- EMR
- Lambda
- Step Functions

DevOps & Engineering
- Git
- GitHub
- GitHub Actions
- CI/CD
- Automated Testing
- Code Reviews
- Agile/Scrum

Data Quality & Operations
- Data Validation
- Data Reconciliation
- Data Profiling
- Data Quality
- Data Observability
- Monitoring
- Performance Optimization
- Cost Optimization

AI-Assisted Development
- GitHub Copilot
- Databricks Assistant
- ChatGPT
- Claude
- AI-assisted engineering automation

Work Location: Remote

📌 Data Engineer (India)
🏢 WebSenor InfoTech
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: data engineer (india) / india