12 Sep
|
CloudBoson
|
Noida
Role & responsibilities
Design, build, and optimize ETL/ELT pipelines using PySpark on Azure Databricks.
Develop real-time and near-real-time ingestion pipelines using Spark Structured Streaming and Azure Event Hubs
Develop and tune complex SQL models in BigQuery for analytics and reporting
Build ingestion pipelines from APIs and third-party sources into Azure Data Lake / BigQuery
Implement incremental loads, CDC, and MERGE/upsert patterns for large datasets.
Experience on data lake/warehouse solutions across GCP and Azure
Build dashboards and reporting layers in Looker Studio
Lead code reviews, mentor engineers, and drive Git-based CI/CD practices for data pipelines
Partner with product/business teams to translate requirements into scalable data models
Optimize cost & performance of already running queries
Carefully plan and implement data load requirement.
Regular monitoring of the data pipelines and systems and maintains the uptime of the overall data availability.
Co-ordinate with stake holders and provide relevant data and business insights.
Required Skills:
5+ years in data engineering
Robust SQL (joins, window functions, query optimization)
Well versed in Python and PySpark (Azure Databricks)
Spark Structured Streaming for real-time data processing
Azure Event Hubs for event ingestion and streaming architecture
BigQuery , AWS DynamoDB, Azure Storage Tables, Azure SQL, Azure Event Hub.
Azure Data Lake (ADLS Gen2)
Multi-cloud experience: GCP + Azure
Looker Studio for BI/reporting
Git version control
REST API integration (auth, pagination, error handling)
Positive to have:
Orchestration tools (Airflow, Databricks Workflows)
Data quality/governance frameworks
Cloud cost optimization experience
📌 Senior Data Engineer Noida
🏢 CloudBoson
📍 Noida