30 Jul
|
Qapitol QA
|
Uttar Pradesh
30 Jul
Qapitol QA
Uttar Pradesh
Data Engineer
Location: Noida, Uttar Pradesh (Work from Office)
Experience: 3 4 years Employment Type: Full-time
About the Role
We are looking for a hands-on Data Engineer to build and operate production data pipelines on our Azure-based lakehouse. You will own ingestion and orchestration through Azure Data Factory, develop transformation logic in Databricks on a medallion architecture (Bronze Silver Gold), and deliver clean, analytics-ready datasets that power BI and decision-making across the business. You will work closely with business analysts and the analytics lead, with real ownership of pipelines end to end.
Key Responsibilities
- Design, build, and maintain batch ETL/ELT pipelines using Azure Data Factory and Azure Databricks (PySpark / Spark SQL)
- Develop and maintain Bronze, Silver, and Gold layers on Delta Lake schema validation, type casting, deduplication, timestamp standardization, and dataquality checks
- Ingest data from transactional systems (MySQL/OMS, Zoho applications, ERP, flat files, REST APIs) into the lakehouse
- Build Gold-layer fact and dimension tables supporting order fulfillment, sales, inventory, margin/COGS, and credit-note analytics
- Production-harden notebooks and jobs parameterization (widgets/config-driven), required-column validation, safe casting, error handling, and logging
- Optimize Spark workloads for performance and cost (partitioning, caching, cluster sizing, Delta OPTIMIZE / Z-ORDER)
- Monitor scheduled jobs, debug pipeline failures, and handle schema drift and baddata scenarios
- Serve reliable, well-modeled datasets to downstream BI tools (Tableau, Zoho Analytics, Metabase/Superset)
- Partner with analysts and business stakeholders to translate reporting requirements into robust data models
- Maintain documentation data dictionaries, pipeline runbooks, and lineage notes
Must-Have Skills
- 3 4 years of hands-on data engineering experience in a production environment
- Strong, demonstrable experience on the Azure data stack Azure Data Factory, Azure Databricks, Azure Synapse, and Microsoft Fabric. Candidates who have worked on these tools will be given robust preference.
- Candidates without direct Azure exposure must have equivalent hands-on depth on similar alternatives e.g., AWS (Glue, EMR, Redshift), GCP (Dataflow/Dataproc, BigQuery, Cloud Composer), Snowflake with Airflow/dbt, or self-managed Spark platforms along with willingness to work on Azure
- Strong SQL: complex joins, window functions, CTEs, and aggregation logic on large datasets
- Proficiency in Python and PySpark / Spark SQL for data transformation
- Solid understanding of lakehouse and medallion architecture concepts; hands-on with Delta Lake (or Iceberg/Hudi equivalents)
- Data modeling fundamentals: star schema, fact/dimension design, slowly changing dimensions
- Version control with Git and basic CI/CD awareness Good to Have
- Hands-on exposure to Microsoft Fabric (Lakehouse, Data Pipelines, OneLake)
- Unity Catalog or other data governance / access-control experience
- Streaming fundamentals: Event Hubs, Kafka, or Spark Structured Streaming
- Experience integrating third-party REST APIs as data sources
- Familiarity with BI tools such as Tableau, Power BI, Metabase, Superset, or Zoho Analytics
- Domain exposure to B2B distribution, e-commerce, retail, or supply chain
- Relevant Microsoft certification (e.g., DP-700 Fabric Data Engineer Associate)
Education
B.E. / B.Tech / MCA / M.Sc. in Computer Science, IT, or a related field or equivalent practical experience.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Data Engineer (Uttar Pradesh)
🏢 Qapitol QA
📍 Uttar Pradesh