2. Key Responsibilities
- Design, develop and maintain robust, scalable ETL/ELT pipelines on Databricks using Python, PySpark and Spark SQL, following medallion (Bronze/Silver/Gold) architecture principles.
- Independently analyze business requirements, translate them into technical designs and solution options, and drive them to completion with minimal handholding.
- Design and implement conceptual, logical and physical data models (dimensional and 3NF) for analytics, reporting and master data use cases.
- Build and manage data ingestion from enterprise sources (SAP S/4HANA, Salesforce, flat files, APIs, streaming) into the Lakehouse.
- Develop and operate workloads on AWS (S3, Lambda, IAM, networking fundamentals) integrated with the Databricks platform.
- Implement data quality checks, validation rules, reconciliation and error-handling frameworks; prepare test scenarios and perform thorough unit and business-level validation of own deliverables before handover.
- Apply Unity Catalog-based governance: access controls, lineage, cataloguing and aligned handling of personal data.
- Optimize Spark jobs and Delta tables for performance and cost (partitioning, Z-ordering, cluster sizing, job orchestration).
- Contribute to CI/CD practices for data pipelines (Git-based version control, code reviews, automated deployment) and produce clear technical documentation.
- Support BI and MDM workstreams by delivering curated, well-modelled datasets for Power BI and Informatica IDMC consumption.
- Provide L2/L3 support for deployed pipelines, troubleshoot production issues and drive root-cause resolution.
3. Primary (Must-Have) Skills
- Python: Strong, hands-on programming for data engineering: clean, modular, well-tested code.
- SQL: Advanced SQL for complex transformations, analysis and performance optimization.
- Apache Spark / PySpark: Deep understanding of Spark architecture, Data frame API, Spark SQL, performance tuning and debugging of distributed jobs.
- Data Modelling:
Solid experience in dimensional modelling (star/snowflake), normalized modelling, slowly changing dimensions and canonical/master data models.
- Databricks: Proven project experience with Databricks workspaces, Delta Lake, Delta Live Tables/Jobs, Workflows and Unity Catalog.
- AWS: Working proficiency with core AWS data services (S3, Glue, Lambda, IAM, CloudWatch) and integration with Databricks.
4.
Other Required
Skills
- Experience with orchestration tools (Databricks Workflows or equivalent).
- Git-based development workflow, code review discipline and CI/CD for data pipelines (Azure DevOps or similar).
- Exposure to integrating with enterprise systems such as SAP S/4HANA and Salesforce (APIs, CDC, extractors) is highly desirable.
- Familiarity with Power BI datasets/semantic models and with MDM concepts (e.g., Informatica IDMC) is an advantage.
- Exposure to streaming technologies (Structured Streaming, Kafka/Kinesis) is a plus.
5. Ways of Working & Behavioral Expectations
- Independent delivery: Works independently end-to-end: clarifies requirements early, proposes designs, and delivers complete, tested solutions without repeated review cycles.
- Ownership: Takes accountability for quality and timelines; proactively prepares test scenarios and validates business logic before submitting work.
- Speed of comprehension: Grasps current requirements and domain logic quickly and converts them into working solutions at the pace the project demands.
- Communication: Communicates progress, risks and blockers clearly and early; documents solutions to a standard others can maintain.
- Teamwork: Collaborates effectively with BI developers, MDM consultants, business SMEs and vendor teams in a multi-vendor environment.
6. Qualifications
- Bachelor's degree in Computer Science, Engineering or a related field (Master's preferred).
- Relevant certifications are an advantage: Databricks Certified Data Engineer (Associate/Professional), AWS Certified Data Engineer / Solutions Architect.
📌 AWS Data Engineer - G Noida (Delhi)
🏢 Coforge
📍 Delhi