02 Oct
|
Intellistaff
|
Bengaluru
02 Oct
Intellistaff
Bengaluru
Experience: 5 - 12+ Years
Location: Bangalore / Chennai / Hyderabad / Pune
Work Model: Hybrid
Hands-on Coding: 60 - 70%
Role OverviewWe are looking for a highly hands-on Data Engineer Lead with strong expertise in Python, PySpark, Apache Spark, Azure Databricks and Azure Data Engineering.Although this is a Lead-level position, the role is strongly hands-on, with approximately 6070% of the role focused on coding and technical implementation.The ideal candidate should combine architect-level design capabilities with strong hands-on Python and data engineering skills.Key
Responsibilities
- Design and develop scalable data ingestion, transformation and validation frameworks using Python, PySpark and Databricks.
- Develop reusable and metadata/configuration-driven data engineering frameworks.
- Design end-to-end solutions using Azure Databricks, ADLS, ADF, Azure SQL/Synapse and Delta Lake.
- Develop robust ETL/ELT pipelines supporting batch, streaming and incremental processing.
- Define architecture for data quality, validation, reconciliation, error handling, quarantine, replay and observability.
- Design solutions for CDC, idempotency, schema evolution, SCD and incremental processing.
- Develop reusable Python frameworks for ingestion, transformation and validation.
- Troubleshoot complex production issues and drive root-cause resolution.
- Optimize Spark/PySpark workloads involving partitioning, joins, shuffles, data skew and query execution.
- Establish Python coding standards, framework patterns, testing practices and CI/CD standards.
- Conduct technical design and code reviews.
- Provide technical leadership to data engineering teams.
- Translate business and functional requirements into scalable technical solutions.
- Collaborate with architects, product teams, business stakeholders and engineering teams.
Mandatory Technical Skills
- Python Advanced / Hands-on
- PySpark Strong
- Apache Spark Strong
- Azure Databricks
- Advanced SQL
- ETL/ELT
- Data ingestion and transformation framework development
- Azure Data Factory
- ADLS
- Azure SQL / Synapse
- Delta Lake
- Spark performance optimization
- Data quality and validation
- Production debugging and troubleshooting
- Git / CI-CD
Lead-Level ExpectationsThe candidate must be able to:
- Code independently and remain hands-on for 6070% of the role.
- Design reusable Python-based data engineering frameworks.
- Make architecture-level technical decisions.
- Explain technical trade-offs behind design decisions.
- Review code and identify performance/design issues.
- Troubleshoot complex production failures.
- Mentor engineers on Python, PySpark, Spark and Databricks.
- Own technical delivery from design through production implementation.
Critical Screening Areas AreaExpected LevelPythonAdvancedPython Framework DevelopmentMandatoryPySparkStrongApache SparkStrongDatabricksStrong Hands-onSQLAdvancedAzureStrong Hands-onETL FrameworksStrongSpark OptimizationCriticalProduction DebuggingCriticalArchitectureStrongData ModellingStrongTechnical LeadershipStrongGood to Have
- CDC / Kafka
- Unity Catalog
- Data Lakehouse / Medallion Architecture
- Oracle
- Azure Synapse
- Data archival / retention / purge
- Insurance / BFSI domain experience
- Enterprise-scale data platform experience
Mandatory Candidate CriteriaThe candidate must have personally developed Python frameworks for data ingestion, transformation, validation or similar data engineering use cases.Candidates who have only:
- Managed a team,
- Performed production monitoring,
- Used existing frameworks,
- Worked mainly on SQL,
- Used Databricks only for notebook execution,without robust hands-on Python framework development experience should not be considered.
Ideal Pipeline Architecture ExposureCandidates should be able to explain architectures such as:Source ADF/Ingestion ADLS Databricks/PySpark Delta SQL/Synapse Downstream ConsumptionThey should also be able to explain how they handled incremental loads, CDC, data quality, failures, reconciliation and performance optimization.
📌 Lead Azure Data Engineer (Bengaluru)
🏢 Intellistaff
📍 Bengaluru