30 Sep
|
Virtual Connect Solutions
|
Bengaluru
30 Sep
Virtual Connect Solutions
Bengaluru
Role Summary
We are hiring a Lead Data Engineer for a strongly Python-first profile, complemented by solid depth across Databricks, data modelling, T-SQL,
and data architecture. You will design, build, and own production-grade Python codebases and frameworks at the core of our Databricks
Lakehouse platform, while also setting technical direction on pipeline architecture, data modelling, and platform performance, and mentoring a team of engineers.
Key Responsibilities
- Design, build, and own production-grade Python codebases for the data platform —
- ETL/ELT pipelines, reusable internal packages,
schema validation, and testing frameworks (pytest, CI/CD).
- Write clean, well-tested, performant Python: apply OOP, concurrency (multiprocessing/asyncio), decorators, and generators to solve real
pipeline-scale problems.
- Profile and optimize slow Python pipelines end to end — from algorithmic/code-level fixes to migrating heavy workloads onto
Databricks/Spark when Python alone won't scale.
- Set and enforce Python engineering standards across the team: code reviews, packaging/versioning, structured logging, and exception
handling patterns.
- Design and deliver Bronze/Silver/Gold Medallion pipelines on Databricks —
- Delta Lake, Structured Streaming/Auto Loader, and Unity
Catalog governance.
- Own dimensional data modelling decisions — star/snowflake schema design and SCD strategy — for the platform's core data models.
- Write and review advanced T-SQL for platform workloads — window functions, MERGE/upsert logic, and execution-plan-based
performance tuning.
- Partner with Data Architects and stakeholders on ETL vs. ELT, orchestration, and cost/performance trade-offs across the platform; mentor
Data Engineers and Senior Data Engineers through design and code reviews. Must-have Qualifications
8–12 years in data engineering, including at least 2 years in a technical leadership or mentoring capacity. This role carries a strong Python emphasis,
Complemented By Hands-on Depth In
Python &
- Data Engineering (Primary Focus)
- Strong hands-on Python: OOP, generators, decorators, concurrency (multiprocessing/asyncio), exception handling, and unit testing
(pytest).
- Experience building production ETL/ELT pipelines, schema validation, and reusable internal Python packages/frameworks.
- Proven ability to profile, debug, and optimize Python code for performance at data-pipeline scale.
Databricks &
- Lakehouse
- Deep, hands-on Databricks experience: Delta Lake (ACID, time travel, OPTIMIZE/ZORDER/VACUUM), Structured Streaming, Auto Loader,
and Unity Catalog.
- Experience tuning Spark jobs at scale — data skew, cluster sizing/autoscaling, and cost optimization.
Data Modelling &
- Warehousing
- Solid grounding in dimensional modelling: star/snowflake schema, SCD Types 1/2/3, fact/dimension grain, and conformed dimensions.
- Working knowledge of Data Vault modelling and Kimball vs. Inmon trade-offs.
T-SQL
- Advanced T-SQL: window functions, MERGE/upsert patterns, indexing strategy, and execution-plan-based query tuning.
Data Architecture
- Proven experience designing Medallion (Bronze/Silver/Gold) architectures and evaluating ETL vs. ELT and batch vs. streaming trade-offs.
- Experience with orchestration (Databricks Workflows, Airflow, or ADF) on Azure or AWS.
GOOD TO HAVE
- PCEP / PCAP (Python Institute) or equivalent Python certification.
- Databricks Certified Data Engineer Professional (or equivalent) certification.
- Experience with data governance/compliance (GDPR/HIPAA) and cost-governance (FinOps) practices.
- Prior experience formally leading or mentoring a team of 3+ engineers.
EDUCATION Bachelor's or Master's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
Soft Skills
Strong stakeholder communication; able to explain technical trade-offs to both engineering and business audiences; comfortable owning ambiguous, enterprise-scale design decisions end to end.
Skills: python,etl,sql
📌 Lead Data Engineer/Associate Data Architect (Bengaluru)
🏢 Virtual Connect Solutions
📍 Bengaluru