07 Aug
|
Bridgestone - Global Capability Center
|
Bengaluru
07 Aug
Bridgestone - Global Capability Center
Bengaluru
Position Summary
The Lead Data Engineer, Enterprise Data Platform (EDP) is a hands-on technical leader responsible for building and operating the ingestion and curation layers of Bridgestone's multi-cloud Enterprise Data Platform. The EDP is a platform-as-a-service: business units build their own data products on top of it, within guardrails and governance standards this role helps define and enforce. The platform runs on a Databricks and Unity Catalog backbone spanning AWS and Azure, delivering curated data to business units, BI and analytics teams, and conversational query experiences.
This role writes code, designs pipelines, and sets the technical standard for a team of Data Engineers. This role partners closely with Platform Engineering, Solution Engineers, Product Owners, and business unit engineering teams. Deep Databricks expertise, strong data engineering fundamentals, and the ability to operate credibly in a large, matrixed enterprise are essential.
Job Duties
- Designs, builds, and operates industrialized data ingestion and transformation pipelines from source systems through the silver layer of a medallion architecture, using Databricks (Spark, Unity Catalog, workflows) as the core platform.
- Serves as the team's Databricks subject matter expert data modeling in Delta, job and cluster design, performance and cost tuning, and Unity Catalog governance patterns.
- Guides, coaches, and technically mentors other Data Engineers; conducts design and code reviews to ensure architecture, platform, and solution standards are followed.
- Builds and maintains pipelines across multiple clouds — primarily AWS (Glue, Step Functions, Redshift, Aurora, Transfer Family, S3) and Azure (Data Factory, ADLS) — and plans the migration of native cloud pipelines to Databricks as a destination technology.
- Defines and enforces the guardrails, standards, and reusable patterns that business units follow when building their own projects on the platform, balancing self-service velocity against governance requirements.
- Partners with the BI and analytics teams to deliver curated, well-modeled,
well-documented data products fit for reporting and for conversational query experiences such as Databricks Genie Space.
- Builds and maintains Databricks Genie Spaces over curated data, including the table selection, instructions, sample queries, and metric definitions that make natural-language answers trustworthy. Reviews what Genie returns, and when an answer is wrong, fixes the underlying data model rather than patching the question.
- Sets the team's working practice for AI-assisted engineering. Decides where AI tooling belongs in pipeline development, code review, testing, and documentation, where it does not, and what review is required before AI-assisted work reaches production.
- Drives data quality, lineage, observability, and access control practices in partnership with data governance and security teams.
- Participates in early design reviews and architectural strategy for data modeling, design, and implementation within the responsible data domains.
- Drives continuous improvement and automation — CI/CD, infrastructure as code, testing, and operational tooling — and solicits recommendations from the team to ensure the right solutions are implemented at the right time.
- Monitors production pipelines, leads triage and root cause analysis for data incidents, and documents RCAs and SOPs to standardize and stabilize daily operations.
- Coordinates with vendors and platform partners (Databricks, AWS, Microsoft) to resolve product defects and escalations.
- Other duties as assigned.
Required Qualifications
- Bachelor's degree in computer science, computer engineering, information systems, or equivalent work experience.
- Minimum of 8 years in data engineering or IT development, including 5 years of cloud data engineering and 3 years leading or technically guiding other engineers.
- Deep, hands-on Databricks experience — Spark (PySpark/SQL), Delta Lake, Unity Catalog,
Databricks workflows, and job/cluster performance and cost optimization. This is the core technology of the platform.
- Hands-on working experience with Databricks Genie, including building Genie Spaces from scratch: curating the underlying tables, writing space instructions and sample queries, defining metrics, and tuning a space until its answers hold up under real business questions.
- Robust data engineering fundamentals: data modeling, dimensional and medallion architecture patterns, partitioning, incremental and CDC ingestion, idempotency, schema evolution, and backfill strategy.
- Demonstrated hands-on pipeline development at enterprise scale — not solely oversight or coordination. This role builds.
- Production experience across multiple clouds, primarily AWS and Azure, including managed data services, storage, networking basics, and identity and access management.
- Expert-level Python and SQL.
- Experience with orchestration at scale (Databricks Workflows, Step Functions, Azure Data Factory, or equivalent) for complex, dependency-heavy pipelines.
- Working knowledge of data governance concepts, tools, and processes in a complex organizational environment — cataloging, lineage, data quality, and fine-grained access control.
- Experience with Git and CI/CD (Azure DevOps or equivalent) to promote code and release packages through environments.
- Experience working in large enterprise environments: matrixed teams, multiple business units with differing data maturity, change management, and formal release processes.
- Ability to perform end-to-end testing and debug issues across distributed cloud services.
- Excellent communication skills — able to explain technical constraints and trade-offs to engineers, business stakeholders, and leadership.
- Practical experience with Agile delivery practices.
- Uses AI tooling as part of a normal working day rather than as an occasional experiment. Candidates will be asked to walk through their own workflow in the interview: which tools, at what points in the work, what they verify before trusting the output, and what actually got faster as a result.
📌 Lead - Data Engineer (Bengaluru)
🏢 Bridgestone - Global Capability Center
📍 Bengaluru