06 Aug
|
Su0026P Global Market Intelligence
|
Chennai
06 Aug
Su0026P Global Market Intelligence
Chennai
We are seeking a hands-on Senior Data Engineer / Data Platform Lead to support our offshore team in India. This role focuses on executing data pipeline onboarding, migration, and support activities within our enterprise data platform. You will work closely with the Databricks team and internal stakeholders to ensure smooth transition, ongoing refinement, and reliable operation of data pipelines and integrations in a Databricks-on-AWS lakehouse environment.
Databricks on AWS positions data engineering around governed ingestion, transformation, and delivery of high-quality data for analytics and downstream use cases.
This is a delivery-focused role requiring strong technical expertise in data engineering, with responsibilities for supporting data onboarding, pipeline maintenance, platform support, and cloud-native operational excellence. You will lead a small team of engineers, providing technical guidance and ensuring high-quality execution across scalable, secure, and governed data workflows leveraging Amazon S3, AWS Glue, AWS IAM, and AWS Lake Formation alongside Databricks capabilities such as Delta Lake, Unity Catalog, and open table format interoperability where appropriate.
Key Responsibilities Data Pipeline Support Migration
- Assist in the migration of existing data pipelines to the Databricks platform, following established patterns and runbooks.
- Support the ongoing operation, troubleshooting, and optimization of pipelines post-migration.
- Implement and follow engineering best practices for data ingestion, transformation, and monitoring.
- Ensure pipelines are reliable, performant, and maintainable.
- Support data ingestion and storage patterns using Amazon S3 as the core cloud object store for lakehouse data layers, with Databricks processing and transformation built on top of that storage foundation.
- Contribute to migration and optimization activities involving AWS Glue for metadata cataloging, schema discovery, or ETL interoperability where required across the enterprise AWS data ecosystem.
- Support analytical integration patterns with downstream platforms such as Amazon Athena when business consumers require governed access beyond Databricks-native consumption paths.
Data Onboarding Asset-Agnostic Support
- Support offshore teams in onboarding recent data assets to the enterprise platform using standardized, asset-agnostic processes.
- Help teams prepare, validate, and publish data assets, ensuring consistency and compliance with platform standards.
- Document onboarding procedures and assist in resolving onboarding issues.
- Help standardize onboarding into Bronze, Silver,
and Gold data layers and support practical lakehouse patterns for ingestion, curation, and consumption in Databricks on AWS.
- Support metadata-driven onboarding workflows using AWS Glue Data Catalog and Databricks governance capabilities to improve discoverability, schema consistency, and reusable ingestion patterns.
- Assist with secure data onboarding by aligning access permissions and environment controls through AWS IAM and AWS Lake Formation, alongside Unity Catalog permissions and governance policies for managed access to datasets.
Platform Integration Support
- Support integration of data pipelines with the enterprise data mastering platform.
- Assist in data quality checks, metadata management, and reconciliation activities.
- Help troubleshoot and resolve issues related to data ingestion and mastering workflows.
- Support interoperability across AWS and Databricks services for ingestion, transformation, and governed publishing of mastered or standardized datasets.
- Contribute to table design and data publication patterns using Delta Lake and, where applicable, Apache Iceberg to support interoperable lakehouse consumption models.
- Assist with auditability and lineage-oriented controls through metadata, access management, and governed publishing standards.
Team Collaboration Delivery
- Work closely with the offshore team to prioritize tasks, meet delivery deadlines, and maintain quality.
- Provide technical guidance, review code, and mentor junior team members.
- Communicate progress, risks, and issues effectively to stakeholders.
- Collaborate with platform, governance, and cloud teams to align engineering delivery with enterprise AWS standards, security controls, and support processes.
- Promote practical engineering discipline across source control, peer review, CI/CD, release management, and environment consistency.
Operational Excellence
- Follow established operational procedures for deployment, monitoring, and incident management.
- Support production support activities, including issue resolution and performance tuning.
- Contribute to documentation and knowledge sharing.
- Support cloud-native monitoring and operational visibility using tools such as Amazon CloudWatch, logging, alerting, and job observability for production data pipelines.
- Contribute to deployment automation and release reliability using CI/CD, infrastructure as code, and cloud-native orchestration patterns, including AWS services such as AWS Step Functions, AWS Lambda, or equivalent enterprise tooling where applicable.
- Help maintain secure and compliant platform operations through access control, secrets handling, environment segregation, and operational readiness practices.
Required Qualifications
- 5+ years of experience in data engineering, pipeline development, or related roles.
- Proven experience supporting data onboarding, migration, and pipeline operations.
- Hands-on expertise with data engineering tools such as Spark, SQL, Python, or similar.
- Experience working with Databricks or similar cloud data platforms.
- Hands-on experience with relevant AWS data engineering services, especially Amazon S3, AWS Glue, and core cloud security concepts such as AWS IAM; familiarity with AWS Lake Formation, Amazon Athena, AWS Lambda, or Amazon Kinesis is strongly preferred for enterprise-scale data platform support. Industry-standard AWS data engineering stacks commonly center on these services for storage, ETL, governance, analytics, and streaming.
- Familiarity with data quality, metadata, lineage, and basic data governance concepts.
- Working knowledge of Delta Lake, Databricks Unity Catalog, and exposure to Apache Iceberg or similar open table/storage formats for lakehouse architectures.
- Understanding of secure data access, governance, and platform controls across Databricks and AWS environments.
- Familiarity with software engineering best practices, including version control, automated testing, CI/CD, infrastructure as code, and monitoring.
- Strong problem-solving skills and ability to troubleshoot production issues.
- Good communication skills and ability to collaborate with global teams.
Preferred Qualifications
- Experience supporting enterprise data onboarding and integration workflows.
- Knowledge of Delta Lake, Apache Iceberg, or similar storage formats.
- Exposure to data mastering or data governance frameworks.
- Experience with AWS Lake Formation, AWS Glue Data Catalog, Amazon Athena, Amazon Kinesis, or AWS Step Functions in support of modern cloud data platforms.
- Experience supporting Databricks on AWS lakehouse implementations, including governed data access with Unity Catalog and cloud storage patterns built on Amazon S3.
- Proficiency in Python and SQL, with experience in PySpark or similar distributed processing frameworks.
- Familiarity with infrastructure as code and cloud deployment automation.
- Certifications in Databricks, AWS, or related data platforms.
📌 Lead, Data Engineering (Chennai)
🏢 Su0026P Global Market Intelligence
📍 Chennai