14 Sep
|
fluid.live
|
Chennai
14 Sep
fluid.live
Chennai
- Data Engineering Lead
Experience: 8-15 Years
Role: Data Engineering Lead
Role Overview
We are looking for an experienced Data Engineering Lead with 8-15 years of experience in designing and delivering scalable data engineering solutions, leading data engineering teams, and driving data architecture and governance initiatives.
The ideal candidate should have strong hands-on expertise in SQL, Python, PySpark, ETL technologies, cloud data platforms, data warehousing, data modeling, orchestration, and distributed data processing. The candidate will be responsible for leading data engineering teams, designing scalable data architectures, optimizing data pipelines, and ensuring high standards of performance, quality, security, and governance.
Key Responsibilities
Data Engineering & Architecture
- Lead the design, development, and implementation of scalable and reliable data engineering solutions.
- Define and implement data architecture, data models, data pipelines, and data integration frameworks.
- Design and optimize batch and real-time data processing solutions across cloud environments.
- Establish technical standards, best practices, and reusable frameworks for data engineering.
- Drive the adoption of modern cloud data platforms and data engineering technologies.
ETL & Data Pipeline Development
- Design, develop, and optimize complex ETL/ELT pipelines using technologies such as Informatica PowerCenter, Databricks, AWS Glue, and Snowflake.
- Lead ETL modernization and migration initiatives from traditional platforms to cloud-based data platforms.
- Optimize pipeline performance, resource utilization, processing time, and overall data flow.
- Design robust orchestration and dependency management using Apache Airflow.
- Implement monitoring, logging, error handling, data validation, and recovery mechanisms for critical pipelines.
Cloud Data Platforms
- Design and implement data solutions across AWS, Azure, and GCP environments.
- Work with cloud-native data services and platforms including AWS, Azure, GCP, and Microsoft Fabric.
- Evaluate cloud technologies and recommend appropriate solutions based on scalability, performance, security, and cost.
- Lead cloud data platform migration, modernization, and optimization initiatives.
Data Warehousing & Data Modeling
- Design and implement enterprise-scale data warehouses, data marts, and analytical data platforms.
- Define and review dimensional, relational, and analytical data models.
- Work with platforms such as Amazon Redshift, SQL Server, and AWS RDS.
- Ensure data models are optimized for performance, scalability, analytics, and reporting requirements.
Data Architecture & Governance
- Define and enforce data architecture standards, governance frameworks, and engineering best practices.
- Establish standards for data quality, security, lineage, metadata, access control, and lifecycle management.
- Collaborate with architecture,
security, analytics, and business teams to ensure data solutions align with enterprise requirements.
- Drive initiatives around data governance, data quality, and regulatory/compliance requirements.
Data Cleanroom
- Lead the design and implementation of Data Cleanroom solutions for secure and privacy-conscious data collaboration.
- Define data processing and access patterns that enable secure analysis while protecting sensitive data.
- Work with stakeholders to establish governance, security, and privacy controls around cleanroom environments.
Performance Tuning & Scalability
- Identify and resolve performance bottlenecks across ETL pipelines, SQL queries, data processing jobs, and cloud data platforms.
- Optimize Spark/PySpark workloads, SQL queries, data storage, partitioning, and resource utilization.
- Design solutions capable of handling increasing data volumes and evolving business requirements.
- Establish performance benchmarks and continuously improve platform scalability and reliability.
Streaming & Messaging
- Design and implement real-time and event-driven data pipelines using Apache Kafka.
- Define appropriate messaging, ingestion, processing, and data delivery patterns.
- Ensure reliability, scalability, and fault tolerance of streaming data solutions.
Reporting & Analytics
- Collaborate with BI and analytics teams to support enterprise reporting and analytical requirements.
- Provide optimized and governed datasets for reporting platforms such as Tableau, Looker, and Power BI.
- Ensure data models and pipelines support consistent, reliable, and performant reporting.
CI/CD & Engineering Practices
- Implement and promote CI/CD practices for data engineering projects.
- Establish best practices around source control, automated testing, code reviews, deployment, and release management.
- Work with DevOps teams to automate data pipeline and platform deployments.
Leadership & Team Management
- Lead, mentor, and guide Data Engineers and other technical team members.
- Provide technical direction, conduct design and code reviews, and establish engineering standards.
- Break down complex requirements into technical solutions and assign tasks effectively across the team.
- Work closely with Product, Business, Analytics, Architecture, and DevOps teams.
- Drive technical decision-making and ensure timely delivery of high-quality data engineering solutions.
- Identify skill gaps within the team and support technical development and knowledge sharing.
- Take ownership of technical deliverables, project execution, and production stability.
Required Technical Skills
Programming Languages
- Strong hands-on experience with:
- SQL
- Python
- PySpark
- Unix Shell / Shell Scripting
- Scala
ETL / Data Engineering
- Strong experience with one or more of:
- Informatica PowerCenter
- Databricks
- AWS Glue
- Snowflake
- Strong understanding of ETL/ELT architecture and optimization.
Cloud Platforms
- Strong experience with one or more of:
- AWS
- Microsoft Azure
- Google Cloud Platform (GCP)
- Microsoft Fabric
- Strong understanding of cloud-based data architecture and services.
Databases
- Experience with:
- Amazon Redshift
- SQL Server
- AWS RDS
- Strong understanding of SQL optimization, database design, and performance tuning.
Data Warehousing & Modeling
- Strong expertise in:
- Data Warehousing
- Data Modeling
- Dimensional Modeling
- Fact and Dimension Tables
- Star and Snowflake Schemas
- Data Marts
- Enterprise Data Architecture
Orchestration
- Strong experience with Apache Airflow.
- Expertise in DAG design, scheduling, dependency management, monitoring, retries, and failure handling.
Messaging & Streaming
- Strong understanding of Apache Kafka and real-time/event-driven data architectures.
BI & Reporting
- Exposure to:
- Tableau
- Looker
- Power BI
CI/CD
- Good understanding of CI/CD concepts, version control, automated testing, deployment automation, and release management.
Leadership & Soft Skills
- Strong technical leadership and team management capabilities.
- Excellent problem-solving and analytical skills.
- Ability to translate business requirements into scalable technical solutions.
- Strong stakeholder management and communication skills.
- Ability to mentor engineers and drive technical excellence.
- Strong ownership and ability to manage multiple priorities in a fast-paced setting.
Preferred Qualifications
- Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, Data Science, or a related field.
- 815 years of experience in Data Engineering, Data Architecture, ETL, or related areas.
- Proven experience leading data engineering teams and large-scale data initiatives.
- Experience with enterprise-scale cloud data migrations and modernization programs.
- Experience working with Data Cleanrooms, Data Governance, and enterprise data architecture is highly preferred.
Key Technologies SQL | Python | PySpark | Unix Shell | Scala | Informatica PowerCenter | Databricks | AWS Glue | Snowflake | AWS | Azure | GCP | Microsoft Fabric | Redshift | SQL Server | AWS RDS | Tableau | Looker | Power BI | Apache Airflow | Apache Kafka | CI/CD | Data Cleanroom | Data Architecture | Data Governance | ETL Optimization | Data Warehousing | Data Modeling | Performance Tuning | Scalability | Team Leadership
📌 Lead Data Engineer (Chennai)
🏢 fluid.live
📍 Chennai