20 Sep
|
DataZymes
|
Bengaluru
20 Sep
DataZymes
Bengaluru
Roles and Responsibilities
- Create and maintain optimal data pipeline architecture for ETL/ELT into structured data
- Assemble large, complex data sets that meet business requirements and create multi-dimensional modelling like Star Schema and Snowflake Schema
- Expert level experience in creating scalable data warehouse including Fact tables, Dimensional tables and ingest datasets into cloud-based tools.
- Identify, design, and implement internal process improvements including automating manual processes and optimizing data delivery
- Collaborate with stakeholders to ensure seamless integration of data with internal data marts, enhancing advanced reporting.
- Setup and maintain data ingestion, streaming, scheduling, and job monitoring automation using AWS services including Lambda, Code Pipeline, Glue, S3, and Redshift.
- Build analytics tools that utilize the data pipeline to provide actionable insight into customer acquisition and operational efficiency.
- Work with stakeholders to assist with data-related technical issues and support their data infrastructure needs.
- Utilize GitHub for version control, code collaboration, and repository management
- Create data tools for analytics and data scientist team members
- Ensure data privacy and compliance with relevant regulations when handling customer data.
- Maintain data quality and consistency within the application
Requirements
1Strong Senior Data Engineer Profile with AWS data-warehousing and pipeline expertise
2Mandatory (Experience):
Must have at least 4+ years of hands-on data engineering experience with the recent 2+ years in AWS cloud data warehouses and AWS cloud services
3Mandatory (Tech skill 1): Must have advanced SQL knowledge and hands-on experience with relational databases and query authoring, plus a cloud data warehouse like AWS Redshift.
4Mandatory (Tech skill 2): Must have expert-level experience creating scalable data warehouses — Fact tables, Dimensional tables, and ingesting datasets into cloud-based tools.
5Mandatory (Tech skill 3): Must have strong multi-dimensional data modelling experience — Star Schema, Snowflake Schema, normalization/de-normalisation, joins, OLAP cube modelling, and schema evolution while maintaining data integrity.
6Mandatory (Tech skill 4): Must have hands-on experience creating and maintaining optimal ETL/ELT data pipeline architecture into structured data.
7Mandatory (Tech skill 5): Must have experience setting up and maintaining data ingestion, streaming, scheduling, and job-monitoring automation using AWS services — Lambda, Glue, S3, Redshift, and Code Pipeline (CI/CD)
8Mandatory (Tech skill 6): Must have experience building and optimizing big-data pipelines, architectures, and datasets,
including data compression into PARQUET and SQL performance tuning.
9Mandatory (Tech skill 7): Must have experience with GitHub for version control, code collaboration, code reviews, branching strategies, and continuous integration.
10Mandatory (Tech skill 8): Must have solid analytical skills across structured and unstructured datasets, with experience performing root-cause analysis to answer business questions and identify improvements.
11Mandatory (Tech skill 9): Must have experience collaborating with cross-functional teams and Global IT to gather requirements and align work with business objectives, and ensuring data privacy/compliance (e.G. GDPR).
12Mandatory (Tech skill 10): Working knowledge of message queuing, stream processing, and highly scalable big-data stores;
familiarity with Agile working models.
13Mandatory (Project Alignment): Current or recent role must clearly demonstrate hands-on work with AWS data platforms and ETL pipeline implementation—resume must describe specific projects, tools used, and candidate's direct scope
14Mandatory (Education): Bachelor’s or master’s degree in Technology and Computer Science background
15Mandatory (Availability): Must be an immediate joiner or currently serving notice period, able to start within the next week
16Mandatory (Note 1) : Role is Hybrid, WFH flexibility as well up to 6 days a month
17Mandatory (Note 2): CTC is inclusive of 20% variable
18Preferred (Domain): Healthcare/Pharmaceutical/Life Sciences industry experience
📌 Hiring: Senior Data Engineer (Bengaluru)
🏢 DataZymes
📍 Bengaluru