18 Sep
|
ZecData Technology
|
India
18 Sep
ZecData Technology
India
We are looking for an experienced Data Engineer with 4–9 years of hands-on experience, with strong expertise in AWS cloud data engineering, Python, SQL, ETL/ELT pipelines, data warehousing, and big data technologies.
The ideal candidate will be responsible for designing, developing, and maintaining scalable data pipelines and data platforms on AWS. The candidate should have strong experience working with large datasets, building reliable data workflows, optimizing data processing, and collaborating with data scientists, analysts, software engineers, and business teams.
Key Responsibilities
- Design, develop, and maintain scalable and reliable data pipelines on AWS.
- Build and maintain ETL/ELT workflows for structured and unstructured data.
- Develop data ingestion pipelines from multiple sources including databases, APIs, files, and streaming systems.
- Transform, clean, validate, and process large volumes of data.
- Develop efficient and optimized SQL queries for data processing and analytics.
- Build and maintain cloud-based data warehouses and data lakes.
- Work closely with Data Scientists, Data Analysts, BI teams, and Software Engineers to provide high-quality data.
- Implement data quality, validation, monitoring, and error-handling mechanisms.
- Optimize data pipelines for performance, scalability, reliability, and cost.
- Troubleshoot pipeline failures and production data issues.
- Develop reusable frameworks and components for data engineering workflows.
- Implement security, access control, encryption, and governance best practices.
- Participate in architecture discussions and contribute to technical design decisions.
- Write technical documentation and maintain data pipeline and system documentation.
- Follow software engineering best practices including Git, code reviews, testing, CI/CD, and Agile methodologies.
Required Technical SkillsProgramming
- Strong hands-on experience with Python.
- Strong proficiency in SQL.
- Good understanding of data structures, algorithms, and software engineering principles.
- Experience writing reusable and production-quality code.
AWS Data Engineering
Strong hands-on experience with multiple AWS services, preferably:
- Amazon S3 – Data Lake and object storage
- AWS Glue – ETL, Data Catalog, and Crawlers
- Amazon Redshift – Data warehousing
- Amazon Athena – SQL-based analytics
- AWS Lambda – Serverless data processing
- Amazon EMR – Big data processing
- Amazon RDS – Relational databases
- Amazon Kinesis – Real-time/streaming data
- AWS Step Functions – Workflow orchestration
- Amazon CloudWatch – Monitoring and logging
- IAM – Access management and security
Candidates should have strong practical experience with AWS services rather than only theoretical knowledge.
ETL / Data Pipelines
- Strong experience developing ETL/ELT pipelines.
- Experience with batch and incremental data processing.
- Understanding of data ingestion, transformation, validation, and loading processes.
- Experience handling pipeline failures, retries, logging, and monitoring.
- Knowledge of scheduling and workflow orchestration.
Big Data Technologies
- Hands-on experience with Apache Spark / PySpark.
- Experience processing large datasets using distributed computing frameworks.
- Understanding of Spark optimization, partitioning, joins, caching, and performance tuning.
- Exposure to Hadoop/Hive is a plus.
Data Warehousing
- Strong understanding of data warehouse concepts.
- Experience with dimensional modeling.
- Knowledge of fact and dimension tables, star schema, and snowflake schema.
- Experience with Amazon Redshift or other cloud data warehouses.
- Understanding of partitioning, indexing, and query optimization.
Databases
Experience with one or more of:
- PostgreSQL
- MySQL
- SQL Server
- Oracle
- MongoDB
- DynamoDB
Strong understanding of relational database concepts and query optimization is required.
Data Lake & Modern Data Architecture
- Experience working with AWS Data Lake architectures.
- Strong understanding of data lake and data warehouse concepts.
- Experience working with formats such as Parquet, JSON, CSV, and Avro.
- Understanding of partitioning and schema evolution.
- Knowledge of data governance, lineage, and data quality practices is preferred.
DevOps & CI/CD
- Experience with Git/GitHub/GitLab/Bitbucket.
- Understanding of CI/CD pipelines.
- Exposure to Docker is preferred.
- Experience with infrastructure/configuration tools such as Terraform or CloudFormation is a plus.
- Familiarity with Linux/Unix environments.
Good to Have
- Experience with Apache Airflow or other workflow orchestration tools.
- Experience with Kafka/Kinesis and real-time data pipelines.
- Knowledge of AWS Lake Formation and AWS Glue Data Catalog.
- Experience with dbt or similar transformation frameworks.
- Knowledge of Terraform.
- Experience with AWS CloudWatch and data pipeline monitoring.
- Understanding of data security and AWS IAM.
- Experience with data governance and data quality frameworks.
- Exposure to machine learning/data science workflows.
Required Experience
- 4–9 years of professional experience in Data Engineering, Big Data Engineering, or a related field.
- Strong hands-on experience with AWS Data Engineering.
- Solid proficiency in Python and SQL.
- Hands-on experience developing and supporting ETL/ELT pipelines.
- Experience working with AWS Glue, S3, Redshift, Athena, and preferably Spark/PySpark.
- Experience working with production data pipelines and troubleshooting data-related issues.
- Strong analytical and problem-solving skills.
Educational Qualification
- Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, Data Science, or a related field.
- Equivalent practical experience may also be considered.
Preferred Tech Stack
Python | SQL | AWS | S3 | AWS Glue | Redshift | Athena | Lambda | EMR | Kinesis | Step Functions | CloudWatch | IAM | PySpark | Apache Spark | Airflow | Kafka | PostgreSQL | MySQL | Docker | Terraform | Git | CI/CD
Pay: ₹90,000.00 - ₹110,000.00 per month
Benefits:
- Provident Fund
Work Location: Remote
📌 Data Engineer - AWS(exp:4-9yrs) (India)
🏢 ZecData Technology
📍 India