AWS Data Engineer (Hyderabad)

AWS Data Engineer (Hyderabad)

17 Sep
|
Qentelli
|
Hyderabad

17 Sep

Qentelli

Hyderabad

Job Summary

We are seeking a Data Engineer to ensure the reliability, accuracy, performance, scalability, and automation of our data platforms and pipelines. This role focuses on validating and supporting end-to-end data ingestion, processing, streaming, analytics workflows, and infrastructure automation built on AWS, Kafka, SQL, Databricks, Python, and Infrastructure as Code (IaC) technologies. You will work closely with Data Engineers, Platform Engineers, Analytics teams, and DevOps teams to embed quality into the data lifecycle while ensuring trusted data, resilient pipelines, automated infrastructure provisioning, and production-ready systems.

Key Responsibilities

- Validate batch and streaming data pipelines for correctness, completeness, consistency, and timeliness.

- Create and maintain data quality checks (nulls, duplicates, schema drift, referential integrity).

- Verify business rules and transformations using SQL-based validations.

- Design, develop, and execute test strategies for Databricks-based data pipelines and analytics workflows.

- Establish data reconciliation and end-to-end traceability between source and downstream systems.

- Test ETL/ELT pipelines built using AWS services (Glue, Lambda, EMR, Step Functions).

- Validate transformations written in SQL and Python.

- Ensure correctness across data ingestion, enrichment, aggregation, and publishing layers.

- Test reprocessing, backfills, and historical data loads.

- Validate ETL/ELT processes built using Apache Spark (PySpark/Scala) in Databricks.

- Validate Kafka-based streaming pipelines for data integrity, ordering, and exactly-once/at-least-once semantics.

- Test producer and consumer logic, serialization formats (Avro, JSON, Protobuf).

- Validate topic configurations, partitions, offsets,



retention policies, and schema changes.

- Simulate and test late arrivals, duplicate events, and consumer failures.

- Test data workflows using AWS S3, Glue, Lambda, Redshift, Athena, Kinesis, DynamoDB, or similar services.

- Validate IAM roles, permissions, and secure data access.

- Verify data lifecycle policies, encryption, and storage optimizations.

- Build and maintain automated data testing frameworks using Python.

- Develop reusable test utilities, fixtures, and synthetic datasets.

- Integrate data tests into CI/CD pipelines for pre-merge, scheduled, and post-deployment validation.

- Enable automated alerts for data quality failures.

- Validate pipeline performance for large-scale datasets.

- Test throughput, latency, and concurrency under peak workloads.

- Validate retry logic, error handling, idempotency, and recovery mechanisms.

- Perform soak, regression, and failover testing.

- Validate data pipeline metrics, logs, and alerts using CloudWatch, Prometheus, Grafana, or equivalent tools.

- Partner with teams to define data SLAs and SLOs.

- Participate in incident response, root-cause analysis, and postmortems related to data quality issues.

- Validate infrastructure provisioning and deployment processes using Infrastructure as Code (IaC) tools such as Terraform, CloudFormation, Pulumi, and Ansible.





- Review and test IaC templates/modules to ensure consistency, security, compliance, and repeatability across environments.

- Support automated provisioning of data platform components including Databricks, Kafka, AWS resources, networking, IAM roles, and storage services through IaC frameworks.

- Collaborate with DevOps and Platform Engineering teams to integrate IaC validation into CI/CD pipelines and release processes.

- Validate infrastructure changes, configuration drift detection, workplace deployments, and disaster recovery procedures.

Required Qualifications

- 5+ years of experience in Data Engineering, Data Quality Engineering, or Data Platform Engineering roles.

- Strong hands-on experience with SQL for complex data validation and analysis.

- Proficiency in Python for automation, scripting, and data validation.

- Experience testing and supporting data pipelines and ETL/ELT workflows.

- Hands-on experience with Kafka or other streaming platforms.

- Solid understanding of AWS data services (S3, Glue, Redshift, Lambda, Athena, Kinesis, DynamoDB, etc.).

- Experience working with large datasets and distributed systems.

- Strong debugging, analytical, and problem-solving skills.

- Hands-on experience with Infrastructure as Code (IaC) tools such as Terraform, CloudFormation, Pulumi, and/or Ansible.

- Experience provisioning and managing cloud infrastructure through code in AWS environments.

- Knowledge of CI/CD practices and tools such as Jenkins, GitHub Actions, GitLab CI/CD, or AWS CodePipeline.

- Experience with version control systems such as Git and infrastructure deployment best practices.

- Understanding of cloud security, IAM, networking, and environment management through IaC.

📌 AWS Data Engineer (Hyderabad)
🏢 Qentelli
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: aws data engineer (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: aws data engineer (hyderabad) / hyderabad