06 Aug
|
Resource Capita It Services
|
Hyderabad
06 Aug
Resource Capita It Services
Hyderabad
We are seeking a Data Quality Engineerto ensure the reliability, accuracy, performance, and scalability of our data platforms and pipelines. This role focuses on validating end-to-end data ingestion, processing, streaming, and analytics workflowsbuilt on AWS, Kafka, SQL, Data Bricks and Python. You will work closely with Data Engineers, Platform Engineers, Analytics teams, and DevOps to embed quality into the data lifecycle ensuring trusted data, resilient pipelines, and production-ready systems.
Key Responsibilities:
- Validate batch and streaming data pipelinesfor correctness, completeness, consistency, and timeliness.
- Create and maintain data quality checks(nulls, duplicates, schema drift, referential integrity).
- Verify business rules and transformationsusing SQL-based validations.
- Design, develop, and executetest strategiesfor Databricks-based data pipelines and analytics workflows
- Establish data reconciliation and end-to-end traceabilitybetween source and downstream systems.
- Test ETL/ELT pipelines built using AWS services(Glue, Lambda, EMR, Step Functions).
- Validate transformations written in SQL and Python.
- Validate ETL/ELT processesbuilt using Apache Spark (PySpark/Scala) in Databricks.
- Validate Kafka-based streaming pipelinesfor data integrity, ordering, and exactly-once/at-least-once semantics.
- Test producer and consumer logic, serialization formats (Avro, JSON, Protobuf).
- Test data workflows using AWS S3, Glue, Lambda, Redshift, Athena, Kinesis, DynamoDB, or similar services.
- Build and maintain automated data testing frameworksusing Python.
- Develop reusable test utilities, fixtures, and synthetic datasets. Integrate data tests into CI/CD pipelinesfor pre-merge, scheduled, and post-deployment validation.
- Enable automated alerts for data quality failures.
- Validate pipeline performance for large-scale datasets.
- Test throughput, latency, and concurrency under peak workloads.
- Validate retry logic, error handling, idempotency, and recovery mechanisms.
- Perform soak, regression, and failover testing Validate data pipeline metrics, logs, and alerts using CloudWatch, Prometheus, Grafana, or equivalent tools.
Required Qualifications:
- 7+ yearsof experience in QA, SDET, or Data Quality Engineering roles.
- Solid hands-on experience with SQLfor complex data validation and analysis.
- Proficiency in Pythonfor test automation and data validation.
- Experience testing data pipelines and ETL/ELT workflows.
- Hands-on experience with Kafka or other streaming platforms.
- Solid understanding of AWS data services(S3, Glue, Redshift, Lambda, Athena, etc.).
- Experience working with large datasets and distributed systems.
- Strong debugging, analytical, and problem-solving skills.
📌 Data Quality Engineer | Hyderabad | Work From Office
🏢 Resource Capita It Services
📍 Hyderabad