13 Aug
|
Naukri Assist
|
Hyderabad
13 Aug
Naukri Assist
Hyderabad
We are seeking a Data Engineer III to design, build, and operate scalable, reliable data pipelines on AWS. This role requires deep expertise in Python, PySpark, and ETL, along with strong hands-on experience using AWS Glue and Amazon S3. You will own end-to-end pipeline development—from ingestion through transformation and publishing—ensuring high data quality, performance, observability, and operational excellence.
Key Responsibilities
- Design, build, and maintain batch (and near-real-time where applicable) ETL pipelines using Python and PySpark.
- Develop and operate AWS-native data workflows using AWS Glue (Jobs, Crawlers, Catalog) and S3 as the core storage layer.
- Implement scalable data transformations (joins, aggregations, windowing, deduplication, enrichment) with a focus on performance and cost efficiency.
- Build robust data quality checks (completeness, accuracy, timeliness), reconciliation, and automated validation.
- Define and manage schemas, partitioning strategies, file formats (e.g., Parquet), and data layout best practices on S3.
- Troubleshoot production failures, optimize job runtimes, manage backfills/reprocessing, and improve pipeline resiliency.
- Partner with analytics, product,
and engineering stakeholders to translate requirements into reliable datasets and data products.
- Establish engineering standards: code reviews, testing strategy, documentation, and operational runbooks.
- Mentor junior engineers and drive continuous improvement in pipeline design, reliability, and maintainability.
Required Qualifications (Must Have)
- 5–9 years of data engineering / backend engineering experience focused on data pipelines.
- Expert-level Python for data engineering (clean code, packaging, debugging, performance tuning).
- Strong PySpark experience in production (Spark fundamentals, distributed processing, optimization).
- Proven experience building and supporting ETL pipelines end-to-end (ingestion transform publish/serve).
- Strong hands-on AWS Glue experience (authoring Glue Spark jobs, Glue Catalog, Crawlers, job orchestration patterns).
- Solid hands-on Amazon S3 experience (partitioning, lifecycle policies concepts, data layout, access patterns).
- Solid understanding of data engineering fundamentals: schema evolution, incremental loads, idempotency, late-arriving data, and backfills.
- Ability to operate pipelines in production: monitoring, alerting, incident response, and root-cause analysis.
📌 Data Engineer (Hyderabad)
🏢 Naukri Assist
📍 Hyderabad