02 Aug
|
Hoonartek
|
India
Pune
About Us
We empower enterprises globally through intelligent, creative, and insightful services for data integration, data analytics and data visualization.
Hoonartek is a leader in enterprise transformation, data engineering and an acknowledged world-class Ab Initio delivery partner.
Using centuries of cumulative experience, research and leadership, we help our clients eliminate the complexities & risk of legacy modernization and safely deliver big data hubs, operational data integration, business intelligence, risk & compliance solutions and traditional data warehouses & marts.
At Hoonartek, we work to ensure that our customers, partners and employees all benefit from our unstinting commitment to delivery, quality and value. Hoonartek is increasingly the choice for customers seeking a trusted partner of vision, value and integrity
How We Work?
Define, Design and Deliver (D3) is our in-house delivery philosophy. It’s culled from agile and rapid methodologies and focused on ‘just enough design’. We embrace this philosophy in everything we do, leading to numerous client success stories and indeed to our own success.
We embrace change, empowering and trusting our people and building long and valuable relationships with our employees, our customers and our partners. We work flexibly, even adopting traditional/waterfall methods where circumstances demand it. At Hoonartek, the focus is always on delivery and value.
Job Description
Essential
- 4+ years of hands-on Apache Spark / PySpark experience in a production data engineering environment
- Strong SQL proficiency for data validation, reconciliation queries, and ad hoc data investigation
- Demonstrable experience running structured data quality checks — using frameworks such as Outstanding Expectations, dbt tests, or custom PySpark assertions
- Ability to read and interpret existing PySpark code confidently without making changes
- Working knowledge of Google Cloud Platform — specifically BigQuery, Google Cloud Storage (GCS), and Dataproc or Databricks on GCP
- Familiarity with distributed data formats: Parquet, Avro, ORC, or Delta Lake
Desirable
- Prior experience on a GCP-funded, Google PSO, or cloud migration validation engagement
- Exposure to data pipeline orchestration tools such as Apache Airflow or Cloud Composer
- Experience with data migration or cloud lift-and-shift validation projects
- Basic Python scripting for automating validation result aggregation and reporting
- Familiarity with Apache Iceberg, Apache Hudi, or lakehouse architecture patterns
5. Ideal Resource Profile
Each of the 2 contracted resources should exhibit the following qualities:
- Self-sufficient practitioner — able to independently navigate existing codebases, pipeline documentation, and GCP environments
- Detail-oriented with strong written communication skills for producing clear, stakeholder-ready validation reports
- Collaborative mindset with experience working within a structured GCP or cloud engagement framework
📌 Spark Engineers (India)
🏢 Hoonartek
📍 India