08 Aug
|
wissen technology
|
Pune
08 Aug
wissen technology
Pune
Key Responsibilities
Build and maintain data transformation pipelines using java Spark
Develop and optimize large-scale/CPU intensive data processing using Apache Spark
Orchestrate workflows using Airflow
Implement data quality checks, testing, and monitoring for pipeline. Good to have exposer into managing metadata, cataloguing, and lineage
Support schema evolution, backfills, and incremental processing
Ensure pipelines meet SLAs for freshness, reliability, and performance
Expertise/working knowledge in Spark and HBase(semantic layer, virtual datasets, Reflections)
Required Skills & Qualifications
Solid hands-on experience with
HBase
Apache Spark
Experience with HBase or similar lakehouse query engines
Airflow
Understanding of data catalogs and lineage (e.g., OpenLineage, DataHub, Apache Polaris , openlineage)
Proficiency in Java
Experience with Git-based development and CI/CD
Nice-to-Have Skills
OpenTable format/Iceberg ,Apache Arrow
CDC-based analytics pipelines
Cloud platforms (AWS)
Kubernetes-based data platforms
Skills:- Java, Apache Spark, Apache Airflow, SQL and Amazon Web Services (AWS)
📌 Data Engineer_JavaSpark (Pune)
🏢 wissen technology
📍 Pune