18 Sep
|
Tech Mahindra
|
Bengaluru
18 Sep
Tech Mahindra
Bengaluru
Role & responsibilities
Design, develop, and maintain robust, scalable, and high-performance data pipelines using Python, Apache Spark, and Shell scripting.
-Build and optimize metadata-driven frameworks for data ingestion, transformation, and processing.
-Develop and manage ETL/ELT workflows, ensuring data quality, integrity, and reliability.
-Integrate and automate data quality checks and validation processes within data pipelines.
-Deploy and manage containerized applications using Docker and orchestrate workloads on Kubernetes.
-Work with up-to-date data lake and warehouse technologies such as Iceberg and Trino.
-Implement real-time data streaming solutions using Kafka.
-Orchestrate complex workflows using Airflow.
-Integrate with data catalog and governance tools such as Datahub and Ranger.
-Monitor and optimize data infrastructure using Prometheus and Grafana.
-Collaborate with cross-functional teams to understand business requirements and deliver data solutions.
-Ensure security, compliance, and best practices in data management and governance.
Preferred candidate profile
-Solid proficiency in Linux, Python, and Shell scripting.
-Hands-on experience with Docker, Kubernetes, and container orchestration.
-Hands-on experience with Minio and Azure Data Lake Storage (ADLS) using S3 protocols.
-Expertise in Apache Spark and distributed data processing.
-Experience with Apache Iceberg, Kafka, Airflow, Datahub, Trino, and Ranger.
-Proficiency in Java for data engineering tasks.
-Familiarity with monitoring and observability tools such as Prometheus and Grafana.
-Solid understanding of data modeling, data warehousing, and big data technologies.
📌 Hiring Pyspark Apache Iceberg Developer / Lead Chennai/ Bangalore Bengaluru
🏢 Tech Mahindra
📍 Bengaluru