The successful candidate will work on batch and real-time data pipelines, leveraging technologies such as Scala, Java, and Apache Spark. They will manage distributed systems and cloud resources, ensure security compliance using HashiCorp Vault, orchestrate workflows with Apache Airflow, and investigate production issues by analyzing logs and monitoring metrics.
Responsibilities
- Develop and maintain batch and streaming data pipelines using Scala and Java.
- Write and execute shell scripts and Yarn commands for Spark job management.
- Manage and optimize big data environments using Apache Spark, EMR, Hadoop, and YARN.
- Utilize cloud storage and container platforms such as Amazon S3 and EKS.
- Implement and manage security features, including HashiCorp Vault and tokenization/encryption protocols.
- Schedule and orchestrate batch workflows using Apache Airflow.
- Conduct root cause analysis and handle production incidents.
- Analyze application, Spark executor, and Dynatrace logs.
- Monitor jobs and EMR clusters using Dynatrace metrics.
- Validate data and set up production alerts.
Requirements
Experience with batch and streaming processing.
- Proficiency in distributed systems management.
- Knowledge of cloud platforms (AWS S3, EKS).
- Security expertise (Vault, tokenization, encryption).
- Workflow orchestration (Airflow).
- Solid analytical and troubleshooting skills for production investigation
📌 Big Data/ Scala Developer (India)
🏢 Kmccorp India
📍 India
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.