Core Responsibilities
- Design and optimize batch/streaming data pipelines using Scala, Spark, and Kafka
- Implement real-time tokenization/cleansing microservices in Java
- Manage production workflows via Apache Airflow (batch scheduling)
- Conduct root-cause analysis of data incidents using Spark/Dynatrace logs
- Monitor EMR clusters and optimize performance via YARN/Dynatrace metrics
- Ensure data security through HashiCorp Vault (Transform Secrets Engine)
- Validate data integrity and configure alerting systems
Mandatory Competencies
- Expertise in distributed data processing (Spark on EMR/Hadoop)
- Proficiency in shell scripting and YARN job management
- Ability to implement format-preserving encryption (tokenization solutions)
- Experience with production troubleshooting (executor logs, metrics, RCA)
Benefits
Benefits
Insurance - Family
Term Insurance
PF
Paid Time Off - 20 days
Holidays - 10 days
Flexi timing
Market-competitive salary
Diverse & Inclusive workspace
📌 Big Data Engineer (Hyderabad)
🏢 Kmccorp India
📍 Hyderabad
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.