- Develop and maintain automation scripts to improve operational workflows for data platforms.
- Troubleshoot production issues across Databricks and ETL pipelines, ensuring timely resolution.
- Support and optimize data processing jobs using PySpark and SQL in cloud-based environments.
- Collaborate with engineering teams to automate repetitive operational tasks and reduce manual effort.
- Contribute to reliability improvements by identifying recurring incidents and driving corrective actions.
Skills Required:
- 11-13 years of relevant experience in data engineering and operational support for distributed systems
- PySpark for building and maintaining Spark-based data processing workloads.
- AWS experience supporting data platform operations and cloud services.
- SQL expertise for querying, debugging, and validating data pipelines.
- Databricks experience supporting enterprise data engineering and production support.
Valuable to Have:
- Informatica experience for ETL development and pipeline support.
- Shell scripting, Oracle, Python, IBM DataStage, and Git for broader platform coverage.