4+ years of overall IT experience, including hands-on experience with Big Data technologies.
- Robust hands-on experience in Python and PySpark.
- Python should be used extensively for application development, ETL
(Extract, Transform, Load) processes, and Data Lake curation.
- Experience in building PySpark applications using Spark DataFrames in
Python with development tools such as Jupyter Notebook and PyCharm (IDE).
- Proven experience in optimizing Spark jobs that process large-scale datasets and high-volume workloads.
- Hands-on experience with version control systems, particularly Git.
- Experience working with AWS Analytics services, including: