- Expertise in PySpark, Python - Knowledge and Skill in Databricks - Solid proficiency in Azure services: Data Factory, Synapse, Data Lake, Databricks. - Hands-on experience with ETL tools and data pipeline orchestration. - Proficiency in Python or Scala for data processing. - Knowledge of SQL and NoSQL databases. - Familiarity with data modeling and data warehousing concepts. - Understanding of security best practices for data in AWS. - Good hands on experience on Python, Numpy , pandas. - Experience in building ETL/ Data Warehouse transformation process. - Experience working with structured and unstructured data.
- Developing Big Data and non-Big Data cloud-based enterprise solutions in PySpark and SparkSQL and related frameworks/libraries, - Developing scalable and re-usable, self-service frameworks for data ingestion and processing, - Integrating end to end data pipelines to take data from data source to target data repositories ensuring the quality and consistency of data, - Knowledge of big data frameworks (Spark, Hadoop)
Good to Have
- Experience with streaming data (Azure Event Hub, Kafka). - Knowledge of big data frameworks (Spark, Hadoop). - Exposure to machine learning data preparation. - Familiarity with Agile methodologies. - Knowledge of data management principles - Knowledge of Microsoft Fabric architecture and services