- Good experience as a data Engineer with Pyspark, Databrick
- Design, develop, and maintain robust and scalable data pipelines for ingesting, transforming, and loading data from various sources into the data lake using PySpark. This includes implementing ETL/ELT processes.
- Solid proficiency in Python and SQL.
- AWS Experience in Managing and optimize data storage within the data lake (e.g., Azure Data Lake Storage, AWS S3) including data partitioning, file formats (Parquet, Delta Lake), and schema evolution.
- Ability to access, clean, transform and enrich complex data from multiple sources.
- Experience with data visualization tools (Tableau, Power BI, etc.) is advantage
- Writing complex SQL queries and PL/SQL procedures to extract, transform, and load (ETL) data from Oracle and SQL Server databases.
- Developing and optimizing database objects like tables, views, stored procedures, functions, and triggers.
- Ensuring data quality and integrity through validation and cleaning processes.
- Deep understanding of telecommunications business processes and systems"