- Understanding and experience with Google Cloud architecture
- Building Data pipelines with Data proc Cluster, Data Flow, Pub sub, Cloud Composer (Airflow), Cloud functions etc.)
- Be fluent in several object-oriented languages, preferably Python.
- Proficiency with build tools like SBT (Straightforward Build Tool)
- Hands on Spark/Scala programming experience, knowledge on optimization techniques in spark ecosystem
- In-depth knowledge of database and data warehouse technologies & concepts from Google (e.g., Big Query, Cloud SQL)
- Be fluent in Python
- Good experience with enterprise-scale data modelling
- Strong understanding of distributed systems architecture.
- Data quality process experience including data cleansing, audits and alerts, triage mechanisms & processes, and referential integrity
- CI/CD: Release, deployment with Cloud build and git flow, GKE, Docker
- A certification such as Google Cloud Professional Cloud Architect, Google Professional Data Engineer would be a plus and should be demonstratable