Your Technical responsibilities:
- Design and develop scalable ETL/ELT pipelines on GCP
- Build and maintain data ingestion and transformation frameworks using BigQuery and GCP services
- Develop Cloud Functions and automation scripts for data movement and processing
- Create and maintain workflows using Cloud Composer/Airflow
- Implement data reconciliation, validation, and data quality controls
- Troubleshoot pipeline failures and performance issues using monitoring and logging tools
- Support migration of on-premises or legacy workloads to GCP
- Create dashboards, reports, and analytics solutions for business users
- Collaborate with business, architecture, and development teams to deliver data-driven solutions.
- Use Data flow or STS to build pipeline from traditional and cloud source to GCP.
- Valuable understanding of cloud design considerations and FinOps
- Proficient in a modern scripting language like Python, Spark and Scala.
- Work with container technology such as Docker, version control systems (Github), build management and CI/CD tools (Concourse, Jenkins)
- Data Pipeline Workflows: Build and maintain data pipelines using tools and services like Dataflow, Dataproc, Apache Beam, Apache Airflow, Cloud Composer, etc. to collect, process, and distribute data.
- CI/CD Implementation: Implement CI/CD pipelines for data applications using services like Cloud Build, Jenkins, or Gitlab. Continuous testing, continuous integration, build, package,
and deployment should be automated across multi-cloud environments.
- Code Management: Take charge of source code repositories like GitHub or Bitbucket, ensuring robust version control practices.
- Data Security & Compliance: Implement security best practices. Handle actions such as encryption, managing and securing service accounts, secret management (like with Secret Manager), authentication, access control, and data compliance necessities (like GDPR).
- Performance Optimization and Cost Control: Monitor, debug, troubleshoot, and optimize performance of Dataflow jobs, and other data infrastructure. Implement strategies for optimizing cost across multiple clouds.
- Data Lakes and Data Warehouses Management: Manage BigQuery or other cloud data warehouse solutions and services for Data Lakes.
Mandatory Requirements (Qualifications)
We are looking for the candidates with the following:
- BE/BTech/MCA with a sound industry experience of 5 to 8 years
- 36+ years of Data Engineering experience with hands-on GCP implementation
- Programming Language- Python, SQL, Spark (PySpark)
- Experience in FS domain- preferably asset wealth management
- Experience with CI/CD pipelines and cloud-native development practices.
- Exposure to Apache Spark, Databricks, and modern data engineering frameworks.
- Strong communication skills
- Google Cloud Certified is add-on
📌 Artificial Intelligence Engineer (Hyderabad)
🏢 EY
📍 Hyderabad