Founded in 1976, CGI is among the largest independent IT and business consulting services firms in the world. With 94,000 consultants and professionals across the globe, CGI delivers an end-to-end portfolio of capabilities, from strategic IT and business consulting to systems integration, managed IT and business process services and intellectual property solutions. CGI works with clients through a local relationship model complemented by a global delivery network that helps clients digitally transform their organizations and accelerate results. CGI Fiscal 2024 reported revenue is CA$14.68 billion and CGI shares are listed on the TSX (GIB.A) and the NYSE (GIB). Learn more at cgi.com.
Job Title: Data Engineer ETL, Python/Scala, Apache Spark
Position: Software Engineer
Experience: 5- 7 Years
Category: Software Development/ Engineering
Shift: 1-10PM (Hybrid)
Main location: Bangalore
Position ID: J0926-0282
Employment Type: Full Time
Education Qualification: Bachelor's degree in Computer Science or related field or higher with minimum 3 years of relevant experience.
Position Description: We are looking for an experienced Data Engineer to join our team. The ideal candidate should be passionate about coding and developing scalable and high-performance applications.
#LI-BK7
Your future duties and responsibilities
Core Data Engineering
• ETL/ELT Architecture: Design and implement complex data pipelines from disparate sources into our centralized data lake/warehouse.
• Big Data Processing: Leverage Apache Spark (via Dataproc) and Python/Scala to handle petabyte scale transformations.
• Advanced SQL: Write and optimize complex queries for data modeling, ensuring high performance in BigQuery.
• Stream & Batch: Develop real time processing solutions using Pub/Sub and Dataflow (Apache Beam).
GCP Infrastructure
• Orchestration: Manage workflow automation using Cloud Composer (Airflow).
• Storage & Databases: Architect solutions across Cloud Storage, Cloud SQL, and Spanner, choosing the right tool for the specific latency and consistency requirements.
• Environment Management: Ensure security, logging, and monitoring are integrated into every pipeline.
Must-Have Skills:
Programming: Mastery of Python or Scala.
Spark Expertise: Proven experience tuning and debugging Spark jobs at scale.
GCP Professional: Deep hands on experience with the full Google data suite (BigQuery, Dataflow, Dataproc, Pub/Sub).
SQL Mastery: Ability to perform complex window functions, CTEs, and query optimization.
Bonus Points (Valuable to Have)
AI/ML Integration: Familiarity with Vertex AI and the GCP AI Platform.
Good-to-Have Skills:
GenAI Literacy: A basic understanding of LLMs, prompt engineering, and how data pipelines feed into RAG (Retrieval Augmented Generation) architectures.
Infrastructure as Code: Experience with Terraform for managing GCP resources.