We are looking for a highly skilled Data Engineer / Senior Data Engineer with strong expertise in Python, SQL, GCP, BigQuery, and Generative AI technologies. The ideal candidate will be responsible for building scalable data pipelines, cloud-native data solutions, and AI-powered applications leveraging RAG (Retrieval-Augmented Generation) frameworks and OpenAI technologies.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines and ETL/ELT workflows.
- Build and optimize AI-powered applications using OpenAI APIs and RAG architectures.
- Develop document ingestion, chunking, embedding, and retrieval pipelines.
- Work with structured and unstructured data at scale.
- Design and optimize data models in BigQuery, PostgreSQL, and Firestore.
- Build and manage workflow orchestration using Apache Airflow / Cloud Composer.
- Develop cloud-native solutions on Google Cloud Platform (GCP).
- Containerize applications using Docker and support CI/CD processes.
- Collaborate with cross-functional teams to deliver high-quality, scalable solutions.
- Ensure code quality through testing, code reviews, and best engineering practices.
Required Skills
- 6+ years of experience in Data Engineering or Backend Engineering.
- Robust programming experience in Python.
- Advanced knowledge of SQL.
- Hands-on experience with Google Cloud Platform (GCP).
- Experience with BigQuery.
- Experience building RAG (Retrieval-Augmented Generation) Pipelines.
- Hands-on experience with OpenAI APIs / LLM-based applications.
- Strong understanding of Chunking, Embeddings, and Vector Search.
- Experience with PostgreSQL and pgvector.
- Experience with Firestore or other NoSQL databases.
- Experience with Apache Airflow / Cloud Composer.
- Proficiency in Docker and Git. [JD_DE | Word]