Job Description
Roles and Responsibilities
/n
/n
Architect and maintain enterprise-grade ELT and ETL data pipelines using Python, PySpark, Kafka, and Databricks to manage large-scale risk data.
/n
Build and deploy GenAI agents utilizing Google ADK, Google Flash 2.5+ LLMs, and Model Context Protocol (MCP) integrated with Human-in-the-Loop workflows.
/n
Design, automate, and deploy microservice integrations for data-intensive applications on Open Shift and Kubernetes using robust CI/CD pipelines.
/n
Implement data federation layers supporting Lambda and Data Mesh architectures via Starburst to enable AI/ML and NLP use cases.
/n
Leverage agentic AI platforms and development assistants such as Devin.AI and Git Hub Copilot with prompt engineering to increase engineering velocity.
/n
Enforce data governance, risk management policies, and regulatory compliance standards across all data platforms.
/n
/n
/n
Preferred Candidate Profile
/n
/n
Work Experience:
8+ years in large-scale application development with 5+ years in a Python and PySpark Data Engineering lead role.
/n
Educational Background: Bachelor's degree in Computer Science, Engineering, or a related field (Master's degree preferred).
/n
Core Technical Skills: Python, PySpark, Databricks, Google ADK, LLMs, FastAPI, Spring Boot, Microservices, Kafka, SQL, Data Mesh, Starburst.
/n
Infrastructure and Cloud: Kubernetes, Open Shift, Docker, Cloud-Native Infrastructure, CI/CD pipelines.
/n
Industry Context: Data engineering experience in Banking Risk, Retail Products, Cards, Mortgage, Deposits, or Wealth Management.
/n
Assumed Requirements / Certifications: Databricks Certified Data Engineer, AWS Certified Data Analytics, or Azure Data Engineer Associate.
/n
/n
📌 Data Engineer Pune
🏢 Straive
📍 Pune