Key Responsibilities
- Pipeline Orchestration: Design, develop, and maintain complex ML workflows using Apache Airflow (Cloud Composer) to automate data ingestion, preprocessing, and model training.
- Lifecycle Management: Administer and scale MLflow for experiment tracking, model packaging, and maintaining a centralized Model Registry across the organization.
- Cloud & Hybrid Ops: Create and optimize training environments for custom ML/LLM models.
- Model Serving & Scaling: Architect high-performance inference endpoints and serve models via FastAPI/Flask with API Gateway.
- Infrastructure Management: Manage auto-scaling CUDA clusters on Google Kubernetes Engine (GKE).
- CI/CD: Manage end-to-end delivery with Continuous Integration & Continuous Delivery (CI/CD).
- Observability & Monitoring: Build dashboards to track model health, latency, and data drift.
Requirements & Skills
- Workflow Management: Experience in managing Apache Airflow and Composer to support the Data Engineering components of grounded AI solutions.
- MLflow: Deep knowledge of MLflow Tracking, Projects,
and Registry. Experience migrating MLflow backends between cloud providers.
- Workflow Tools: Familiarity with Vertex AI Pipelines and Azure DevOps for automation.
- GCP AI Services: Practical experience with Vertex AI (Workbench, Model Garden, Feature Store) and BigQuery ML.
- Containerization: Expert-level Docker and Kubernetes (GKE/AKS) skills. Must understand K8s operators and resource management for ML workloads.
- Infrastructure as Code (IaC): Proficiency in Terraform to manage reproducible cloud environments.
- Programming: Advanced Python skills with a focus on software engineering best practices (unit testing, modular design).
- Data Engineering: Experience with Change Data Capture (CDC), Spark/PySpark, and optimizing data flow from BigQuery to training nodes.
- Access Control: Knowledge of IAM roles, VPC Service Controls, and securing ML endpoints.
- Experience with LLMOps (managing large