Key Responsibilities
Pipeline Orchestration: Design, develop, and maintain complex ML workflows using Apache Airflow (Cloud Composer) to automate data ingestion, preprocessing, and model training.
Lifecycle Management: Administer and scale MLflow for experiment tracking, model packaging, and maintaining a centralized Model Registry across the organization.
Cloud & Hybrid Ops: Create and optimize training settings for custom ML/LLM models.
Model Serving & Scaling: Architect high-performance inference endpoints and serve models via FastAPI/Flask with API Gateway.
Infrastructure Management: Manage auto-scaling CUDA clusters on Google Kubernetes Engine (GKE).
CI/CD: Manage end-to-end delivery with Continuous Integration & Continuous Delivery (CI/CD).
Observability & Monitoring: Build dashboards to track model health, latency, and data drift.
Requirements & Skills
Workflow Management: Experience in managing Apache Airflow and Composer to support the Data Engineering components of grounded AI solutions.
MLflow: Deep knowledge of MLflow Tracking, Projects,
and Registry. Experience migrating MLflow backends between cloud providers.
Workflow Tools: Familiarity with Vertex AI Pipelines and Azure DevOps for automation.
GCP AI Services: Practical experience with Vertex AI (Workbench, Model Garden, Feature Store) and BigQuery ML.
Containerization: Expert-level Docker and Kubernetes (GKE/AKS) skills. Must understand K8s operators and resource management for ML workloads.
Infrastructure as Code (IaC): Proficiency in Terraform to manage reproducible cloud environments.
Programming: Advanced Python skills with a focus on software engineering best practices (unit testing, modular design).
Data Engineering: Experience with Change Data Capture (CDC), Spark/PySpark, and optimizing data flow from BigQuery to training nodes.
Access Control: Knowledge of IAM roles, VPC Service Controls, and securing ML endpoints.
Experience with LLMOps (managing large