10 Aug
|
METAPLORE SOLUTIONS
|
Bengaluru
10 Aug
METAPLORE SOLUTIONS
Bengaluru
: Senior DevOps & Platform Architect (MLOps)
About The Role We are seeking a Senior DevOps & Platform Architect with 17+ years of experience designing, building, and scaling enterprise-grade infrastructure and platform engineering solutions. The ideal candidate brings deep expertise across DevOps, Cloud Architecture, Site Reliability Engineering, and MLOps , and can architect platforms that support both traditional software delivery and machine learning workflows at scale.
This is a senior technical leadership role — you will define platform strategy, set architectural standards, and guide engineering teams in building resilient, secure, and highly automated infrastructure for both application and ML/AI workloads.
Key Responsibilities
Platform & DevOps Architecture
- Architect and lead the design of scalable, secure, and highly available cloud infrastructure (AWS/Azure/GCP).
- Define and drive DevOps strategy, including CI/CD pipelines, Infrastructure as Code (IaC), and automation frameworks across the organization.
- Own platform architecture decisions — containerization, orchestration, service mesh, networking, and observability.
- Establish and enforce standards for security, compliance, cost optimization, and disaster recovery.
- Drive infrastructure modernization initiatives (monolith-to-microservices, on-prem to cloud, cloud-to-multi-cloud).
- Lead capacity planning, performance tuning, and reliability engineering (SRE practices — SLIs/SLOs/error budgets).
MLOps & ML Platform Engineering
- Design and build end-to-end MLOps pipelines — model training, versioning, validation, deployment, and monitoring.
- Architect scalable ML infrastructure supporting experiment tracking, feature stores, model registries, and automated retraining pipelines.
- Enable CI/CD for ML (CI/CD/CT) — integrating model testing, validation, and rollback strategies into deployment workflows.
- Collaborate with Data Science and ML Engineering teams to productionize models efficiently, ensuring reproducibility and scalability.
- Implement monitoring/observability for deployed models (data drift, model drift, performance degradation).
- Optimize GPU/compute resource utilization for training and inference workloads (including on Kubernetes).
Leadership & Collaboration
- Act as a technical leader and mentor for DevOps, SRE, and Platform engineering teams.
- Partner with engineering, data science, security, and product leadership to align platform strategy with business goals.
- Drive architectural reviews, design documents, and technical roadmaps.
- Evaluate and introduce new tools/technologies to improve platform maturity and developer experience.
- Own incident management processes and post-mortem culture for platform/production issues.
Required Skills & Qualifications
- 14+ years of experience in DevOps, Site Reliability Engineering, Cloud/Platform Architecture, or related roles.
- Proven experience architecting large-scale, production-grade infrastructure on AWS,
Azure, or GCP (multi-cloud experience a plus).
- Strong hands-on expertise with Kubernetes and container orchestration at scale.
- Deep experience with Infrastructure as Code (Terraform, CloudFormation, Pulumi).
- Strong background in CI/CD tooling (Jenkins, GitLab CI, GitHub Actions, ArgoCD, Spinnaker).
- Hands-on MLOps experience with tools such as MLflow, Kubeflow, SageMaker, Vertex AI, Azure ML, or similar.
- Experience with model serving frameworks (KServe, Seldon, TorchServe, Triton Inference Server) and feature stores (Feast or similar).
- Proficiency in scripting/automation (Python, Bash, Go).
- Solid knowledge of observability stacks (Prometheus, Grafana, ELK/EFK, Datadog, New Relic).
- Experience with service mesh and API gateway technologies (Istio, Envoy, Kong).
- Solid understanding of networking, security best practices, and compliance frameworks (SOC2, HIPAA, ISO 27001 as applicable).
- Experience with GPU infrastructure and distributed training frameworks is a strong plus.
- Excellent communication skills with a track record of technical leadership and cross-functional collaboration.
Good To Have (Preferred)
- Certifications: AWS/Azure/GCP Solutions Architect (Professional level), CKA/CKAD, or equivalent.
- Experience with data pipeline orchestration (Airflow, Dagster, Prefect).
- Familiarity with LLMOps / Generative AI deployment patterns (vector databases, RAG pipelines, model gateways).
- Experience building internal developer platforms (IDPs) or self-service infrastructure tooling.
- Background in FinOps / cloud cost optimization at scale.
📌 Senior Architect - DevOps and ML Op's (Bengaluru)
🏢 METAPLORE SOLUTIONS
📍 Bengaluru