01 Oct
|
LTM
|
Coimbatore
Role Overview
We are seeking a highly experienced AI Infrastructure Architect to design, build, and govern scalable, secure, resilient, and cost-efficient AI platforms across AWS, Microsoft Azure, and Google Cloud Platform. The role will lead end-to-end architecture for AI/ML workloads, including data platforms, model training, fine-tuning, inference, MLOps, GPU infrastructure, observability, security, and governance for enterprise and production-grade use cases.
The ideal candidate combines deep cloud infrastructure expertise, hands-on AI/ML platform knowledge, and strong enterprise architecture leadership across multi-cloud environments.
Key Responsibilities AI/ML Infrastructure Architecture
· Lead the design of end-to-end AI infrastructure for model experimentation, training, fine-tuning, inference, deployment, monitoring, and ongoing operations.
· Architect scalable platforms for batch and real-time ML workloads, LLM-based solutions, Generative AI pipelines, and enterprise AI applications.
· Define standards for model experimentation, versioning, registry, promotion, lifecycle management, and retirement.
· Create reusable reference architectures, blueprints, guardrails, and design patterns for AI workloads.
· Ensure platforms meet scalability, availability, performance, disaster recovery, and operational requirements.
Multi-Cloud Platform Design: AWS, Azure & GCP
· Architect cloud-native and cloud-agnostic AI platforms across AWS SageMaker, EKS, EC2 GPU, S3 and IAM; Azure Machine Learning, AKS, Azure OpenAI and GPU VM series; and GCP Vertex AI, GKE and TPU/GPU infrastructure.
· Define workload placement principles based on capability, security, latency, resilience, portability, cost, and strategic vendor alignment.
· Enable workload portability and standardized operating practices across cloud environments.
· Define hybrid and multi-cloud AI operating models, including connectivity, identity, observability, governance, and disaster recovery.
MLOps, DevOps & Platform Engineering
· Establish MLOps frameworks for CI/CD and continuous training of models, pipelines, features, and AI applications.
· Design automation for model lifecycle management, experiment tracking, model registry, validation, deployment, rollback, and retraining.
· Implement monitoring for service health, model performance, drift, data quality, latency, throughput, reliability, and cost.
· Integrate AI delivery pipelines with enterprise DevOps, platform engineering, security, and change-management standards.
Data & Compute Architecture
· Design scalable data ingestion, feature stores, training datasets, data lakes,
and data access patterns for AI/ML workloads.
· Architect accelerator strategies using NVIDIA GPUs, TPUs, and fit-for-purpose inference compute.
· Optimize utilization, performance, scheduling, capacity, and cost for training and inference workloads.
· Define storage, networking, caching, and distributed-compute patterns for large-scale AI platforms.
Security, Governance & Compliance
· Define AI security architecture covering identity, privileged access, data access, network isolation, secrets and key management, encryption, supply-chain security, and tenant/workload isolation.
· Implement governance controls for model usage, data privacy, lineage, approvals, responsible AI, risk management, and compliance.
· Align AI platforms with enterprise security architecture, regulatory obligations, audit requirements, and internal governance frameworks.
· Embed security-by-design, policy-as-code, traceability, and evidence collection into platform workflows.
Leadership & Advisory
· Act as the technical authority for AI infrastructure and platform architecture decisions.
· Guide cloud architects, platform engineers, data engineers, ML engineers, security teams, and application teams.
· Support AI platform roadmaps, cloud strategy, capability assessments, investment decisions, and architecture reviews.
· Communicate architectural choices, trade-offs, risks, and recommendations to business, engineering, and leadership stakeholders.
· Mentor teams and promote reusable engineering practices and architecture standards.
Core Technical Skills Cloud Platforms & Architecture
· Advanced architecture expertise across AWS, Microsoft Azure, and GCP.
· Strong experience in cloud networking, IAM, security architecture, landing zones, resilience, and multi-cloud governance.
AI/ML Platforms
· Hands-on experience designing and deploying enterprise AI/ML infrastructure.
· Expertise with Azure Machine Learning, AWS SageMaker, and GCP Vertex AI.
· Experience with Generative AI and LLM platforms supporting training, fine-tuning, evaluation, inference, and monitoring.
Infrastructure & Platform Engineering
· Kubernetes expertise across EKS, AKS, and GKE.
· GPU/accelerator infrastructure architecture, scheduling, performance tuning, capacity management,
and cost optimization.
· Infrastructure as Code using Terraform, ARM/Bicep and/or CloudFormation.
· Containerization, platform automation, service mesh, networking, and observability.
MLOps & Automation
· CI/CD and continuous training for ML pipelines and AI applications.
· Model registry, experiment tracking, feature/pipeline versioning, deployment automation, and inference scaling.
· Monitoring, logging, ing, drift detection, reliability engineering, and performance tuning.
Data Systems
· Large-scale data platforms for AI/ML workloads, including batch and streaming architectures.
· Feature stores, data ingestion, data quality, lineage, governance, and secure data-access patterns.
· Strong understanding of distributed systems and high-performance computing concepts.
Preferred Qualifications
· Experience designing and governing enterprise AI platforms at scale.
· Exposure to Responsible AI frameworks, model risk management, and AI governance operating models.
· Solid background in cost optimization and FinOps for GPU-intensive AI workloads.
· Consulting, client-facing advisory, or architecture review experience.
· Experience supporting regulated industries.
· Relevant certifications in AWS, Azure, GCP, Kubernetes, enterprise architecture, security, or AI/ML.
Education & Experience
· Bachelor's or Master's degree in Computer Science, Engineering, Information Technology, or a related field.
· 12+ years of overall experience in infrastructure, cloud, platform engineering, or enterprise architecture.
· Proven experience leading AI/ML infrastructure architecture initiatives from strategy through production adoption.
· Strong architectural judgement, written and verbal communication, and stakeholder-management capability.
Key Competencies
· Enterprise architecture leadership
· Strategic thinking and decision-making
· Multi-cloud architecture and governance
· AI platform engineering and MLOps
· Security, resilience, compliance, and cost management
· Executive communication and stakeholder influence
· Technical mentorship and cross-functional collaboration
Job Location: Coimbatore
Mandatory Skills : AI/GenAI Research, Application Rearchitecting, Architecture Patterns and Styles, Cost Benefit Analysis Method, Migration Planning,
AI Infrastructure Architect, AI Platform Architect, ML Infrastructure Architect, Multi-Cloud Architect, AWS SageMaker, Azure Machine Learning, Vertex AI, Kubernetes, EKS, AKS, GKE, Generative AI, LLM, MLOps, Terraform, GPU Infrastructure, Responsible AI, AI Governance, Platform Engineering.
📌 AI Infrastructure Architect (AWS AI Services | AWS Cloud Architecture) (Coimbatore)
🏢 LTM
📍 Coimbatore