27 Sep
|
LTM
|
Coimbatore
Role Description Role Overview
We are seeking a highly experienced AI Infrastructure Architect to design, build, and govern scalable, secure, resilient, and cost-efficient AI platforms across AWS, Microsoft Azure, and Google Cloud Platform. The role will lead end-to-end architecture for AI/ML workloads, including data platforms, model training, fine-tuning, inference, MLOps, GPU infrastructure, observability, security, and governance for enterprise and production-grade use cases.
The ideal candidate combines deep cloud infrastructure expertise, hands-on AI/ML platform knowledge, and strong enterprise architecture leadership across multi-cloud environments.
Key Responsibilities AI/ML Infrastructure Architecture
- Lead the design of end-to-end AI infrastructure for model experimentation, training, fine-tuning, inference, deployment, monitoring, and ongoing operations.
- Architect scalable platforms for batch and real-time ML workloads, LLM-based solutions, Generative AI pipelines, and enterprise AI applications.
- Define standards for model experimentation, versioning, registry, promotion, lifecycle management, and retirement.
- Create reusable reference architectures, blueprints, guardrails, and design patterns for AI workloads.
- Ensure platforms meet scalability, availability, performance, disaster recovery, and operational requirements.
Multi-Cloud Platform Design: AWS, Azure & GCP
- Architect cloud-native and cloud-agnostic AI platforms across AWS SageMaker, EKS, EC2 GPU, S3 and IAM; Azure Machine Learning, AKS, Azure OpenAI and GPU VM series; and GCP Vertex AI, GKE and TPU/GPU infrastructure.
- Define workload placement principles based on capability, security, latency, resilience, portability, cost, and strategic vendor alignment.
- Enable workload portability and standardized operating practices across cloud environments.
- Define hybrid and multi-cloud AI operating models, including connectivity, identity, observability, governance, and disaster recovery.
MLOps, DevOps & Platform Engineering
- Establish MLOps frameworks for CI/CD and continuous training of models, pipelines, features, and AI applications.
- Design automation for model lifecycle management, experiment tracking, model registry, validation, deployment, rollback, and retraining.
- Implement monitoring for service health, model performance, drift, data quality, latency, throughput, reliability, and cost.
- Integrate AI delivery pipelines with enterprise DevOps, platform engineering, security,
and change-management standards.
Data & Compute Architecture
- Design scalable data ingestion, feature stores, training datasets, data lakes, and data access patterns for AI/ML workloads.
- Architect accelerator strategies using NVIDIA GPUs, TPUs, and fit-for-purpose inference compute.
- Optimize utilization, performance, scheduling, capacity, and cost for training and inference workloads.
- Define storage, networking, caching, and distributed-compute patterns for large-scale AI platforms.
Security, Governance & Compliance
- Define AI security architecture covering identity, privileged access, data access, network isolation, secrets and key management, encryption, supply-chain security, and tenant/workload isolation.
- Implement governance controls for model usage, data privacy, lineage, approvals, responsible AI, risk management, and compliance.
- Align AI platforms with enterprise security architecture, regulatory obligations, audit requirements, and internal governance frameworks.
- Embed security-by-design, policy-as-code, traceability, and evidence collection into platform workflows.
Leadership & Advisory
- Act as the technical authority for AI infrastructure and platform architecture decisions.
- Guide cloud architects, platform engineers, data engineers, ML engineers, security teams, and application teams.
- Support AI platform roadmaps, cloud strategy, capability assessments, investment decisions, and architecture reviews.
- Communicate architectural choices, trade-offs, risks, and recommendations to business, engineering, and leadership stakeholders.
- Mentor teams and promote reusable engineering practices and architecture standards.
Core Technical Skills Cloud Platforms & Architecture
- Advanced architecture expertise across AWS, Microsoft Azure, and GCP.
- Solid experience in cloud networking, IAM, security architecture, landing zones, resilience, and multi-cloud governance.
AI/ML Platforms
- Hands-on experience designing and deploying enterprise AI/ML infrastructure.
- Expertise with Azure Machine Learning, AWS SageMaker, and GCP Vertex AI.
- Experience with Generative AI and LLM platforms supporting training, fine-tuning, evaluation, inference, and monitoring.
Infrastructure & Platform Engineering
- Kubernetes expertise across EKS, AKS, and GKE.
- GPU/accelerator infrastructure architecture, scheduling, performance tuning, capacity management, and cost optimization.
- Infrastructure as Code using Terraform, ARM/Bicep and/or CloudFormation.
- Containerization, platform automation, service mesh, networking, and observability.
MLOps & Automation
- CI/CD and continuous training for ML pipelines and AI applications.
- Model registry, experiment tracking, feature/pipeline versioning, deployment automation, and inference scaling.
- Monitoring, logging, ing, drift detection, reliability engineering, and performance tuning.
Data Systems
- Large-scale data platforms for AI/ML workloads, including batch and streaming architectures.
- Feature stores, data ingestion, data quality, lineage, governance, and secure data-access patterns.
- Strong understanding of distributed systems and high-performance computing concepts.
Preferred Qualifications
- Experience designing and governing enterprise AI platforms at scale.
- Exposure to Responsible AI frameworks, model risk management, and AI governance operating models.
- Strong background in cost optimization and FinOps for GPU-intensive AI workloads.
- Consulting, client-facing advisory, or architecture review experience.
- Experience supporting regulated industries.
- Relevant certifications in AWS, Azure, GCP, Kubernetes, enterprise architecture, security, or AI/ML.
Education & Experience
- Bachelor's or Master's degree in Computer Science, Engineering, Information Technology, or a related field.
- 12+ years of overall experience in infrastructure, cloud, platform engineering, or enterprise architecture.
- Proven experience leading AI/ML infrastructure architecture initiatives from strategy through production adoption.
- Strong architectural judgement, written and verbal communication, and stakeholder-management capability.
Key Competencies
- Enterprise architecture leadership
- Strategic thinking and decision-making
- Multi-cloud architecture and governance
- AI platform engineering and MLOps
- Security, resilience, compliance, and cost management
- Executive communication and stakeholder influence
- Technical mentorship and cross-functional collaboration
Job Location: Coimbatore
📌 AI Infrastructure Architect (AWS AI Services | AWS Cloud Architecture) (Coimbatore)
🏢 LTM
📍 Coimbatore