Project Role : AI Infrastructure Architect
Project Role Description : Architect and build custom Artificial Intelligence (AI) infrastructure/hardware solutions. Optimize AI infrastructure/hardware performance power consumption cost and scalability of computational stack. Advise on AI infrastructure technology and vendor evaluation selection and full stack integration.
Must have skills : AI Agents & Workflow Integration
Good to have skills : Microsoft Azure Data Services
Minimum 5 year(s) of experience is required
Educational Qualification : 15 years full time education
Role Summary / Description
AI Powered Tech Talent
As a hands-on Engineer in AI Infrastructure Architecture you will design build automate monitor and optimize AI/ML infrastructure on Microsoft Azure for reliable scalable and cost-effective model development and production workloads. you will work on moderately complex infrastructure components under guidance from senior architects and engineers contributing to GPU/accelerated compute environments model deployment pipelines observability security and operational reliability for AI-driven business solutions.
Key Responsibilities
Write review and debug code scripts and infrastructure-as-code for Azure AI infrastructure automation monitoring and deployment tooling.
Configure and provision Azure compute resources for AI/ML workloads including Azure VMs Azure Kubernetes Service Azure Machine Learning compute Azure Storage and supporting networking/security services.
Support deployment automation and CI/CD pipelines for AI systems models and applications using tools such as Git Bicep/ARM/Terraform Azure DevOps/GitHub Actions Docker Kubernetes and workflow orchestration tooling.
Deploy and operate AI services model-serving components and data pipelines while applying reliability security cost-efficiency and scalability practices.
Monitor infrastructure and model-serving health using Azure Monitor Log Analytics and related observability tools troubleshoot issues across compute storage networking containers and application layers.
Collaborate with data scientists ML engineers platform engineers and architects to integrate AI models into enterprise systems while meeting compliance and operational requirements.
Document reusable patterns configuration standards and runbooks for Azure-based AI infrastructure.
Required Qualifications
Bachelors degree in Computer Science Computer Engineering Information Technology or a related engineering field.
Minimum 2 years of experience coding building monitoring or troubleshooting AI/ML infrastructure data platforms model deployment pipelines or cloud/platform engineering solutions.
Strong understanding of AI/ML concepts and the compute storage networking security and deployment foundations required to run AI workloads.
Minimum 2 years of proficiency in programming or scripting languages such as Python Java C Bash or PowerShell.
Experience with CI/CD infrastructure-as-code containers Kubernetes workflow orchestration and operational monitoring tools.
Robust problem-solving ability communication skills and collaboration mindset in a fast-paced engineering environment.
Required Skills/ Experience
Hands-on experience with Azure services relevant to AI infrastructure such as Azure VMs AKS Azure Machine Learning Azure Storage Azure Virtual Network Azure Monitor Log Analytics and Azure DevOps/GitHub Actions.
Experience designing or operating GPU/accelerated compute distributed training setups containerized deployments and model-serving workloads.
Working knowledge of Bicep/ARM/Terraform Docker Kubernetes CI/CD pipelines and observability practices.
Ability to optimize infrastructure for performance reliability scalability cost and security.
Understanding of MLOps patterns including experiment tracking model registry model deployment monitoring and rollback approaches.
Good to Have Skills
Azure certification such as Azure Administrator Azure Developer Azure Solutions Architect Azure AI Engineer or Azure Data Engineer.
Exposure to industry use cases in BFSI healthcare retail/e-commerce telecom manufacturing or public sector where AI infrastructure must meet compliance reliability and data-governance expectations.
Familiarity with large language model infrastructure vector databases retrieval pipelines GPU scheduling or model optimization techniques.
Knowledge of security controls FinOps practices incident management and production support processes for enterprise AI platforms.