30 Jul
|
Accenture
|
Pune
Project Role: AI Infrastructure Architect
Project Role Description:
Architect and build custom Artificial Intelligence (AI) infrastructure/hardware solutions. Optimize AI infrastructure/hardware performance, power consumption, cost, and scalability of computational stack. Advise on AI infrastructure technology and vendor evaluation, selection, and full stack integration.
Must have skills:
Machine Learning Operations
Good to have skills:
Microsoft Azure Data Services
Minimum Experience Required:
12 year(s) of experience is required
Educational Qualification:
15 years full-time education
Role Summary / Description
AI Powered Tech Talent
As a Senior Engineer in AI Infrastructure Architecture for Azure, you will own significant portions of the end-to-end architecture and engineering of optimized compute infrastructure for large-scale AI and machine learning systems. You will design scalable distributed training environments, model-serving foundations, automation patterns, and operational controls that align with client standards, SLAs, security, compliance, and cost-efficiency expectations. You will bring industry experience across enterprise AI adoption, cloud modernization, regulated workloads, FinOps, and production reliability, while mentoring engineers and partnering with architects to translate business requirements into robust Azure-based AI infrastructure solutions.
Key Responsibilities
- Own end-to-end architecture and design of optimized Azure compute infrastructure for large-scale AI/ML systems, including distributed training, GPU/accelerated compute, container platforms, and model-serving environments.
- Design and tune large-scale Azure GPU clusters and distributed training systems using services such as Azure VMs, AKS, Azure Machine Learning, Azure Storage, Azure NetApp Files, Azure Virtual Network, Entra ID, Azure Monitor, and Azure DevOps/GitHub Actions, including accelerator selection, networking, and high-throughput storage design.
- Serve as an authoritative AI infrastructure expert on Azure, applying deep knowledge of Azure AI/ML services, accelerators, networking, security, and cost levers.
- Develop and evaluate architecture alternatives, weighing trade-offs across compute, networking, storage, orchestration, model serving, observability, security, compliance, cost, and operational complexity.
- Lead architecture assessments and reviews of existing and proposed environments,
identifying gaps, risks, bottlenecks, and optimization opportunities, and recommending remediation actions.
- Drive architecture decision-making by documenting rationale, trade-offs, assumptions, and dependencies so decisions are transparent, defensible, and aligned with business SLAs and standards.
- Define and maintain AI infrastructure roadmap inputs, capacity planning models, scaling strategies, cost forecasts, and performance improvement opportunities.
- Design deployment, automation, and CI/CD strategies for reliable, repeatable, and scalable releases of AI systems, models, data pipelines, and platform components into production.
- Establish AI monitoring and observability practices across InfraOps and MLOps, including SLAs, SLOs, alerting, performance/cost tracking, and continuous optimization.
- Integrate AI/ML systems into enterprise environments while ensuring interoperability, security, compliance, regulatory alignment, and adherence to client standards.
- Collaborate with clients, stakeholders, architects, and engineering teams to align infrastructure decisions with business outcomes and translate requirements into actionable architecture standards.
- Set technical direction for workstreams, mentor engineers, review designs/code, and promote engineering best practices across the team.
Required Qualifications
- Bachelor's degree in Computer Science, Computer Engineering, Information Technology, or a related engineering field.
- Minimum 4 years of experience coding, building, monitoring, troubleshooting, designing, and operating AI/ML infrastructure, cloud platforms, data platforms, model deployment pipelines, or large-scale engineering solutions.
- Strong understanding of AI/ML concepts and the computing infrastructure required to deploy, run, and optimize production AI workloads.
- Minimum 4 years of proficiency in programming or scripting languages such as Python, Java, C++, Bash, PowerShell, or equivalent engineering languages.
- Experience with data pipeline and workflow management tools such as Apache Airflow, Kubeflow, managed orchestration services, or platform-native workflow tooling.
- Strong problem-solving skills and ability to work in a fast-paced engineering or client delivery environment.
- Excellent communication, collaboration, and stakeholder alignment skills.
- Minimum 4 years of experience in AI/ML infrastructure engineering or related roles on a hyperscaler or enterprise platform for deploying large-scale solutions.
- Proven experience leading AI projects or engineering workstreams and managing priorities across multiple initiatives.
- Demonstrated experience evaluating and selecting AI technologies, frameworks, cloud services, and architecture patterns.
Required Skills/Experience
- Strong hands-on experience with Azure AI infrastructure services including Azure VMs, AKS, Azure Machine Learning, Azure Storage, Azure NetApp Files, Azure Virtual Network, Entra ID, Azure Monitor, and Azure DevOps/GitHub Actions.
- Experience architecting GPU/accelerated compute, distributed training, model serving, high-throughput storage, container platforms, and secure cloud networking.
- Strong working knowledge of Bicep/ARM/Terraform, CI/CD, Docker, Kubernetes, InfraOps, MLOps, observability, and incident response practices.
- Ability to optimize Azure AI infrastructure for performance, power, cost, scalability, security, reliability, and compliance.
- Experience producing architecture decision records, reference implementations, standards, runbooks, and reusable infrastructure patterns.
Valuable to Have Skills
- Azure certifications such as Azure Solutions Architect Expert, Azure DevOps Engineer Expert, Azure AI Engineer, or Azure Data Engineer credentials.
- Industry experience in BFSI, healthcare, retail/e-commerce, telecom, manufacturing, energy, or public sector environments where AI infrastructure must meet compliance, security, reliability, and cost-control requirements.
- Exposure to LLM infrastructure, vector databases, retrieval pipelines, GPU scheduling, high-performance storage, low-latency model serving, and model optimization techniques.
- Knowledge of enterprise architecture governance, FinOps, infrastructure partner/vendor collaboration, and production support operating models.
Locations: Job No. ATCI-5700875-S2061799 | Pune | Required Skill: Machine Learning Operations
📌 AI Infrastructure Architect (Pune)
🏢 Accenture
📍 Pune