30 Jul
|
Accenture
|
Secunderabad
30 Jul
Accenture
Secunderabad
Project Role: AI Infrastructure Architect
Project Role Description: Architect and build custom Artificial Intelligence (AI) infrastructure/hardware solutions. Optimize AI infrastructure/hardware performance, power consumption, cost and scalability of computational stack. Advise on AI infrastructure technology and vendor evaluation, selection and full stack integration.
Must have skills: Machine Learning Operations
Good to have skills: Amazon Web Services (AWS)
Minimum 5 year(s) of experience is required
Educational Qualification: 15 years full time education
Role Summary / Description
AI Powered Tech Talent
As a hands-on Engineer in AI Infrastructure Architecture, you will design, build, automate, monitor and optimize AI/ML infrastructure on AWS for reliable, scalable and cost-effective model development and production workloads. You will work on moderately complex infrastructure components under guidance from senior architects and engineers, contributing to GPU/accelerated compute environments, model deployment pipelines, observability, security and operational reliability for AI-driven business solutions.
Key Responsibilities
- Write, review and debug code, scripts and infrastructure-as-code for AWS AI infrastructure, automation, monitoring and deployment tooling.
- Configure and provision AWS compute resources for AI/ML workloads, including GPU-enabled instances, Amazon EC2, Amazon EKS, Amazon SageMaker and supporting storage/networking services.
- Support deployment automation and CI/CD pipelines for AI systems, models and applications using tools such as Git, Terraform/CloudFormation, Docker, Kubernetes and workflow orchestration tooling.
- Deploy and operate AI services, model-serving components and data pipelines while applying reliability, security, cost-efficiency and scalability practices.
- Monitor infrastructure and model-serving health using AWS CloudWatch and related observability tools; troubleshoot issues across compute, storage, networking, containers and application layers.
- Collaborate with data scientists, ML engineers, platform engineers and architects to integrate AI models into enterprise systems while meeting compliance and operational requirements.
- Document reusable patterns, configuration standards and runbooks for AWS-based AI infrastructure.
Required Qualifications
- Bachelor's degree in Computer Science, Computer Engineering, Information Technology or a related engineering field.
- Minimum 2 years of experience coding, building, monitoring or troubleshooting AI/ML infrastructure, data platforms, model deployment pipelines or cloud/platform engineering solutions.
- Robust understanding of AI/ML concepts and the compute, storage, networking, security and deployment foundations required to run AI workloads.
- Minimum 2 years of proficiency in programming or scripting languages such as Python, Java, C++, Bash or PowerShell.
- Experience with CI/CD, infrastructure-as-code, containers, Kubernetes,
workflow orchestration and operational monitoring tools.
- Strong problem-solving ability, communication skills and collaboration mindset in a fast-paced engineering environment.
Required Skills/ Experience
- Hands-on experience with AWS services relevant to AI infrastructure such as EC2, EKS, SageMaker, S3, IAM, VPC, CloudWatch and related DevOps services.
- Experience designing or operating GPU/accelerated compute, distributed training setups, containerized deployments and model-serving workloads.
- Working knowledge of Terraform or CloudFormation, Docker, Kubernetes, CI/CD pipelines and observability practices.
- Ability to optimize infrastructure for performance, reliability, scalability, cost and security.
- Understanding of MLOps patterns including experiment tracking, model registry, model deployment, monitoring and rollback approaches.
Good to Have Skills
- AWS certification such as AWS Certified Solutions Architect, Developer, DevOps Engineer or Machine Learning specialty/associate level.
- Exposure to industry use cases in BFSI, healthcare, retail/e-commerce, telecom, manufacturing or public sector where AI infrastructure must meet compliance, reliability and data-governance expectations.
- Familiarity with large language model infrastructure, vector databases, retrieval pipelines, GPU scheduling or model optimization techniques.
- Knowledge of security controls, FinOps practices, incident management and production support processes for enterprise AI platforms.
Locations: Job No. ATCI-5700819-S2061817 | Hyderabad | Required Skill: Machine Learning Operations
📌 AI Infrastructure Architect (Secunderabad)
🏢 Accenture
📍 Secunderabad