04 Aug
|
Ignatiuz
|
Indore
Position Summary
We are seeking an experienced AI Infrastructure and Platform Architect to design, optimize, and manage scalable AI infrastructure and platforms across on-premises, cloud, and hybrid environments.
The ideal candidate should have strong experience with GPU-based systems, AI/ML platforms, infrastructure architecture, performance optimization, capacity planning, and production support. The role will work closely with AI/ML developers, DevOps engineers, data engineers, and solution architects to improve the performance, reliability, scalability, and cost efficiency of AI solutions.
Key Responsibilities
Design and manage AI infrastructure for model training, fine-tuning, inference, computer vision, Generative AI, and LLM workloads.
Define CPU, GPU, RAM, VRAM, storage, networking, cooling, and power requirements.
Review existing hardware and platform performance and recommend upgrades or optimizations.
Perform capacity planning to support future workloads and minimize frequent hardware changes.
Build and maintain AI platforms using Linux, Docker, Kubernetes, GPU orchestration, and cloud services.
Configure and manage NVIDIA drivers, CUDA, cuDNN, TensorRT, and related AI acceleration technologies.
Monitor system health, GPU utilization, memory usage, storage performance, and network throughput.
Diagnose infrastructure failures, system crashes, performance bottlenecks, and platform outages.
Implement monitoring, alerting, backup, disaster recovery, security, and operational best practices.
Prepare architecture documents, hardware specifications, technical recommendations, and operational runbooks.
Support production deployment, troubleshooting, and continuous platform improvement.
AI Solution Optimization The candidate should also be capable of:
Reviewing the end-to-end AI solution and identifying performance, architecture, and infrastructure gaps.
Recommending improvements to scalability, reliability, maintainability, and cost efficiency.
Supporting AI/ML developers with model training and experimentation environments.
Helping reduce training time through GPU optimization, distributed training, resource tuning, and efficient data pipelines.
Providing guidance on model accuracy, evaluation, hyperparameter tuning, and experimentation practices.
Improving model-serving and inference performance.
Mentoring existing team members on AI infrastructure and production-readiness best practices.
Required Skills
AI infrastructure and GPU-based computing
NVIDIA GPU architecture, CUDA, cuDNN, NCCL, and TensorRT
Linux administration
Docker and Kubernetes
PyTorch, TensorFlow, Hugging Face, or similar frameworks
Cloud and on-premises AI platforms
Infrastructure sizing and capacity planning
Performance monitoring and troubleshooting
High-performance storage and networking
MLOps, CI/CD, automation, and Infrastructure as Code
Monitoring tools such as Prometheus, Grafana, NVIDIA DCGM, or OpenTelemetry
Qualifications
Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Information Technology,
Engineering, or a related field.
Robust overall experience in infrastructure, cloud, platform engineering, architecture, or AI systems.
A minimum of 3 years of direct, hands-on experience specifically working with AI infrastructure, machine learning platforms, GPU environments, or production AI workloads.
Proven experience designing, deploying, supporting, or optimizing production AI systems.
Strong problem-solving, troubleshooting, communication, and technical documentation skills.
The candidate's total professional experience may be significantly higher. However, at least three years should involve genuine, relevant, hands-on work with AI systems and platforms.
Added Advantage
Preference will be given to candidates who have experience with:
AI solution architecture
Large Language Models and Generative AI
RAG and agentic AI systems
Distributed model training
Computer vision and edge AI
Model-serving platforms
AI performance benchmarking
FinOps and infrastructure cost optimization
High Performance Computing environments
Experience Validation
Candidates should be able to explain their direct contribution to AI projects, including:
AI infrastructure or platforms they designed or managed
GPU and hardware-sizing decisions
Model training or inference environments supported
Performance issues diagnosed and resolved
Improvements achieved in training time, utilization, reliability, or cost
Production AI workloads they deployed or maintained
General DevOps, cloud, or system administration experience without direct AI or machine learning exposure will not be sufficient for this position.
📌 AI Infrastructure and Platform Architect (Indore)
🏢 Ignatiuz
📍 Indore