04 Aug
|
Ignatiuz
|
Indore
Position Summary
We are seeking an experienced AI Infrastructure and Platform Architect to design, optimize, and manage scalable AI infrastructure and platforms across on-premises, cloud, and hybrid environments.
The ideal candidate should have strong experience with GPU-based systems, AI/ML platforms, infrastructure architecture, performance optimization, capacity planning, and production support. The role will work closely with AI/ML developers, DevOps engineers, data engineers, and solution architects to improve the performance, reliability, scalability, and cost efficiency of AI solutions.
Key Responsibilities
- Design and manage AI infrastructure for model training, fine-tuning, inference, computer vision, Generative AI, and LLM workloads.
- Define CPU, GPU, RAM, VRAM, storage, networking, cooling, and power requirements.
- Review existing hardware and platform performance and recommend upgrades or optimizations.
- Perform capacity planning to support future workloads and minimize frequent hardware changes.
- Build and maintain AI platforms using Linux, Docker, Kubernetes, GPU orchestration, and cloud services.
- Configure and manage NVIDIA drivers, CUDA, cuDNN, TensorRT, and related AI acceleration technologies.
- Monitor system health, GPU utilization, memory usage, storage performance, and network throughput.
- Diagnose infrastructure failures, system crashes, performance bottlenecks, and platform outages.
- Implement monitoring, alerting, backup, disaster recovery, security, and operational best practices.
- Prepare architecture documents, hardware specifications, technical recommendations, and operational runbooks.
- Support production deployment, troubleshooting, and continuous platform improvement.
AI Solution Optimization The candidate should also be capable of:
- Reviewing the end-to-end AI solution and identifying performance, architecture, and infrastructure gaps.
- Recommending improvements to scalability, reliability, maintainability, and cost efficiency.
- Supporting AI/ML developers with model training and experimentation environments.
- Helping reduce training time through GPU optimization, distributed training, resource tuning, and efficient data pipelines.
- Providing guidance on model accuracy, evaluation, hyperparameter tuning, and experimentation practices.
- Improving model-serving and inference performance.
- Mentoring existing team members on AI infrastructure and production-readiness best practices.
Required Skills
- AI infrastructure and GPU-based computing
- NVIDIA GPU architecture, CUDA, cuDNN, NCCL, and TensorRT
- Linux administration
- Docker and Kubernetes
- PyTorch, TensorFlow, Hugging Face, or similar frameworks
- Cloud and on-premises AI platforms
- Infrastructure sizing and capacity planning
- Performance monitoring and troubleshooting
- High-performance storage and networking
- MLOps, CI/CD, automation, and Infrastructure as Code
- Monitoring tools such as Prometheus, Grafana, NVIDIA DCGM, or OpenTelemetry
Qualifications
- Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Information Technology, Engineering,
or a related field.
- Robust overall experience in infrastructure, cloud, platform engineering, architecture, or AI systems.
- A minimum of 3 years of direct, hands-on experience specifically working with AI infrastructure, machine learning platforms, GPU environments, or production AI workloads.
- Proven experience designing, deploying, supporting, or optimizing production AI systems.
- Strong problem-solving, troubleshooting, communication, and technical documentation skills.
The candidate's total professional experience may be significantly higher. However, at least three years should involve genuine, relevant, hands-on work with AI systems and platforms.
Added Advantage
Preference will be given to candidates who have experience with:
- AI solution architecture
- Large Language Models and Generative AI
- RAG and agentic AI systems
- Distributed model training
- Computer vision and edge AI
- Model-serving platforms
- AI performance benchmarking
- FinOps and infrastructure cost optimization
- High Performance Computing environments
Experience Validation
Candidates should be able to explain their direct contribution to AI projects, including:
- AI infrastructure or platforms they designed or managed
- GPU and hardware-sizing decisions
- Model training or inference environments supported
- Performance issues diagnosed and resolved
- Improvements achieved in training time, utilization, reliability, or cost
- Production AI workloads they deployed or maintained
General DevOps, cloud, or system administration experience without direct AI or machine learning exposure will not be sufficient for this position.
📌 AI Infrastructure and Platform Architect (Indore)
🏢 Ignatiuz
📍 Indore