02 Aug
|
Ignatiuz
|
Indore
Job Description
Position Summary
We are seeking an experienced AI Infrastructure and Platform Architect to design, optimize, and manage scalable AI infrastructure and platforms across on-premises, cloud, and hybrid environments.
The ideal candidate should have strong experience with GPU-based systems, AI/ML platforms, infrastructure architecture, performance optimization, capacity planning, and production support. The role will work closely with AI/ML developers, Dev Ops engineers, data engineers, and solution architects to improve the performance, reliability, scalability, and cost efficiency of AI solutions.
Key Responsibilities
- Design and manage AI infrastructure for model training, fine-tuning, inference, computer vision, Generative AI, and LLM workloads.
- Define CPU, GPU, RAM, VRAM, storage, networking, cooling, and power requirements.
- Review existing hardware and platform performance and recommend upgrades or optimizations.
- Perform capacity planning to support future workloads and minimize frequent hardware changes.
- Build and maintain AI platforms using Linux, Docker, Kubernetes, GPU orchestration, and cloud services.
- Configure and manage NVIDIA drivers, CUDA, cuDNN, TensorRT, and related AI acceleration technologies.
- Monitor system health, GPU utilization, memory usage, storage performance, and network throughput.
- Diagnose infrastructure failures, system crashes, performance bottlenecks, and platform outages.
- Implement monitoring, alerting, backup, disaster recovery, security, and operational best practices.
- Prepare architecture documents, hardware specifications, technical recommendations, and operational runbooks.
- Support production deployment, troubleshooting, and continuous platform improvement.
AI Solution Optimization
The candidate should also be capable of:
- Reviewing the end-to-end AI solution and identifying performance, architecture, and infrastructure gaps.
- Recommending improvements to scalability, reliability, maintainability, and cost efficiency.
- Supporting AI/ML developers with model training and experimentation environments.
- Helping reduce training time through GPU optimization, distributed training, resource tuning, and efficient data pipelines.
- Providing guidance on model accuracy, evaluation, hyperparameter tuning, and experimentation practices.
- Improving model-serving and inference performance.
- Mentoring existing team members on AI infrastructure and production-readiness best practices.
Required Skills
- AI infrastructure and GPU-based computing
- NVIDIA GPU architecture, CUDA, cuDNN, NCCL, and TensorRT
- Linux administration
- Docker and Kubernetes
- PyTorch, Tensor Flow, Hugging Face, or similar frameworks
- Cloud and on-premises AI platforms
- Infrastructure sizing and capacity planning
- Performance monitoring and troubleshooting
- High-performance storage and networking
- MLOps, CI/CD, automation, and Infrastructure as Code
- Monitoring tools such as Prometheus, Grafana, NVIDIA DCGM, or Open Telemetry
Qualifications
- Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Information Technology, Engineering, or a related field.
- Strong overall experience in infrastructure, cloud, platform engineering, architecture, or AI systems.
- A minimum of 3 years of direct, hands-on experience specifically working with AI infrastructure, machine learning platforms, GPU environments, or production AI workloads.
- Proven experience designing, deploying, supporting, or optimizing production AI systems.
- Robust problem-solving, troubleshooting, communication, and technical documentation skills.
The candidate's total professional experience may be significantly higher. However, at least three years should involve genuine, relevant, hands-on work with AI systems and platforms.
Added Advantage
Preference will be given to candidates who have experience with:
- AI solution architecture
- Large Language Models and Generative AI
- RAG and agentic AI systems
- Distributed model training
- Computer vision and edge AI
- Model-serving platforms
- AI performance benchmarking
- Fin Ops and infrastructure cost optimization
- High Performance Computing environments
Experience Validation
Candidates should be able to explain their direct contribution to AI projects, including:
- AI infrastructure or platforms they designed or managed
- GPU and hardware-sizing decisions
- Model training or inference environments supported
- Performance issues diagnosed and resolved
- Improvements achieved in training time, utilization, reliability, or cost
- Production AI workloads they deployed or maintained
General Dev Ops, cloud, or system administration experience without direct AI or machine learning exposure will not be sufficient for this position.
Requirements
AI Infrastructure, AI Platform Architect, MLOps, Kubernetes, Docker, CUDA, NVIDIA GPU, TensorRT, Linux, AWS, Prometheus, Grafana, Machine Learning Infrastructure, GPU Optimization, Production Support
📌 AI Infrastructure and Platform Architect (Indore)
🏢 Ignatiuz
📍 Indore