08 Aug
|
Nameless
|
Indore
Position Summary
We are seeking an experienced AI Infrastructure and
Platform Architect to design, optimize, and manage scalable AI
infrastructure and platforms across on-premises, cloud, and hybrid
environments.
The ideal candidate should have strong experience with
GPU-based systems, AI/ML platforms, infrastructure architecture, performance
optimization, capacity planning, and production support. The role will work
closely with AI/ML developers, DevOps engineers, data engineers, and solution
architects to improve the performance, reliability, scalability, and cost
efficiency of AI solutions.
Key Responsibilities
- Design
and manage AI infrastructure for model training, fine-tuning, inference,
computer vision, Generative AI, and LLM workloads.
- Define
CPU, GPU, RAM, VRAM, storage, networking, cooling, and power requirements.
- Review
existing hardware and platform performance and recommend upgrades or
optimizations.
- Perform
capacity planning to support future workloads and minimize frequent
hardware changes.
- Build
and maintain AI platforms using Linux, Docker, Kubernetes, GPU
orchestration, and cloud services.
- Configure
and manage NVIDIA drivers, CUDA, cuDNN, TensorRT, and related AI
acceleration technologies.
- Monitor
system health, GPU utilization, memory usage, storage performance, and
network throughput.
- Diagnose
infrastructure failures, system crashes, performance bottlenecks, and
platform outages.
- Implement
monitoring, alerting, backup, disaster recovery, security, and operational
best practices.
- Prepare
architecture documents, hardware specifications, technical
recommendations, and operational runbooks.
- Support
production deployment, troubleshooting, and continuous platform
improvement.
AI Solution Optimization
The candidate should also be capable of:
- Reviewing
the end-to-end AI solution and identifying performance, architecture, and
infrastructure gaps.
- Recommending
improvements to scalability, reliability, maintainability, and cost
efficiency.
- Supporting
AI/ML developers with model training and experimentation environments.
- Helping
reduce training time through GPU optimization, distributed training,
resource tuning, and efficient data pipelines.
- Providing
guidance on model accuracy, evaluation, hyperparameter tuning, and
experimentation practices.
- Improving
model-serving and inference performance.
- Mentoring
existing team members on AI infrastructure and production-readiness best
practices.
Required Skills
- AI
infrastructure and GPU-based computing
- NVIDIA
GPU architecture, CUDA, cuDNN, NCCL, and TensorRT
- Linux
administration
- Docker
and Kubernetes
- PyTorch,
TensorFlow, Hugging Face, or similar frameworks
- Cloud
and on-premises AI platforms
- Infrastructure
sizing and capacity planning
- Performance
monitoring and troubleshooting
- High-performance
storage and networking
- MLOps,
CI/CD, automation, and Infrastructure as Code
- Monitoring
tools such as Prometheus, Grafana, NVIDIA DCGM, or OpenTelemetry
Qualifications
- Bachelor's
or Master's degree in Computer Science, Artificial Intelligence,
Information Technology, Engineering,
or a related field.
- Strong
overall experience in infrastructure, cloud, platform engineering,
architecture, or AI systems.
- A
minimum of 3 years of direct, hands-on experience specifically working
with AI infrastructure, machine learning platforms, GPU environments, or
production AI workloads.
- Proven
experience designing, deploying, supporting, or optimizing production AI
systems.
- Solid
problem-solving, troubleshooting, communication, and technical
documentation skills.
The candidate's total professional experience may be
significantly higher. However, at least three years should involve genuine,
relevant, hands-on work with AI systems and platforms.
Added Advantage
Preference will be given to candidates who have experience
with:
- AI
solution architecture
- Large
Language Models and Generative AI
- RAG
and agentic AI systems
- Distributed
model training
- Computer
vision and edge AI
- Model-serving
platforms
- AI
performance benchmarking
- FinOps
and infrastructure cost optimization
- High
Performance Computing environments
Experience Validation
Candidates should be able to explain their direct
contribution to AI projects, including:
- AI
infrastructure or platforms they designed or managed
- GPU
and hardware-sizing decisions
- Model
training or inference environments supported
- Performance
issues diagnosed and resolved
- Improvements
achieved in training time, utilization, reliability, or cost
- Production
AI workloads they deployed or maintained
General DevOps, cloud, or system administration experience
without direct AI or machine learning exposure will not be sufficient for this
position.
📌 AI Infrastructure and Platform Architect (Indore)
🏢 Nameless
📍 Indore