01 Oct
|
Trine Infotech
|
India
01 Oct
Trine Infotech
India
About the Role
We are looking for an Infrastructure AI Engineer to join our SRE team and build, deploy, and optimize high-performance cloud and bare-metal infrastructure for AI/ML and LLM workloads.
Key Responsibilities
Design and manage scalable infrastructure for AI model training and real-time inference.
Manage GPU/TPU clusters, Kubernetes, Ray, Slurm, and distributed storage.
Build AIOps capabilities using Prometheus, Datadog, OpenTelemetry and automated anomaly detection.
Develop self-healing infrastructure and automated remediation workflows.
Define and monitor SLOs, SLAs, and Error Budgets.
Build CI/CD and MLOps infrastructure for continuous model deployment and validation.
Optimize GPU/CPU utilization, infrastructure performance, and cloud costs.
Perform capacity planning, load testing, and stress testing for large-scale inference workloads.
Required Skills
6+ years in SRE,
Systems Engineering, Infrastructure Engineering, or Cloud Platform Engineering, preferably with AI/ML workloads.
Strong hands-on experience with Kubernetes, containers, and GPU provisioning.
Advanced experience with Terraform, Pulumi, Ansible, or CloudFormation.
Programming experience in Python, Go, or C++.
Solid knowledge of Prometheus, Grafana, OpenTelemetry, and up-to-date observability platforms.
Valuable to Have
Experience with vLLM, TensorRT-LLM, or Triton Inference Server.
Knowledge of DeepSpeed, Megatron-LM, or Ray Cluster.
Understanding of NVLink, InfiniBand, RoCE, GPU interconnects, and memory optimization.
Exposure to Pinecone, Milvus, Qdrant, or other vector databases.
📌 Infrastructure Ai Engineer Gurugram (India)
🏢 Trine Infotech
📍 India