08 Oct
|
Axiom Global Technologies
|
India
08 Oct
Axiom Global Technologies
India
Role & responsibilities :
AI Infrastructure Engineer
Owns the GPU fleet, model serving, the Kubernetes clusters in the US and India, the deploy path, reliability, and cloud cost.
Must have
- GPU-backed model inference run in production on Kubernetes, with on-call ownership of that service.
- vLLM in production: tuned under real load, quantized models (AWQ or similar), concurrency and context-length limits, and upgrades rolled out without an outage.
- CUDA and GPU fundamentals: the memory model, profiling with Nsight or equivalent, reading a utilization graph or an out-of-memory error and knowing what happened, and managing the driver, CUDA toolkit, PyTorch, and vLLM version matrix.
- Kubernetes at depth: GPU node pools, scheduling, device plugins, autoscaling (KEDA or equivalent), and persistent volumes for model caches.
- Ownership of the CI/CD path from build to cluster: container builds, registries, and immutable image tags.
- Azure in production, ideally AKS with GPU node pools, including quota, SKU availability by region, and VNet networking. Deep AWS or GCP experience is acceptable.
- Observability: Prometheus and Grafana, alert design, and a record of turning incidents into alerts.
- PostgreSQL operations: pooling, migrations under load, replication or failover.
- Robust Python and shell. Terraform or equivalent infrastructure as code.
- Fluent written and spoken English. Works autonomously with clean, reviewable pull requests.
📌 AI Infrastructure Engineer (India)
🏢 Axiom Global Technologies
📍 India