03 Sep
|
AtlasOne
|
Bengaluru
03 Sep
AtlasOne
Bengaluru
Role summary
Atlas deploys Sage, our fine-tuned BFSI-specific SLM, in bank and NBFC environments, including on-prem and VPC-controlled setups. This role owns capacity planning, load testing, and infrastructure architecture for these deployments.
Responsibilities
- Own production readiness for Sage deployments, on-prem and cloud. Load testing at a defined multiple of expected volume is required before any go-live, regardless of client timeline.
- GPU Infrastructure & Serving Stack: Own GPU architecture and capacity planning for our serving stack (vLLM/TGI), including instance selection, AWQ memory profiling, and throughput benchmarking across GPU classes.
- Latency Diagnosis & Tuning: Diagnose latency across queueing, prefill (TTFT), and decode (TPOT) stages, determining which bottlenecks are fixable by scaling versus engine/architectural reconfiguration.
- Build and maintain a standard load-testing and capacity-planning process applied to every deployment.
- Handle bank-controlled VPC deployment requirements: network isolation/air-gapped environments, data residency, audit logging.
- Knowledge Transfer & Team Resilience:
Drive active cross-skilling with the existing DevOps engineer to eliminate single-point-of-failure risks.
- Hold go/no-go authority on production readiness, independent of client-driven timelines.
Requirements
- Experience running GPU-backed inference infrastructure in production under real load.
- Direct experience with vLLM, TGI, or comparable LLM serving frameworks in production.
- Experience designing and running load tests that accurately predicted production behaviour.
- On-prem or VPC-constrained deployment experience. Cloud-only experience isn't a substitute.
- Comfortable owning a production-readiness call and pushing back on a launch date when the system isn't ready.
Nice to have
- Experience in a regulated industry where downtime has compliance consequences.
- Familiarity with GPU instance economics (e.g. A10G vs L4 class) sufficient to make deployment recommendations directly.
Engagement
Open to fractional, contract, or full time. Full-time conversion expected as on-prem deployment volume grows.
📌 Senior LLMops / Platform Engineer - AI Systems (Bengaluru)
🏢 AtlasOne
📍 Bengaluru