18 Sep
|
Uplers
|
Bengaluru
Experience : 2.00 + years
Salary : INR 2000000-3000000 / year (based on experience)
Expected Notice Period : 30 Days
Shift : (GMT+05:30) Asia/Kolkata (IST)
Opportunity Type : Hybrid ()
Placement Type : Full Time Permanent position(Payroll and Compliance to be managed by: A technology company)
(*Note: This is a requirement for one of Uplers' client - A technology company)
What do you need for this opportunity?
Must have skills required: Familiarity with GPU, Tgi, vLLM/SGLang, GPU infrastructure, LLM framework, VPC, on -prem, L4, TCP/UDP/QUIC/L4S, TTFT, Tpot, A10G A technology company is Looking for: Senior LLMops / Platform Engineer - AI Systems About The Job Role summary They deploy Sage, our fine-tuned BFSI-specific SLM, in bank and NBFC environments, including on-prem and VPC-controlled setups. This role owns capacity planning, load testing, and infrastructure architecture for these deployments.
Responsibilities
- Own production readiness for Sage deployments, on-prem and cloud. Load testing at a defined multiple of expected volume is required before any go-live, regardless of client timeline.
- GPU Infrastructure & Serving Stack: Own GPU architecture and capacity planning for our serving stack (vLLM/TGI), including instance selection, AWQ memory profiling, and throughput benchmarking across GPU classes.
- Latency Diagnosis & Tuning: Diagnose latency across queueing, prefill (TTFT), and decode (TPOT) stages, determining which bottlenecks are fixable by scaling versus engine/architectural reconfiguration.
- Build and maintain a standard load-testing and capacity-planning process applied to every deployment.
- Handle bank-controlled VPC deployment requirements: network isolation/air-gapped environments, data residency, audit logging.
- Knowledge Transfer & Team Resilience:
Drive active cross-skilling with the existing DevOps engineer to eliminate single-point-of-failure risks.
- Hold go/no-go authority on production readiness, independent of client-driven timelines.
Requirements
- Experience running GPU-backed inference infrastructure in production under real load.
- Direct experience with vLLM, TGI, or comparable LLM serving frameworks in production.
- Experience designing and running load tests that accurately predicted production behaviour.
- On-prem or VPC-constrained deployment experience. Cloud-only experience isn''''t a substitute.
- Comfortable owning a production-readiness call and pushing back on a launch date when the system isn''''t ready.
Nice to have
- Experience in a regulated industry where downtime has compliance consequences.
- Familiarity with GPU instance economics (e.g. A10G vs L4 class) sufficient to make deployment recommendations directly.
How to apply for this opportunity?
- Step 1: Click On Apply! And Register or Login on our portal.
- Step 2: Complete the Screening Form & Upload updated Resume
- Step 3: Increase your chances to get shortlisted & meet the client for the Interview!
About Uplers: Our goal is to make hiring reliable, simple, and fast. Our role will be to help all our talents find and apply for relevant contractual onsite opportunities and progress in their career. We will support any grievances or challenges you may face during the engagement. (Note: There are many more opportunities apart from this on the portal. Depending on the assessments you clear, you can apply for them as well).
So, if you are ready for a new challenge, a outstanding work environment, and an opportunity to take your career to the next level, don't hesitate to apply today. We are waiting for you!
📌 Senior LLMops / Platform Engineer - AI Systems (Bengaluru)
🏢 Uplers
📍 Bengaluru