26 Sep
|
InfoCepts
|
Vasanth Nagar
26 Sep
InfoCepts
Vasanth Nagar
Remote: Hybrid
Purpose of the Position: The AI Platform Engineer- MLOps/LLMOps will be responsible for designing, implementing, and operationalizing enterprise-grade AI, Machine Learning, and Generative AI solutions. The role will focus on deploying and monitoring AI applications, establishing governance frameworks, and enabling secure, reliable, and cost-effective AI operations across cloud environments. The individual will provide technical leadership and drive best practices for AI engineering, MLOps, and LLMOps initiatives.Key Result Areas and Activities:Reliability and Availability of AI services and solutionsDesign scalable AI/ML and Generative AI deployment architectures.Define enterprise standards for model deployment, monitoring, and lifecycle management.Design and operationalize RAG, agentic AI, prompt management, vector databases, and LLM integration frameworks.MLOps & LLMOps ImplementationEstablish CI/CD pipelines for machine learning and LLM-based applications.Automate model training, deployment, versioning, and rollback processes.Optimize LLM performance, scalability, and cost efficiency including Inference cost managementModel Governance & ReliabilityImplement frameworks for model monitoring, evaluation, drift detection, observability, and compliance.Ensure responsible AI and governance standards are followed.Technical Leadership & Stakeholder CollaborationMentor AI engineers,
junior platform engineers and data scientists.Collaborate with business and technology teams to translate AI use cases into production-ready solutions.Essential Skills:LLM serving and inference optimization and running an inference gateway that routes across providers and open-weight models with failoverLLM observability and tracing: end-to-end request tracing through retrieval, prompt, tool calls and generation using LangSmith or Azure AI FoundryEvaluation as a platform capability: building the harness product teams use to regression-test AI behaviour, run LLM-as-judge and offline evals, and gate releases on measured quality rather than judgement callsPrompt, model and config lifecycle management: versioning prompts and system messages like code, safe rollout and rollback, A/B and shadow testing of model or prompt changes in productionToken and inference cost engineering: spend attribution by team, feature and tenant; caching and model-routing strategies; GPU utilisation and right-sizingRetrieval infrastructure operations: running and tuning vector or hybrid search at scale, embedding pipelines, index refresh strategies, and retrieval quality monitoringKubernetes and GPU infrastructure in production: Docker, autoscaling, node pools, GPU scheduling and sharing, plus infrastructure as code (Terraform or Bicep) and GitOps (Argo CD or Flux)Strong Python and the software engineering discipline to build self-service platform tooling and internal SDKs that engineers actually adoptClassical MLOps foundations: model registry and versioning (MLflow, Azure ML, or SageMaker), automated retraining and redeployment pipelines, drift and data quality monitoring, and pipeline orchestration with robust SQL
📌 Machine Learning (Vasanth Nagar)
🏢 InfoCepts
📍 Vasanth Nagar