26 Sep
|
InfoCepts
|
Vasanth Nagar
26 Sep
InfoCepts
Vasanth Nagar
Purpose of the Position: The AI Platform Engineer- MLOps/LLMOps will be responsible for designing, implementing, and operationalizing enterprise-grade AI, Machine Learning, and Generative AI solutions. The role will focus on deploying and monitoring AI applications, establishing governance frameworks, and enabling secure, reliable, and cost-effective AI operations across cloud environments. The individual will provide technical leadership and drive best practices for AI engineering, MLOps, and LLMOps initiatives.
Key Result Areas and Activities:
Reliability and Availability of AI services and solutions
Design scalable AI/ML and Generative AI deployment architectures.
Define enterprise standards for model deployment, monitoring, and lifecycle management.
Design and operationalize RAG, agentic AI, prompt management, vector databases, and LLM integration frameworks.
MLOps & LLMOps Implementation
Establish CI/CD pipelines for machine learning and LLM-based applications.
Automate model training, deployment, versioning, and rollback processes.
Optimize LLM performance, scalability, and cost efficiency including Inference cost management
Model Governance & Reliability
Implement frameworks for model monitoring, evaluation, drift detection, observability, and compliance.
Ensure responsible AI and governance standards are followed.
Technical Leadership & Stakeholder Collaboration
Mentor AI engineers, junior platform engineers and data scientists.
Collaborate with business and technology teams to translate AI use cases into production-ready solutions.
Essential Skills:
LLM serving and inference optimization and running an inference gateway that routes across providers and open-weight models with failover
LLM observability and tracing: end-to-end request tracing through retrieval, prompt, tool calls and generation using LangSmith or Azure AI Foundry
Evaluation as a platform capability: building the harness product teams use to regression-test AI behaviour, run LLM-as-judge and offline evals, and gate releases on measured quality rather than judgement calls
Prompt, model and config lifecycle management: versioning prompts and system messages like code, secure rollout and rollback, A/B and shadow testing of model or prompt changes in production
Token and inference cost engineering: spend attribution by team, feature and tenant; caching and model-routing strategies; GPU utilisation and right-sizing
Retrieval infrastructure operations: running and tuning vector or hybrid search at scale, embedding pipelines, index refresh strategies, and retrieval quality monitoring
Kubernetes and GPU infrastructure in production: Docker, autoscaling, node pools, GPU scheduling and sharing, plus infrastructure as code (Terraform or Bicep) and GitOps (Argo CD or Flux)
Strong Python and the software engineering discipline to build self-service platform tooling and internal SDKs that engineers actually adopt
Classical MLOps foundations: model registry and versioning (MLflow, Azure ML, or SageMaker), automated retraining and redeployment pipelines, drift and data quality monitoring, and pipeline orchestration with strong SQL
📌 Machine Learning (Vasanth Nagar)
🏢 InfoCepts
📍 Vasanth Nagar