24 Sep
|
InfoCepts
|
Bengaluru
24 Sep
InfoCepts
Bengaluru
Purpose of the Position: The AI Platform Engineer- MLOps/LLMOps will be responsible for designing, implementing, and operationalizing enterprise-grade AI, Machine Learning, and Generative AI solutions. The role will focus on deploying and monitoring AI applications, establishing governance frameworks, and enabling secure, reliable, and cost-effective AI operations across cloud environments. The individual will provide technical leadership and drive best practices for AI engineering, MLOps, and LLMOps initiatives.
Key Result Areas and Activities:
- Reliability and Availability of AI services and solutions
- Design scalable AI/ML and Generative AI deployment architectures.
- Define enterprise standards for model deployment, monitoring, and lifecycle management.
- Design and operationalize RAG, agentic AI, prompt management, vector databases, and LLM integration frameworks.
- MLOps & LLMOps Implementation
- Establish CI/CD pipelines for machine learning and LLM-based applications.
- Automate model training, deployment, versioning, and rollback processes.
- Optimize LLM performance, scalability, and cost efficiency including Inference cost management
- Model Governance & Reliability
- Implement frameworks for model monitoring, evaluation, drift detection, observability, and compliance.
- Ensure responsible AI and governance standards are followed.
- Technical Leadership & Stakeholder Collaboration
- Mentor AI engineers, junior platform engineers and data scientists.
- Collaborate with business and technology teams to translate AI use cases into production-ready solutions.
- Essential Skills:
- LLM serving and inference optimization and running an inference gateway that routes across providers and open-weight models with failover
- LLM observability and tracing: end-to-end request tracing through retrieval, prompt, tool calls and generation using LangSmith or Azure AI Foundry
- Evaluation as a platform capability: building the harness product teams use to regression-test AI behaviour, run LLM-as-judge and offline evals, and gate releases on measured quality rather than judgement calls
- Prompt, model and config lifecycle management: versioning prompts and system messages like code, safe rollout and rollback, A/B and shadow testing of model or prompt changes in production
- Token and inference cost engineering: spend attribution by team, feature and tenant; caching and model-routing strategies; GPU utilisation and right-sizing
- Retrieval infrastructure operations: running and tuning vector or hybrid search at scale, embedding pipelines, index refresh strategies, and retrieval quality monitoring
- Kubernetes and GPU infrastructure in production: Docker, autoscaling, node pools, GPU scheduling and sharing, plus infrastructure as code (Terraform or Bicep) and GitOps (Argo CD or Flux)
- Solid Python and the software engineering discipline to build self-service platform tooling and internal SDKs that engineers actually adopt
- Classical MLOps foundations: model registry and versioning (MLflow, Azure ML, or SageMaker), automated retraining and redeployment pipelines, drift and data quality monitoring, and pipeline orchestration with strong SQL
📌 Machine Learning (Bengaluru)
🏢 InfoCepts
📍 Bengaluru