08 Oct
|
Observance Solutions Private
|
India
08 Oct
Observance Solutions Private
India
Job Summary
We are seeking a Data Science Ops / AI & GenAI Ops Engineer to support the deployment, monitoring, governance, and maintenance of Data Science, Machine Learning, AI, and Generative AI solutions in production.
The engineer will work at the intersection of Data Science, AI Engineering, Data Engineering, and ML Operations, ensuring that models, LLM-based applications, AI agents, RAG pipelines, prompt workflows, and AI services are reliably integrated into operational environments.
The role has a strong focus on GenAI Operations, model lifecycle management, production monitoring, observability, CI/CD, automation, reliability, and incident management on AWS.
Logistics
- Location: Offshore – India
- Contract Type: Full-time
- Start Date: Immediate
- Duration: Until FY27-Q4, with possible extension
- Role: Data Science Ops / AI & GenAI Ops
- Shift Duration: 9 hours per day
- Time Zone: CST
- Required Coverage: Candidate must be available to work until 12:00 PM CST
Roles & ResponsibilitiesData Science & AI Pipeline Development
- Design, develop, and maintain production pipelines for Data Science, Machine Learning, AI, and GenAI solutions on AWS.
- Build reusable and scalable pipelines for model training, validation, deployment, inference, and monitoring.
- Support integration of AI/ML solutions with enterprise applications and operational workflows.
GenAI & LLM Operations
- Deploy and maintain LLM-based applications, RAG solutions, AI agents, prompt workflows, embeddings, vector search, and API-based AI services.
- Support GenAI applications built using Amazon Bedrock and related AWS AI services.
- Manage prompt versions, model configurations, embeddings, knowledge bases, and AI workflows.
- Implement operational controls around prompt quality, hallucination detection, guardrails, latency, throughput, and cost.
- Support evaluation and continuous improvement of GenAI applications in production.
CI/CD & Orchestration
- Design and maintain CI/CD pipelines using GitHub Actions, Apache Airflow, or similar orchestration tools.
- Automate model and AI application deployment processes.
- Implement automated testing, validation, deployment, and rollback mechanisms.
- Manage dependencies and deployment workflows across development, testing, and production environments.
Model Lifecycle Management
- Support the complete ML/AI model lifecycle, including experimentation, model registration, deployment, monitoring, retraining, and retirement.
- Work with AWS SageMaker, AWS Batch, and other AWS services for model execution and deployment.
- Support experiment tracking and model registry processes.
- Maintain reproducibility, version control, and governance of models and AI assets.
Observability & Monitoring
- Implement monitoring and observability for:
- Model performance and drift
- Data quality
- Prediction quality
- Prompt quality
- Hallucination indicators
- API latency
- Throughput
- Availability
- Token consumption
- Infrastructure utilization
- AI/GenAI operational costs
- Develop dashboards, alerts, and operational metrics using appropriate monitoring and visualization tools.
Production Operations & Incident Management
- Monitor AI/ML workloads and production pipelines.
- Troubleshoot and resolve production incidents, model failures, pipeline issues, and integration problems.
- Perform root-cause analysis and implement corrective and preventive actions.
- Ensure production environments meet reliability, scalability, performance, and security requirements.
Collaboration & Documentation
- Collaborate with Data Scientists, Data Engineers, AI Engineers, ML Engineers, and DevOps teams.
- Establish and document standards, operating procedures, deployment processes,
and troubleshooting guides.
- Help define best practices for productionizing Data Science and GenAI solutions.
- Support governance and operational readiness of AI/ML solutions.
Performance & Cost Optimization
- Optimize AI/ML workloads for performance, reliability, scalability, and cost-effectiveness.
- Monitor AWS resource utilization and identify opportunities for cost optimization.
- Optimize inference workloads, pipelines, storage, compute, and GenAI API usage.
Must-Have SkillsProgramming & Data Processing
- Python
- PySpark
- SQL
Data Platforms
- Databricks
- Delta Lake
Machine Learning & AI
- Scikit-learn
- TensorFlow
- PyTorch
- Amazon Bedrock
- LLMs
- RAG architectures
- Machine Learning model deployment concepts
AI & GenAI Engineering
- Prompt Engineering
- Fine-Tuning
- AI Agents
- Amazon Bedrock Knowledge Bases
- Guardrails
- Embeddings
- Vector search
- LLM application development and operations
Model Development & Deployment
- AWS SageMaker
- AWS Batch
- Experiment tracking
- Model registry
- Model lifecycle management
Workflow Orchestration
- Apache Airflow
Data Storage
- AWS S3
Cloud
- AWS
DevOps & CI/CD
- Git
- GitHub Actions
- CI/CD principles and automation
Valuable-to-Have Skills
- Experience with Vector Databases such as OpenSearch, Pinecone, Weaviate, or similar technologies.
- Strong Data Engineering experience.
- Experience with ETL/ELT pipelines.
- Knowledge of data quality and data governance practices.
- Experience with Tableau, Power BI, Datadog, or similar visualization/observability platforms.
- Experience with GenAI evaluation frameworks and LLM observability tools.
- Familiarity with MLOps/LLMOps platforms and practices.
- Experience with AWS security, IAM, networking, and infrastructure automation.
Role Focus
Primary Focus: Data Science Operations with a strong emphasis on Generative AI / LLM Operations.
Pay: ₹900,000.00 - ₹1,200,000.00 per year
Work Location: Remote
📌 Data Science Ops / AI & GenAI Ops Engineer (India)
🏢 Observance Solutions Private
📍 India