08 Oct
|
VY SYSTEMS PRIVATE
|
Bengaluru
08 Oct
VY SYSTEMS PRIVATE
Bengaluru
Role Overview
We are looking for an experienced Data Scientist – Agentic AI with strong expertise in Python, Machine Learning, Generative AI, Large Language Models (LLMs), RAG and Agentic AI.
The ideal candidate should have hands-on experience in developing, fine-tuning, evaluating and deploying machine learning and GenAI solutions. The candidate should be comfortable working with open-source LLMs, LangChain/LangGraph, PySpark and AI observability/tracing frameworks.
The role involves building intelligent AI systems that can reason, use tools, retrieve information and execute multi-step tasks using Agentic AI architectures.
Mandatory Technical Skills
1. Data Science & Python
- 5+ years of experience in Data Science / Machine Learning / AI.
- Strong programming experience in Python.
- Robust understanding of data analysis, feature engineering and statistical techniques.
- Experience with Python ML and data science libraries such as:
- NumPy
- Pandas
- Scikit-learn
- Matplotlib / Seaborn
- Good understanding of data preprocessing, exploratory data analysis and experimentation.
2. Machine Learning & Statistics
- Strong understanding of Machine Learning fundamentals.
- Experience with supervised and unsupervised learning techniques.
- Knowledge of:
- Regression
- Classification
- Clustering
- Feature Engineering
- Model Selection
- Hyperparameter Tuning
- Cross-validation
- Strong understanding of Statistics / ML fundamentals.
- Ability to interpret model performance and statistical results.
3. Generative AI / LLM
- Strong hands-on experience with Generative AI and Large Language Models (LLMs).
- Understanding of Transformer architecture and modern LLM-based applications.
- Experience working with commercial or open-source LLMs.
- Strong understanding of:
- Prompt Engineering
- Context Management
- Embeddings
- Tokenization
- LLM inference
- Hallucination mitigation
4. RAG – Retrieval Augmented Generation
- Strong hands-on experience developing RAG applications.
- Experience with:
- Document ingestion
- Chunking
- Embeddings
- Vector search
- Semantic search
- Retrieval pipelines
- Context retrieval
- Reranking
- Ability to optimize RAG pipelines for relevance, accuracy and latency.
- Experience integrating LLMs with enterprise knowledge sources.
5. Agentic AI
- Hands-on experience building Agentic AI / AI Agent solutions.
- Understanding of agent architecture and multi-step reasoning workflows.
- Experience with:
- AI Agents
- Multi-Agent systems
- Tool Calling
- Function Calling
- Agent orchestration
- Planning and reasoning workflows
- Memory
- Workflow automation
- Ability to build agents that can interact with tools, APIs, databases and external systems.
6. LangChain / LangGraph
- Strong hands-on experience with LangChain and/or LangGraph.
- Experience building LLM workflows and agent-based applications.
- Understanding of:
- Chains
- Agents
- Tools
- State management
- Graph-based workflows
- Agent orchestration
- Retrieval workflows
- Experience designing scalable Agentic AI workflows.
7. LLM Fine-Tuning
- Hands-on experience with LLM fine-tuning.
- Understanding of techniques such as:
- Supervised Fine-Tuning (SFT)
- Parameter-Efficient Fine-Tuning
- LoRA
- QLoRA
- Experience preparing datasets for fine-tuning.
- Ability to evaluate fine-tuned models against baseline models.
- Understanding of model optimization and inference considerations.
8. BERT / LLaMA / Open-Source LLMs
Experience working with one or more open-source / transformer-based models such as:
- BERT
- LLaMA / Llama
- Mistral
- Gemma
- Qwen
- Other open-source LLMs
Candidate should understand model loading, inference, fine-tuning and evaluation.
9. PySpark
- Strong experience with PySpark for large-scale data processing.
- Experience working with large datasets and distributed data processing.
- Knowledge of:
- Data transformations
- Data cleaning
- Aggregations
- Joins
- Spark SQL
- Performance optimization
- Ability to build scalable data processing pipelines.
10. Model Validation & Evaluation
- Experience validating and evaluating ML and GenAI models.
- Understanding of traditional ML evaluation metrics.
- Experience evaluating LLM/RAG applications using relevant quality metrics.
- Ability to compare model performance and identify areas for improvement.
- Experience with:
- Accuracy
- Precision
- Recall
- F1 Score
- ROC-AUC
- Retrieval metrics
- LLM response quality
- Groundedness / relevance
- Experience designing evaluation datasets and test cases is preferred.
11. AI Tracing / Observability
- Experience with AI/LLM tracing and observability.
- Ability to monitor AI applications in production.
- Experience tracking:
- LLM requests/responses
- Latency
- Token usage
- Errors
- Retrieval performance
- Agent/tool execution
- Model performance
- Exposure to tools/frameworks such as LangSmith, OpenTelemetry, Arize Phoenix, MLflow or similar is preferred.
12. Model Deployment
- Experience deploying ML/LLM/GenAI solutions into production.
- Exposure to cloud and/or on-premise model deployment.
- Experience with model serving, APIs and production inference.
- Knowledge of deployment environments such as:
- AWS
- Azure
- GCP
- On-premise infrastructure
- Experience with Docker, APIs and CI/CD is an advantage.
Key Responsibilities
- Design, develop and deploy Data Science, Machine Learning and GenAI solutions.
- Build production-ready RAG and Agentic AI applications.
- Develop intelligent agents capable of tool calling, reasoning and multi-step task execution.
- Build LLM-powered applications using LangChain/LangGraph.
- Work with open-source LLMs including BERT, LLaMA and other transformer-based models.
- Fine-tune LLMs for specific business use cases.
- Develop scalable data processing pipelines using PySpark.
- Perform data analysis, feature engineering and statistical modeling.
- Develop and maintain model validation and evaluation frameworks.
- Evaluate ML and LLM models using appropriate performance and quality metrics.
- Implement AI tracing, monitoring and observability for production GenAI systems.
- Deploy models and AI applications in cloud or on-premise environments.
- Optimize model performance, response quality, latency and cost.
- Troubleshoot issues related to model inference, retrieval, agents and LLM workflows.
- Collaborate with Data Scientists, ML Engineers, Software Engineers and business stakeholders.
- Convert business requirements into scalable AI/ML solutions.
Good to Have
- Experience with Vector Databases such as:
- FAISS
- Pinecone
- Weaviate
- Milvus
- Chroma
- Azure AI Search
- Experience with MLflow or similar ML lifecycle tools.
- Experience with Docker/Kubernetes.
- Experience with REST APIs / FastAPI.
- Knowledge of cloud AI/ML services.
- Experience with MLOps / LLMOps.
- Experience with multi-agent frameworks other than LangChain/LangGraph.
- Experience working with enterprise GenAI applications.
Ideal Candidate Profile
The ideal candidate should be a Data Scientist / ML Engineer with strong GenAI and Agentic AI experience, rather than a pure Python developer.
A strong candidate would typically have:
Data Science + Python + ML + Statistics + GenAI/LLM + RAG + Agentic AI + LangChain/LangGraph + LLM Fine-Tuning + Open-Source LLMs + PySpark + Model Evaluation + AI Observability + Model Deployment.
Core Mandatory Skills
Data Science, Python, Machine Learning, Statistics/ML Fundamentals, GenAI/LLM, RAG, Agentic AI, LangChain/LangGraph, LLM Fine-Tuning, BERT/LLaMA/Open-Source LLMs, PySpark, Model Validation/Evaluation, AI Tracing/Observability, Cloud/On-Prem Model Deployment.
Skills:- Data Science, Python, Machine Learning (ML), Statistics/ML Fundamentals, GenAI/LLM, Agentic AI, RAG, LangChain/LangGraph,, AI Tracing/Observability., Model Validation/Evaluation, PySpark, BERT/LLaMA/Open-Source LLMs and LLM Fine-Tuning,
📌 Data Scientist â Agentic AI (Bengaluru)
🏢 VY SYSTEMS PRIVATE
📍 Bengaluru