Role OverviewWe are looking for a versatile Data Scientist who can build robust data pipelines — from web scraping to AI/ML deployment. The ideal candidate is comfortable working across the full data and AI lifecycle: extracting data at scale, transforming and storing it reliably, and using it to train, deploy, and maintain machine learning models in production.
Key Responsibilities
- Design, build, and maintain scalable web scraping scripts/pipelines using Python (e.g., BeautifulSoup, Scrapy, Selenium, Playwright)
- Handle dynamic websites, pagination, anti-bot mechanisms, proxies, and rate-limiting strategies
- Clean, transform, and normalize scraped data (structured/unstructured) before storage
- Design MongoDB schemas and collections optimized for the type of data being handled
- Implement logic to identify and update unique/duplicate records efficiently (upserts, deduplication strategies)
- Schedule and monitor data pipelines/jobs (via cron, Airflow, or similar orchestration tools)
- Ensure data quality, consistency, and integrity across pipelines, including error logging, retries, and failure recovery
- Support end-to-end AI/ML lifecycle: data collection, preprocessing, feature engineering, model selection, training, and validation
- Fine-tune machine learning/deep learning models and evaluate performance against business requirements
- Package and deploy models into production environments (APIs, batch pipelines, etc.)
- Implement and maintain MLOps practices — versioning (models & data), CI/CD for ML, monitoring model performance/drift
- Collaborate with cross-functional teams to integrate AI models with existing data pipelines, including scraped/transformed data as model input
Required Skills
Core:
- Strong proficiency in Python (writing clean, modular, production-grade code)
- Hands-on experience with MongoDB (schema design, aggregation pipelines, indexing, upsert/dedup logic)
- Web scraping tools/libraries: BeautifulSoup, Scrapy, Selenium, Playwright, or similar
- Data transformation/manipulation using Pandas / NumPy
AI/ML:
- Understanding of the end-to-end AI/ML workflow — data prep, training, evaluation, deployment
- Familiarity with ML/DL frameworks: Scikit-learn, TensorFlow, PyTorch
- Exposure to MLOps tools: MLflow, DVC, Docker, Kubernetes (basic), CI/CD pipelines
- Experience with model deployment (REST APIs via FastAPI/Flask, or cloud ML services)
Positive to Have:
- Experience with cloud platforms (AWS / GCP / Azure) for storage, compute, and ML services
- Familiarity with LLMs / NLP (e.g., Hugging Face, LangChain) if relevant to use case
- Knowledge of proxy rotation, CAPTCHA-handling, and anti-scraping evasion techniques
- Experience with workflow orchestration (Airflow, Prefect)
- Version control (Git) and Agile development practices
Soft Skills
- Strong problem-solving mindset, especially around handling unstructured/messy data
- Ability to independently manage multiple workstreams across data engineering and AI/ML
- Good documentation habits for pipelines and models
- Comfortable working in a fast-paced, evolving tech environment
Qualifications
- Bachelor's/Master's degree in Computer Science, Data Science, Engineering, or related field
- Prior experience with production-grade web scraping and/or ML systems is a solid plus
- Portfolio/GitHub showcasing scraping projects and/or ML models is preferred
📌 Data Scientist (India)
🏢 VCBay
📍 India