Urgent Opening For Senior Machine Learning Engineer (NLP)- Pune

Urgent Opening For Senior Machine Learning Engineer (NLP)- Pune

12 Sep
|
Virtusa
|
Pune

12 Sep

Virtusa

Pune

Dear Candidate,

Please find the job description below,

Senior Machine Learning Engineer (NLP)

An experienced, production-focused Machine Learning Engineer specializing in Natural Language Processing (NLP) and high-performance backend systems for the academic publishing domain.

Responsible for the full ML lifecycle — from data labeling and model training to deploying concurrent, production-grade inference APIs that process complex textual data at scale.

Areas of Knowledge & Expertise

1. Advanced Natural Language Processing & Information Extraction

Building, extending, and training custom spaCy (v3.6+) pipelines and components (NER, SpanCat, SpanFinder, custom tokenizers). Experience with SOTA frameworks such as Thinc, Flair, and zero-shot architectures like GLiNER. Knowledge of embedding generation, vector spaces, and up-to-date embedding models (e.g., google/embeddinggemma-300m). Mastery of fuzzy matching (Levenshtein distance), regex, and structured data formats (XML, PDF, DOCX, LaTeX), with a focus on academic manuscript structures and metadata standards.

2. Deep Learning & Model Optimization

Proficiency in PyTorch 2.x and PyTorch Lightning for reproducible model training across hardware accelerations (CUDA, MPS, CPU). Hands-on experience with the Hugging Face ecosystem for transfer learning and fine-tuning of LLMs and encoder-based transformers. Solid foundation in statistics, numerical computing (NumPy, SciPy, Pandas), and classical ML algorithms (e.g., DBSCAN clustering, Scikit-Learn pipelines).

3. High-Performance ML Operations & Backend Engineering

Advanced Python development using Asyncio and FastAPI to build high-throughput,



low-latency REST APIs. Understanding of the GIL, multi-threading, and memory/thread safety when serving heavy ML models in production web servers.

4. MLOps, Cloud & Data Lifecycle

Managing the data lifecycle with tools like Prodigy and spaCy Projects for active learning and gold-standard datasets. Using MLflow for experiment tracking, model registry, and reproducibility. Deploying and managing applications on GCP (Compute Engine, GKE, Vertex AI, Gemini API). Writing robust test suites (pytest, pytest-asyncio) and managing CI/CD via GitHub Actions, Jenkins, Docker, and Kubernetes. Monitoring via ELK and Dynatrace, and handling data streams via Kafka.

Technologies

Core

- Python 3.7+

- FastAPI, Uvicorn, Starlette, Pydantic v2 + pydantic-settings

- pytest, pytest-asyncio (unit + functional tests)

- asyncio, aiohttp

- Project packaging (setuptools + pyproject.toml)

- REST API development

ML

- PyTorch 2.x (CPU/GPU/MPS)

- CUDA 12.*+

- PyTorch Lightning

- Hugging Face Transformers

- Sentence_transformers

Classical ML

- scikit-learn, NumPy, Pandas, SciPy, matplotlib

- JupyterLab — exploration and training notebooks

- Basic knowledge of Linear Algebra

NLP

- Custom NLP pipeline design (spaCy, incl. transformer-encoder pipelines)





- Custom spaCy components: NER, SpanCat, Dependency parser, Sentencizer, SpanFinder, custom tokenizers/matchers

- Thinc, Flair NLP, GLiNER

- Retrieval and embedding models (e.g., google/embeddinggemma-300m)

Data processing

- spaCy Projects

- Prodigy

Text processing

- Fuzzy search

- Regex

- XML processing

LLM

- LLM-assisted data annotation with quality guards

- LLM inference servers: vLLM, Hugging Face TGI, llama.cpp

- Async Python (AsyncOpenAIClient) for high-throughput dataset processing

Cloud & integration

- Docker

- GitHub Actions

- GCP: Artifact Registry, GCS, Compute Engine, GKE, Logging, Vertex AI, Gemini API

- Jenkins, K8s, Dynatrace, ELK

- Apache Kafka

- MLflow

- Git

General knowledge

- NER and span classification

- Sequence labeling and document-level classification

- Transfer learning / fine-tuning transformers

- Train/eval/deploy lifecycle

- Data augmentation, information extraction, relation extraction, entity linking, information retrieval

- Fuzzy matching, clustering (DBSCAN)

- Precision/recall tradeoffs in information extraction

- Thread safety serving models in multi-threaded web servers

- XML processing; basic knowledge of doc(x), LaTeX, PDF formats

- Academic publishing domain — manuscript structure, metadata standards

Other tech as a plus

- Java, JavaScript

Kindly share the details below.
Name:
Total Experience
Relevant Experience
Contact No:
Email ID :
Current Location :
Preferred Location :
Current Company:
CTC:
ECTC:
Notice Period:

Regards,
Priyanka Bhosale

📌 Urgent Opening For Senior Machine Learning Engineer (NLP)- Pune
🏢 Virtusa
📍 Pune

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: urgent opening for senior machine learning engineer (nlp)- pune / pune

Subscribe to this job alert:

Get the latest job offers by email for: urgent opening for senior machine learning engineer (nlp)- pune / pune