Data Scientist / Data Engineer (Hyderabad)

Data Scientist / Data Engineer (Hyderabad)

17 Sep
|
Zeominds It Solutions
|
Hyderabad

17 Sep

Zeominds It Solutions

Hyderabad

Job Designation: Data Scientist NLP / Transformer Models / LLM

Role Overview:

As a Data Scientist NLP / Transformer Models, you will work on the training, fine-tuning, evaluation, optimization, and deployment of Transformer-based NLP systems and LLMs. You will collaborate closely with engineering and research teams to develop scalable AI models and contribute to advanced multilingual AI solutions.

The ideal candidate should be comfortable working with model training and fine-tuning pipelines, large-scale datasets, tokenization, Transformer architectures, Seq2Seq and encoder-decoder models, GPU-based training environments, NLP evaluation, and model optimization.

Experience: 2+ Yrs

Eligibility Criteria

- 2+ years of professional experience in Data Science, Machine Learning, NLP, or a closely related field.
- Strong hands-on experience with Transformers, LLMs, NLP, and Seq2Seq architectures.
- Strong understanding of attention mechanisms and Transformer architecture.
- Proven experience training and fine-tuning models using custom datasets.
- Hands-on experience with PyTorch and/or TensorFlow.
- Solid experience with Hugging Face Transformers and related tooling.
- Understanding of BPE, SentencePiece, and other tokenization techniques.
- Experience with NLP evaluation metrics such as BLEU, perplexity, and accuracy.
- Experience working with GPU-based model training and optimization.
- Knowledge of distributed training concepts and GPU optimization.
- Experience with Pinecone or Milvus for embeddings and semantic search.
- Strong Python programming and problem-solving skills.

Salary: 25% - 30% hike on last drawn CTC and based upon current market standards

Key Responsibilities

- Lead the architecture and development of Transformer-based AI systems.
- Drive technical direction for NLP, LLM, and multilingual AI initiatives.
- Train, fine-tune, and optimize Transformer models using large-scale custom datasets.
- Work with Seq2Seq and encoder-decoder architectures for machine translation and text generation.
- Develop and optimize model training pipelines for GPU-based environments.
- Apply LoRA, QLoRA, PEFT, quantization, and distributed training techniques.
- Design and improve tokenization pipelines using BPE, SentencePiece, and related techniques.
- Evaluate models using BLEU, perplexity, accuracy, and custom benchmarks.
- Analyze model performance and identify opportunities for optimization.
- Work with embeddings, semantic search, and vector databases for NLP and LLM applications.
- Collaborate with platform and infrastructure teams on scalable GPU infrastructure.
- Contribute to research,



experimentation, proof-of-concepts, and production AI systems.
- Mentor and guide junior ML engineers and researchers.
- Document experiments, findings, model performance, and technical approaches.

Required Technical Skills

- Python; PyTorch / TensorFlow
- Hugging Face; Transformers; LLMs; Seq2Seq / Encoder-Decoder Models; Attention Mechanisms
- BPE / SentencePiece; NLP evaluation and benchmarking
- LoRA; QLoRA; PEFT; Quantization; DeepSpeed; FSDP; Distributed Training
- AWS; GPU Infrastructure; Docker; Kubernetes
- Pinecone; Milvus; Embeddings; Semantic Search

Strongly Preferred Skills

- LoRA / QLoRA / PEFT
- DeepSpeed / FSDP
- Quantization and model optimization
- RLHF / SFT
- Multilingual model development
- Machine Translation systems
- Speech or sequence modeling
- Large-scale NLP datasets and distributed model training
- GPU performance optimization
- AI/ML/NLP research projects or publications
- Experience taking research models into production

What We Are Looking For

- Strong ownership and builder mindset
- Deep understanding of Transformer architectures
- Strong Python and ML engineering skills
- Ability to independently design and execute experiments
- Strong analytical and problem-solving abilities
- Research-oriented thinking
- Ability to work with large datasets and GPU infrastructure
- Strong communication and collaboration skills
- Ability to mentor junior ML engineers and researchers
- Ability to bridge AI research and production engineering

Job Designation: Data Engineer Experience: 2-3 yrs

Eligibility Criteria

- Bachelor's or Master's degree in Computer Science, Data Science, Information Technology, Artificial Intelligence, or a related field.
- 23 years of experience in Data Engineering, NLP Engineering, AI Data Processing, or a similar role.
- Experience working with large-scale datasets and distributed computing frameworks.

Salary: 25% - 30% hike on last drawn CTC and based upon current market standards

Key Responsibilities

- Design, develop, and maintain scalable data pipelines for processing large volumes of structured and unstructured data.
- Build document ingestion and processing workflows for PDFs, scanned documents, HTML pages,



and other text sources.
- Develop OCR, PDF parsing, HTML parsing, and text extraction pipelines.
- Create document chunking and preprocessing frameworks for NLP and LLM applications.
- Work with Hugging Face Transformers and NLP libraries for text processing and document understanding.
- Develop and optimize ETL pipelines using Python, Apache Spark, and Spark SQL.
- Design and maintain Vector Database pipelines for embedding generation, storage, and retrieval.
- Implement text normalization, sentence segmentation, deduplication, and data quality workflows.
- Build data masking, classification, and categorization solutions.
- Collaborate closely with AI/ML Engineers to prepare datasets for model training, fine-tuning, and inference.
- Optimize large-scale data processing workflows for scalability, performance, and cost efficiency.
- Maintain CI/CD pipelines and follow software engineering best practices.
- Monitor, troubleshoot, and improve production-grade data processing systems.

Required Technical Skills

- Strong proficiency in Python.
- Hands-on experience with NLP concepts, including:

- Tokenization

- Text Processing

- Hugging Face Transformers

- Experience with

- PDF Parsing

- OCR

- HTML Parsing

- Text Extraction

- Document Chunking
- Solid knowledge of Apache Spark and Spark SQL.
- Experience building ETL and data processing pipelines.
- Working knowledge of Vector Databases.
- Familiarity with Git, version control, and CI/CD workflows.
- Strong debugging, analytical, and problem-solving skills.

Preferred Skills

- Text Normalization
- Sentence Segmentation
- Exact and Near Deduplication
- Data Masking
- Data Classification & Categorization
- Embedding Generation and Retrieval Pipelines
- Large-scale Document Processing Systems
- Retrieval-Augmented Generation (RAG) Pipelines

Performance Optimization Experience in one or more of the following areas is an added advantage:

- CPU Distribution and Parallel Processing
- Pre-batch Generation Techniques
- Chunking Optimization Strategies
- Stream Processing vs Batch Processing
- GPU & CPU Parallel Distribution
- CUDA Optimization
- PyTorch Performance Tuning
- Spark Performance Optimization

What We Are Looking For

- Strong programming fundamentals.
- Good logical and analytical thinking.
- Interest in Data Engineering and AI/ML.
- Ability to learn quickly and work in a team environment.
- Good communication and problem-solving skills.
- Candidates who are willing to work on-site in Hyderabad.

Application Deadline: 21st September 2026

📌 Data Scientist / Data Engineer (Hyderabad)
🏢 Zeominds It Solutions
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: data scientist / data engineer (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: data scientist / data engineer (hyderabad) / hyderabad