11 Sep
|
Ailogic Neural Network
|
Hyderabad
11 Sep
Ailogic Neural Network
Hyderabad
About Us /n AiLogic Neural Network Pvt Ltd is an AI-driven product company focused on building advanced language technology solutions, including machine translation, document intelligence, and large-scale NLP systems. We are looking for a highly motivated Data Engineer to join our growing AI team and contribute to the development of scalable data processing pipelines for NLP and LLM applications. /n Roles & Responsibilities /n /n
- Design, develop, and maintain scalable data pipelines for processing large volumes of structured and unstructured data.
/n
- Build document ingestion and processing workflows for PDFs, scanned documents, HTML pages, and other text sources.
/n
- Implement OCR, PDF parsing, HTML parsing, and text extraction pipelines.
/n
- Develop document chunking and preprocessing frameworks for NLP and LLM-based applications.
/n
- Work with Hugging Face models and NLP libraries for text processing tasks.
/n
- Create and optimize data transformation workflows using Python, Apache Spark, and Spark SQL.
/n
- Develop and manage Vector Database pipelines for embedding storage and retrieval.
/n
- Implement text normalization, sentence segmentation, deduplication, and data quality processes.
/n
- Design and implement data masking, classification, and categorization solutions.
/n
- Collaborate with AI/ML engineers to prepare datasets for model training and inference.
/n
- Optimize large-scale data processing workflows for performance, scalability, and cost efficiency.
/n
- Maintain CI/CD pipelines and follow software engineering best practices.
/n
- Monitor, troubleshoot, and improve production data processing systems.
/n /n Mandatory Skills /n /n
- Solid experience with Python programming.
/n
- Hands-on experience in NLP concepts such as:
/n
- Tokenization
/n
- Text Processing
/n
- Hugging Face Transformers
/n
- Experience in:
/n
- PDF Parsing
/n
- OCR
/n
- HTML Parsing
/n
- Text Extraction
/n
- Document Chunking
/n
- Experience with Apache Spark and Spark SQL.
/n
- Working knowledge of Vector Databases.
/n
- Good understanding of Git and CI/CD practices.
/n
- Experience building data pipelines and ETL workflows.
/n
- Strong debugging and problem-solving skills.
/n /n Preferred Skills /n /n
- Text Normalization
/n
- Sentence Segmentation
/n
- Exact Deduplication and Near Deduplication
/n
- Data Masking
/n
- Data Classification & Categorization
/n
- Embedding Generation and Retrieval Pipelines
/n
- Large-scale Document Processing Systems
/n
- RAG (Retrieval-Augmented Generation) Pipelines
/n /n Performance Optimization Skills /n /n
- CPU Distribution and Parallel Processing
/n
- Pre-batch Generation Techniques
/n
- Chunking Optimization Strategies
/n
- Stream Processing vs Batch Processing
/n
- GPU and CPU Parallel Distribution
/n
- CUDA Optimization
/n
- PyTorch Performance Tuning
/n
- Spark Performance Optimization
/n /n Qualifications /n /n
- Bachelor's or Master's degree in Computer Science, Data Science, Information Technology, or a related field.
/n
- 2–3 years of experience in Data Engineering, NLP Engineering, or AI Data Processing.
/n
- Experience working with large-scale datasets and distributed computing frameworks.
/n
📌 Data Engineer (Hyderabad)
🏢 Ailogic Neural Network
📍 Hyderabad