Seeking a strong Data Engineer / AI Engineer with expertise in building and operationalizing large-scale AI and NLP solutions on cloud platforms. The ideal candidate should have hands-on experience integrating AI/LLM models into production workflows, developing scalable data pipelines, and processing large volumes of multilingual unstructured text.
Key strengths should include:
• Proficiency in Python and SQL with experience deploying AI/NLP solutions such as document classification, entity extraction, NER, PII masking, de-identification, hybrid search, and LLM integrations.
• Solid knowledge of Apache Airflow for orchestrating end-to-end data pipelines and automating batch processing workflows.
• Experience working with AWS services including S3, Athena, Glue, Fargate, EKS, SQS, and Step Functions.
• Capability to design and maintain large-scale document processing systems handling complex JSON structures, embedded documents, and multilingual content.
• Familiarity with vector search and retrieval systems, including embeddings, pgvector, PostgreSQL/Aurora, GIN indexes,
and full-text search.
• Experience with ML lifecycle management using MLflow, Databricks/Azure Databricks, model deployment, monitoring, and evaluation frameworks.
• Strong DevOps practices including GitHub-based development, CI/CD pipelines, schema management, and production support.
Responsibilities
What You Will Do
AI Module Integration & Inference Pipelines
• Integrate and adjust inference pipelines for NLP modules including document classification, entity extraction, de-identification (DEID), and LLM-based early trend detection
• Connect DS-coded AI modules into end-to-end production workflows via Airflow DAGs on AWS EKS
• Build and tune hybrid search pipelines combining GTE multilingual dense embeddings with GIN lexical search on Aurora PostgreSQL
• Integrate with OpenAI-based API platform for multilingual query expansion and LLM-driven trend detection
Document Processing & Pars
📌 AI MLOPS/LLMOps Engineer (Gurugram)
🏢 EXL
📍 Gurugram