Job Description:
Job Purpose
Responsibilities
Design and deploy RAG pipelines using LangChain and vector databases
Build microservices to serve LLMs via APIs; optimize for performance and cost
Use AI assistants for code generation, testing, and documentation—refining outputs through expert review
Track experiments via MLflow and deploy using CI/CD, Docker, and Kubernetes
Define metrics for generative AI performance (e.g. latency, accuracy, user satisfaction)
Share findings and lead technical discussions with NYC stakeholders
Collaboration Style
Daily video stand-ups during overlap window
Asynchronous updates and documentation via Microsoft Teams
Robust emphasis on both written and oral communication
Knowledge and Experience
5+ years in software development, 2+ with GenAI or NLP
Proficiency in Python (FastAPI, asyncio); robust with Hugging Face, OpenAI APIs, LangChain
Experience with vector databases (e.g. Pinecone, Weaviate) and MLflow
Robust judgment in using GenAI tools like ChatGPT and Copilot productively
Excellent written and oral English communication
Familiarity with financial-market data (e.g. order books, trading events)
Nice-to-Have Knowledge and Experience
M.S. or Ph.D. in ML, AI, or related field
Peer-reviewed publications in related fields
Schedule
Working Hours: 12:00 PM - 8:00 PM IST (with overlap from 8:30 AM - 12:00 PM EST)