Job Description:
Job Purpose
- Responsibilities
- Design and deploy RAG pipelines using LangChain and vector databases
- Build microservices to serve LLMs via APIs; optimize for performance and cost
- Use AI assistants for code generation, testing, and documentation—refining outputs through expert review
- Track experiments via MLflow and deploy using CI/CD, Docker, and Kubernetes
- Define metrics for generative AI performance (e.g. latency, accuracy, user satisfaction)
- Share findings and lead technical discussions with NYC stakeholders
- Collaboration Style
- Daily video stand-ups during overlap window
- Asynchronous updates and documentation via Microsoft Teams
- Strong emphasis on both written and oral communication
- Knowledge and Experience
- 5+ years in software development, 2+ with GenAI or NLP
- Proficiency in Python (FastAPI, asyncio); robust with Hugging Face, OpenAI APIs, LangChain
- Experience with vector databases (e.g. Pinecone, Weaviate) and MLflow
- Strong judgment in using GenAI tools like ChatGPT and Copilot productively
- Excellent written and oral English communication
- Familiarity with financial-market data (e.g. order books, trading events)
- Nice-to-Have Knowledge and Experience
- M.S. or Ph.D. in ML, AI, or related field
- Peer-reviewed publications in related fields
- Schedule
Working Hours: 12:00 PM - 8:00 PM IST (with overlap from 8:30 AM - 12:00 PM EST)