DESCRIPTION:
Selection Monitoring team is responsible for making the biggest catalog on the planet even bigger. In order to drive expansion of the Amazon catalog, we develop advanced ML/AI technologies to process billions of products and algorithmically find products not already sold on Amazon. We work with structured, semi-structured and Visually Rich Documents using deep learning, NLP and image processing.
You will encounter many challenges, including:
- Scale (build models to handle billions of pages),
- Accuracy (requirements for precision and recall)
- Speed (generate predictions for millions of recent or changed pages with low latency)
- Diversity (models need to work across different languages, market places and data sources)
You will help us to
- Build a scalable system which can algorithmically extract information from world wide web.
- Intelligently cluster web pages, segment and classify regions, extract relevant information and structure the data available on semi-structured web.
- Build systems that will use existing Knowledge Base to perform open information extraction at scale from visually rich documents.
Key job responsibilities
- Use AI, NLP and advances in LLMs/SLMs and agentic systems to create scalable solutions for business problems.
- Efficiently Crawl web, Automate extraction of relevant information from large amounts of Visually Rich Documents and optimize key processes.
- Design, develop, evaluate and deploy, innovative and highly scalable ML models, esp. leveraging latest advances in RL-based fine tuning methods like DPO, GRPO etc.
- Work closely with software engineering teams to drive real-time model implementations.
- Establish scalable, efficient, automated processes for large scale model development, model validation and model maintenance.
- Lead projects and mentor other scientists, engineers in the use of ML techniques.
- Publish innovation in research forums.
BASIC QUALIFICATIONS:
- Experience programming in Java, C++, Python or r