05 Aug
|
Morningstar
|
Mumbai
05 Aug
Morningstar
Mumbai
Senior Data Scientist
Experience: Not Available to Not Available years
Location: Mumbai
Skills: machine learning, deep learning, GenAI, natural language processing, NLP, data analysis, ETL, APIs, data visualization, data cleaning, exploratory data analysis, Python, SQL, regression, classifiers, ensemble techniques, transformers, Open Source LLMs, Chat GPT, prompt engineering, statistical analysis, time series analysis, hypothesis testing
About the Role:
As a member of the Product and Engineering team at PitchBook, you will be part of a team of big thinkers, innovators, and problem solvers who strive to deepen the positive impact we have on our customers and our company every day. We value curiosity and the drive to find better ways of doing things. We thrive on customer empathy, which remains our focus when creating excellent customer experiences through product innovation. We know that greatness is achieved through collaboration and diverse points of view, so we work closely with partners around the globe. As a team, we assume positive intent in each other’s words and actions, value constructive discussions, and foster a respectful working environment built on integrity, growth, and business value. We invest heavily in our people, who are eager to learn and constantly improve. Join our team and grow with us! PitchBook’s Data Science and Machine Learning team has a clear mission: leverage cutting-edge machine learning and cloud technologies to automatically research private markets and improve the navigability of our platform. The Senior Data Scientist is responsible for using machine learning, deep learning, GenAI and natural language processing (NLP) to collect a high-volume of data for the PitchBook Platform and surface insights for our users. This role requires a strong desire to learn and deliver value through scalable machine learning/AI systems. A strong motivation to succeed is critical and everyone has the opportunity to shape the long-term direction of our team.
Primary Job Responsibilities:
• Leverage a sound understanding of data to improve decision-making,
and deliver insights through various data analysis techniques, including advanced skills in acquiring, processing, and wrangling data from multiple sources using tools and techniques like ETL, APIs, and data visualization.
• Apply various modeling techniques, while fusing data sources using pre-processing methods like transformation and normalization including employing data cleaning techniques for both structured and unstructured data, conducting exploratory data analysis, and extracting insights to inform business decisions through iterative exploration and hypothesis testing.
• Source additional information and solutions through research and relevant libraries, including assisting in applying best practice model fit testing, tuning, and validation techniques to assess model performance, considering data attributes like dataset size and partitioning.
• Develop and optimize machine learning & deep learning models such as regression, classifiers, ensemble techniques (bagging, boosting, stacking etc.), transformers based Open Source LLMs and Chat GPT including prompt engineering for NLP tasks such as Summarization, Sentiment Analysis, Information Extraction and collaborate with cross-functional teams to integrate solutions into products, besides staying updated with AI/ML advancements especially in GenAI.
• Proficiency in advanced statistical analysis and methods like regression analysis, time series analysis, and hypothesis testing.
• Creatively solve problems and develop effective solutions, along with an understanding of business goals and industry trends to align data projects with strategic objectives.
• Strong verbal and written communication skills to convey complex data insights to stakeholders,
combined with efficient project management to ensure timely and within-budget completion.
• Familiarity and practical skills with various data science and analytics tools, programming skills in languages like Python, and the ability to interpret moderately complex scripts.
Skills and Qualifications:
• Bachelor's/ Master's degree in Computer Science, Machine Learning, Statistics, or related fields.
• 4 + years of experience building predictive models (Classical ML models, Neural Network based Deep Learning models, Transformer based GenAi models etc.) using Python in a production environment.
• 4 + years of demonstrated experience with natural language processing (NLP).
• 4 + years of demonstrated experience with Python.
• Demonstrate experience using complex SQL queries for data extraction and transformation.
• Strong communication and data presentation skills.
• Strong problem-solving ability.
• Ability to communicate complex analysis in a clear, accurate, and actionable manner.
• Experience working closely with software development team(s) is a plus.
• Experience building scalable systems in production is a plus.
• Knowledge of financial services, investment, banking, venture capital, or private equity is desired but not required.
Working Conditions:
This employee will work in a standard office setting. Some flexibility with work-from-home is a possibility. Employees in this position use a computer on an ongoing basis throughout the day. Morningstar is an equal opportunity employer. Morningstar's hybrid work environment gives you the opportunity to collaborate in-person each week as we've found that we're at our best when we're purposely together on a regular basis. In most of our locations, our hybrid work model is four days in-office each week. A range of other benefits are also available to enhance flexibility as needs change. No matter where you are, you'll have tools and resources to engage meaningfully with your global colleagues. I10_MstarIndiaPvtLtd Morningstar India Private Ltd. (Delhi) Legal Entity
📌 Senior Data Scientist (Mumbai)
🏢 Morningstar
📍 Mumbai