10 Sep
|
hiringhood
|
Hyderabad
10 Sep
hiringhood
Hyderabad
Job Description
Data Scientist – AI/ML & Identity Management
n
Location: Hyderabad, India
n
Experience: 7–15 years
n
Employment Type: Full-time, Work from Office
n
Role Overview
n
We are looking for a Data Scientist with hands-on Python coding skills to develop and deploy AI/ML solutions for identity management. The role involves working with real-world identity data, developing NLP and name-matching algorithms, building ML models, and integrating them into production applications.
n
The candidate will also be required to understand, analyze, enhance, and extend an existing Python codebase. Since access to external LLM-based coding assistants such as ChatGPT or Claude may be limited due to the confidentiality of data, the candidate must have strong independent coding and problem-solving abilities.
n
Key Responsibilities
n
n
- Develop algorithms for name matching, entity resolution, record linkage, duplicate detection, and identity matching.
n
- Apply NLP and text similarity techniques including Jaro-Winkler, Levenshtein, TF-IDF, n-grams, token-based similarity, phonetic matching, embeddings, and transformer-based approaches.
n
- Perform data preparation, feature engineering, statistical analysis, model training, validation, and performance optimization.
n
- Develop hybrid solutions combining deterministic rules, fuzzy matching, statistical methods, and machine learning.
n
- Write clean, efficient, maintainable Python code and develop reusable data science and ML components.
n
- Understand and work extensively with an existing Python codebase, identify issues, optimize algorithms, and implement new functionality.
n
- Develop and deploy ML models as production services/APIs and work with engineering teams to integrate them into enterprise applications.
n
- Analyze model performance and address issues such as false positives, false negatives, scalability, latency, and model explainability.
n
- Research and evaluate new AI/NLP techniques relevant to identity verification, fraud detection, and identity intelligence.
n
n
Required Skills
n
n
- Strong hands-on Python programming and software development skills.
n
- Natural Language Processing and advanced text matching.
n
- Robust foundation in Data Science, Machine Learning, statistics, and predictive modelling.
n
- Strong knowledge of NLP, text processing, similarity algorithms, and entity matching.
n
- Experience with Pandas, NumPy, Scikit-learn, SciPy, and related Python ML/NLP libraries.
n
- Experience in identity management, digital identity, KYC/e-KYC, or biometric applications.
n
- Ability to independently understand and modify an existing codebase and develop solutions with limited reliance on LLM coding assistants.
n
- Experience in fraud detection and anomaly detection models with respect to identity fraud is an added advantage.
n
- Experience with customer/entity resolution and duplicate identity detection.
n
- Experience with multilingual NLP, transliteration, and matching of names across different languages/scripts.
n
- Experience in ML model deployment, REST APIs, Docker, AWS, or MLOps.
n
n
Education
n
Bachelor's or Master's degree in Computer Science, Data Science, Artificial Intelligence, Statistics, Mathematics, Engineering, or a related discipline.
📌 Data Scientist (Hyderabad)
🏢 hiringhood
📍 Hyderabad