04 Sep
|
At Dawn Technologies
|
Gurugram
04 Sep
At Dawn Technologies
Gurugram
About The Company
The tech space is crowded, but most solutions feel like they’re cut from the same cloth — missing the mark on what businesses truly need. Organizations are chasing innovation, but too often, it comes at the expense of flexibility, independence, and real technological depth. At Dawn Technologies is a niche company laser-focused on delivering specialized tech solutions in Gen AI, Data Engineering, and Backend Systems. What sets us apart is our commitment to the technologists — the creators, problem-solvers, and innovators who bring true value to transformation. We believe that teamwork plus talent equals exceptional results. We are looking for teammates who are great problem solvers and enjoy collaborating within a team to achieve big outcomes.
Data Scientist – AI/ML
Experience: 5–8 years
Employment Type: Contractual (6 Months, extendable based on performance/business need)
Location: Hybrid from India or Vietnam
About the Role
We're looking for an experienced Data Scientist to support an AI-powered classification platform built around a single LLM agent (using disambiguation, lookup, and determination functions) operating inside a large enterprise client's Azure-based AI infrastructure. The immediate priority is instrumenting real usage to replace planning assumptions with measured data, and rigorously evaluating whether the LLM agent meets the accuracy bar required for production use in a regulated, audit-sensitive domain. A conditional later phase may introduce a classical ML hybrid engine if the LLM-only approach doesn't clear that bar — this candidate should be equipped for that work if it materializes, but it is not the day-one focus.
This is a hands-on role for a practitioner comfortable working close to a live evaluation and instrumentation effort, not a greenfield model-building engagement.
Key Responsibilities
Near-term priority — Discovery & Capability Validation
- Design and implement instrumentation to convert planning assumptions — token consumption, cache/reuse hit rate, model tier selection,
retrieval-context sizing — into measured figures from representative traffic
- Analyze usage data to validate or revise pre-launch cost and volume estimates
- Design and build the evaluation framework and ground-truth test set used to determine whether the LLM agent can classify reliably and independently against a defined rules framework
- Define and calibrate accuracy thresholds and confidence scoring that determine whether a secondary classical modeling phase is needed
- Analyze classification errors and failure modes to distinguish disambiguation gaps, retrieval gaps, and reasoning gaps
- Evaluate explainability/audit quality of agent-generated decision records against compliance requirements — outputs need to satisfy audit trail standards, not just accuracy metrics
Conditional — Hybrid Classical Modeling Phase (if triggered by evaluation results)
- Design, train, and validate classical ML classifiers for structured, hierarchical classification tasks
- Establish retraining cadence, drift detection, and confidence-band calibration
- Produce model explainability outputs (e.g., SHAP/LIME) suitable for regulatory audit review
Ongoing
- Build and maintain MLOps pipelines for evaluation, monitoring, and (if the classical modeling phase triggers) model deployment and retraining, within Azure environment or AWS
- Collaborate with engineering, platform, and compliance stakeholders to translate compliance and accuracy requirements into evaluation and modeling design
- Document methodology, evaluation results, and modeling decisions clearly for both technical and business/compliance audiences
Required Skills & Qualifications
- 5–8 years of skilled experience in Data Science / AI / ML roles
- LLM evaluation methodology — building test sets, defining accuracy/quality metrics, and rigorously assessing LLM-based system performance (this is core to the role, not a bonus skill)
- Practical experience with RAG systems and prompt engineering, including retrieval quality assessment
- Strong proficiency in Python and relevant ML/data libraries (scikit-learn, PyTorch, or similar)
- Solid understanding of statistics, probability, and ML algorithms (supervised/unsupervised, deep learning) — applicable if the classical modeling phase triggers
- Experience with SQL and working with structured/semi-structured datasets at moderate scale
- Hands-on MLOps pipeline experience — building and operating pipelines for evaluation, monitoring, and deployment (MLflow, Airflow, Kubeflow, or cloud-native equivalents)
- Working experience with Azure ML tooling (required — target deployment environment); AWS ML tooling experience also valuable
- Familiarity with containerized deployment (Docker; Kubernetes a plus)
- Experience with data visualization tools (Tableau, Power BI, or Matplotlib/Seaborn) for communicating evaluation results
- Strong analytical rigor and comfort working with ambiguous, evolving requirements
- Excellent communication skills — ability to present evaluation findings to both engineering and compliance/business stakeholders
- Bachelor's/Master's degree in Computer Science, Data Science, Statistics, Mathematics, or related field
Good to Have
- Experience with NLP or classical text classification (relevant if the classical modeling phase triggers)
- Experience with regulated/compliance-sensitive domains (audit trails, explainability requirements — e.g., pharma, finance, legal, trade/logistics)
- Prior experience in contract/consulting roles with fast ramp-up
- Familiarity with Spark for batch data processing (not required at this project's scale)
Contract Details
- Duration: 6 months (extendable based on project needs and performance)
- Engagement Type: Contractual
- Start Date: Immediate / 30 days
📌 Data Scientist – AI/ML (Gurugram)
🏢 At Dawn Technologies
📍 Gurugram