01 Aug
|
Stryker
|
Gurugram
**What you will do:** + **Build and own end-to-end data pipelines on Azure Databricks:** from ingestion through Bronze/Silver/Gold medallion transformation to curated Gold datasets, including incremental and historical loads from diverse sources. + **Own the delivery pipeline:** repository structure, branching strategy, and YAML-based CI/CD in Azure DevOps - and manage promotion of code, data and reports across Dev, QA, UAT and Production. + **Design and maintain the Power BI consumption layer:** semantic models, datasets and reports over lakehouse data - including workspace promotion and capacity management. + **Design and implement data quality, validation and reconciliation frameworks:** proving that automated output matches the manual baseline it replaces and handling schema drift without silent failure. + **Maintain lineage, audit trails and validation evidence:** to the standard required by FDA 21 CFR Part 11 and EU MDR, and contribute to test strategy and UAT execution alongside business SMEs. + **Develop and deploy machine learning and Generative AI solutions:** to develop analytics-ready insight for various use cases and business objectives. + **Monitor, tune and cost-optimise production workloads:** cluster sizing and autoscaling, Spark and Delta performance tuning, alerting, and incident triage when pipelines fail. + **Partner with the central platform team** to specify infrastructure requirements precisely, diagnose provisioning gaps, and unblock dependencies before they reach the critical path. **Translate business and regulatory requirements into scoped solutions** with SMEs and stakeholders, quantify the financial impact,
and present findings to technical and non-technical audiences. **Contribute to the platform's evolution toward Microsoft Fabric / OneLake** , assessing what transfers cleanly and what must be rebuilt. **What you need:** Must have skills + **7-10 years of experience. Hands-on Azure Databricks depth:** Delta Lake, medallion architecture, PySpark, cluster and job management, Unity Catalog, and performance optimisation (partitioning, caching, broadcast joins, AQE, OPTIMIZE / Z-ORDER). + **The wider Azure data stack:** ADLS Gen2, Azure Data Factory, Synapse, with strong SQL and ETL/ELT design including incremental ingestion patterns. SQL Server and Key Vault exposure expected. + **Unstructured and siloed data sourcing:** ingestion and remediation of unstructured content from distributed sources (scanned or flattened documents, Excel-to-PDF snapshots that resist standard OCR, SharePoint repositories, attachments) using OCR, document intelligence and AI techniques. **CI/CD and environment ownership:** Git, Azure DevOps, YAML pipelines, secrets and service principal management, and multi-environment promotion across Dev/QA/UAT/Prod. + **Power BI to a build-and-own standard:** semantic models, Power Query/M, DAX, and deployment pipelines, including where a direct lake-to-BI connection is unavailable.
+ **Expert Python and applied machine learning:** at least two of time series forecasting, classification, computer vision or predictive modelling, plus practical LLM/GenAI experience. + **Production data quality and governance discipline:** validation frameworks, data lineage, audit logging, and role-based security. **Demonstrable hands-on delivery ownership:** within a centrally governed enterprise platform, with the financial acumen to size and defend the business impact of what you build. Preferred skills + **Regulated-environment exposure:** medical devices, pharmaceuticals, healthcare or manufacturing, with familiarity with validation, audit or compliance lifecycles (FDA 21 CFR Part 11, EU MDR, GxP). + **Microsoft Fabric / OneLake experience:** assessed as transition capability, not a gate. **Applied NLP beyond document processing:** embeddings, semantic similarity, fuzzy matching and clustering for text standardisation. Stryker is a global leader in medical technologies and, together with its customers, is driven to make healthcare better. The company offers innovative products and services in MedSurg, Neurotechnology, Orthopaedics and Spine that help improve patient and healthcare outcomes. Alongside its customers around the world, Stryker impacts more than 150 million patients annually. Stryker Corporation is an equal opportunity employer. Qualified applicants will receive consideration for employment without regard to race, ethnicity, color, religion, sex, gender identity, sexual orientation, national origin, disability, or protected veteran status. Stryker is an EO employer - M/F/Veteran/Disability.
📌 Staff Data Scientist (Gurugram)
🏢 Stryker
📍 Gurugram