23 Sep
|
Thinkgrid Labs
|
India
23 Sep
Thinkgrid Labs
India
Thinkgrid Labs designs and builds custom web, mobile, cloud, data, and AI solutions. Our team brings together software engineers, architects, data engineers, and UI/UX designers to deliver reliable systems for clients around the world.
We’re expanding our data practice and building a modern data platform for a US health insurer, with Microsoft Fabric as the primary platform. You’ll design, build, and operate reliable data pipelines—from source ingestion and transformation to analytics-ready lakehouse and warehouse models. Hands-on Microsoft Fabric experience is strongly preferred, and we welcome engineers with strong Databricks or comparable cloud data platform experience who are ready to apply their skills in Fabric.
Job Title: Data Engineer – Microsoft Fabric & Cloud Data Platforms
Location: Remote
Working Hours: 2 PM IST to 11 PM IST
Experience Required: 6–10 years in data engineering, including 2+ years working with cloud data platforms
Education: Bachelor’s or Master’s degree in Computer Science, Health Informatics, or Business
Preferred Certifications: Relevant Microsoft Fabric/Azure, Databricks, or equivalent data engineering certifications
Who are you?
- Strong in SQL and Data Modeling: You work confidently with complex relational databases and design practical data models for integration, reporting, and analytics.
- Experienced with Cloud Data Platforms: You’ve built production solutions on Microsoft Fabric, Databricks, or comparable platforms. Experience with Fabric Mirroring, OneLake, Lakehouse/Warehouse, and Fabric Data Factory is particularly valuable.
- A Reliable Pipeline Builder: You design ingestion and transformation pipelines that handle incremental loads, CDC,
schema changes, and recovery from failures.
- Comfortable with Python and Spark: You use code and notebooks to process data, implement reusable workflows, and validate results.
- Focused on Data Quality: You build checks, reconciliation, and clear data contracts into your work so downstream teams can trust the data.
- Operationally Minded: You monitor pipeline health, troubleshoot production issues, document runbooks, and balance performance with cost.
- Security Conscious: You handle sensitive data responsibly, including PII/PHI, and apply appropriate access controls and governance practices.
What will you be doing?
- Build Data Pipelines: Develop and maintain ingestion and ETL/ELT pipelines across relational databases, files, APIs, and other sources, primarily using Microsoft Fabric.
- Develop Lakehouse and Warehouse Solutions: Organize data across raw, refined, and curated layers, using OneLake and Fabric Lakehouse/Warehouse to support reliable downstream consumption.
- Transform and Model Data: Build reusable transformations and analytics-ready models that translate source data into useful, well-documented datasets.
- Manage Incremental Data and Change: Implement CDC, replication, or watermark-based loading as appropriate. Handle deletes, late-arriving data, backfills, and schema evolution safely,
using Fabric Mirroring where suitable.
- Orchestrate Resilient Workflows: Manage dependencies, scheduling, retries, idempotency, and failure recovery using Fabric Data Factory, notebooks, and appropriate orchestration tools.
- Ensure Quality and Observability: Implement data validation, reconciliation, logging, metrics, lineage, and alerts to maintain data accuracy and pipeline reliability.
- Optimize Performance and Cost: Tune queries, Spark workloads, partitioning, file sizes, and compute usage to meet delivery expectations efficiently.
- Collaborate and Document: Work with platform architects, security teams, analysts, and business stakeholders on requirements, governance, and delivery. Maintain pipeline documentation, data definitions, SLAs, and operational runbooks.
Must-have skills
- Solid SQL and experience working with large, complex relational databases
- Production experience with Microsoft Fabric, Databricks, or a comparable cloud data platform
- Python and Apache Spark for data processing, transformation, and validation
- Experience building and operating ingestion and ETL/ELT pipelines
- Data modeling and lakehouse or data warehouse fundamentals
- CDC/incremental loading, schema evolution, and data quality practices
- Pipeline orchestration, monitoring, troubleshooting, and performance tuning
- Git-based development workflows and familiarity with CI/CD for data solutions
Benefits
- 5-day work week
- Health Insurance
- 100% remote setup with flexible work culture and international exposure
- Opportunity to work on mission-critical healthcare projects impacting providers and patients globally
📌 Data Engineer – Microsoft Fabric & Cloud Data Platforms (India)
🏢 Thinkgrid Labs
📍 India