Location: Remote, India
Employment: Full-time, with employment administered through an Employer of Record in India
Working hours: 2:00–10:00 PM IST, fixed year-round
Reports to: Data Solutions Architect
Annual fixed gross salary: ₹45,00,000–₹60,00,000, based on relevant experience and demonstrated expertise, plus a performance-based incentive plan and benefits.
About GSMS
Data is central to how we compete—and win.
GSMS is a pharmaceutical company supplying medications to U.S. government healthcare, including the Department of Veterans Affairs and Department of Defense. Our data and analytics capabilities are a core competitive advantage, helping the business identify market opportunities, make pricing and portfolio decisions, and improve sales, manufacturing, and distribution.
Our teams work directly with the people using these insights. The platforms and solutions we build inform decisions across the organization every day.
We are modernizing our data platform on Microsoft Fabric as part of a broader, company-wide AI transformation. This work will support trusted enterprise reporting, self-service analytics, advanced analytics, AI agents, and workflow automation.
The opportunity
We are looking for a Principal Data Architect to take ownership of the detailed architecture, enterprise data models, and technical standards behind our data platform.
You will join during our Microsoft Fabric migration, helping validate and evolve an established architectural direction and translating it into dependable production implementations. Beyond the migration, you will guide how the platform grows, how new data sources and business domains are integrated, and how enterprise data remains consistent, reliable, secure, and useful.
This is a hands-on individual contributor role. You will build prototypes, write SQL and PySpark, review implementations, and help resolve difficult technical problems. You should be comfortable making architectural decisions and demonstrating why they work.
The Data Solutions Architect provides broader solution direction and major architectural approvals; you own detailed data architecture and modeling decisions within that direction. The Lead Data Engineer leads day-to-day engineering delivery.
What you will own
?Platform architecture and technical direction
Validate and evolve our Microsoft Fabric medallion architecture, with bronze and silver data in the lakehouse, transformations primarily in PySpark notebooks, and curated gold data in Fabric Warehouse.
Define practical standards for ingestion, transformation, storage, orchestration, integration, and data consumption.
Document architectural decisions, alternatives, and trade-offs so teams can implement them consistently.
?Enterprise data modeling
Develop and maintain conceptual, logical, and physical data models across business domains. Establish consistent approaches to fact grain, conformed dimensions, business keys, historical changes, and shared business definitions.
Work with business stakeholders and the BI analytics team to ensure the warehouse supports accurate, reusable analytics. Partner with the team responsible for Power BI semantic models and reporting to align upstream data structures with downstream needs.
?Migration and production transition
Guide architectural aspects of the transition from SQL Server and Azure Data Factory to Microsoft Fabric. Help resolve migration design questions, validate reconciliation approaches, and support cutover planning and production readiness while protecting business continuity.
?Hands-on implementation and design review
Build prototypes and reference implementations in SQL, Python, and PySpark to validate important design decisions. Review engineering implementations for correctness, maintainability, performance, and adherence to agreed standards.
Partner with engineers to troubleshoot complex issues involving data models, transformations, processing performance, and cross-system dependencies.
?Long-term platform reliability and evolution
Establish architectural patterns for incremental processing, schema changes, recoverability, monitoring, data quality, and controlled deployments.
Guide performance and capacity optimization across lakehouse and warehouse workloads, balancing business needs, reliability, and cost.
Work with engineering, quality, and security stakeholders to incorporate access controls, lineage, auditability, and governance into platform design.
Evaluate new capabilities against real business needs and strengthen the data foundation for self-service analytics, advanced analytics, and AI-enabled solutions.
What you bring
- Typically 10 or more years of progressive experience in data engineering, enterprise data warehousing, or data architecture, including substantial ownership of architectural decisions for production platforms.
Demonstrated scope and depth matter more than a specific year count.
- Direct, hands-on production experience with Microsoft Fabric, including Lakehouse, PySpark notebooks, and Fabric Warehouse.
- Solid enterprise and dimensional data modeling skills, including fact and dimension design, grain definition, conformed dimensions, historical data handling, and integration across business domains.
- Strong SQL and T-SQL skills, with the ability to develop and review Python and PySpark implementations.
- Experience modernizing or migrating production data platforms, including reconciliation, dependency management, and transition to operational support.
- Practical experience with SQL Server and Azure Data Factory.
- Evidence of making and defending architectural trade-offs involving performance, reliability, maintainability, security, and cost.
- Experience with Git-based development, code reviews, automated deployment practices, and separation of development, testing, and production environments.
- The ability to work directly with business stakeholders, clarify requirements, explain technical decisions, and influence implementation across teams.
- Clear written and spoken English for collaboration with colleagues in India and the United States.
Additional experience we value
- Microsoft Purview or comparable data catalog, lineage, and governance capabilities.
- Designing data foundations that support Power BI semantic models, self-service analytics, and AI-enabled applications.
- Automated data testing, observability, and reusable engineering frameworks.
- Using AI or LLM-assisted development within governed processes for implementation, review, and validation.
- Relevant Microsoft certifications or experience in pharmaceutical, healthcare, distribution, or other complex operational environments.
Working at GSMS
- You will be a dedicated member of the GSMS team, with employment administered through an Employer of Record in India.
- Your day-to-day work, priorities, and team relationships will be with GSMS.
- This is a fully remote position with fixed working hours of 2:00–10:00 PM IST throughout the year. Production support responsibilities take place within agreed working hours.
- The position requires availability throughout those hours and cannot be held concurrently with another full-time job. Other outside work must comply with the applicable employment agreement and conflict-of-interest policy.
- Benefits include comprehensive medical insurance with eligible family coverage options, life insurance, paid leave and holidays, and remote-work support.
📌 Principal Data Architect — Microsoft Fabric & Enterprise Data Warehousing (India)
🏢 GSMS
📍 India