09 Sep
|
Takeda
|
Bengaluru
By clicking the “Apply” button, I understand that my employment application process with Takeda will commence and that the information I provide in my application will be processed in line with Takeda’s Privacy Notice and Terms of Use. I further attest that all information I submit in my employment application is true to the best of my knowledge.
: Principal Data Engineer: About the Role: We are seeking a highly experienced Principal Data Engineer: to lead the design, architecture, and implementation of enterprise-scale data platforms that power Takeda's R&D;, Clinical, Regulatory, Commercial, and Enterprise Analytics ecosystems.
As a technical leader, you will drive the strategic direction of data engineering, define architecture standards, and deliver scalable, secure, and compliant data solutions using Databricks, AWS, PySpark, Delta Lake, and up-to-date DataOps practices: . You will partner with business stakeholders, solution architects, product owners, data scientists, and engineering teams to build reliable data products that accelerate innovation and data-driven decision making across Takeda.
This role requires deep expertise in modern cloud data platforms, distributed data processing, software engineering practices, and cross-functional leadership. You will mentor engineering teams, establish best practices, and ensure enterprise data solutions meet quality, scalability, security, and regulatory requirements.
Key Responsibilities: Data Architecture & Platform Leadership:
- Lead architecture, design, and implementation of enterprise data platforms on Databricks and AWS.
- Define and govern enterprise data engineering standards, reference architectures, and reusable frameworks.
- Design scalable Lakehouse architectures using Delta Lake, Unity Catalog, Databricks Workflows, and AWS cloud services.
- Collaborate with enterprise architects and business stakeholders to translate business requirements into scalable technical solutions.
- Drive platform modernization initiatives, cloud migration programs, and data transformation roadmaps.
- Establish best practices for data modeling, metadata management, data lineage, and governance.
- Evaluate emerging technologies and recommend improvements to Takeda's data ecosystem.
Data Engineering & Solution Delivery:
- Design and build high-performance batch and streaming (Optional) data pipelines using PySpark, Spark SQL, and Databricks.
- Architect ingestion frameworks supporting structured, semi-structured, and unstructured data from internal and external systems.
- Lead implementation of medallion architecture patterns (Bronze, Silver, Gold) to support trusted enterprise data products.
- Optimize large-scale data processing workloads to improve performance, reliability, and cost efficiency.
- Design and implement reusable ETL/ELT frameworks and accelerator components.
- Establish data contracts and engineering standards to ensure consistency and reliability across platforms.
- Ensure data solutions are scalable, maintainable, and aligned with enterprise architecture principles.
Cloud Engineering & AWS Platform Management:
- Architect and implement cloud-native data solutions using AWS services including
- S3
- IAM
- Glue
- Lambda
- ECS/EKS
- Step Functions
- EventBridge
- CloudWatch
- Secrets Manager
- KMS
- Redshift
- Design secure multi-account architectures and governance models.
- Establish infrastructure automation practices using Terraform, CloudFormation, or AWS CDK.
- Drive optimization of cloud resources through cost management, workload tuning, and automation.
- Partner with cloud platform teams to ensure operational excellence and security compliance.
DataOps, CI/CD & Engineering Excellence:
- Lead adoption of software engineering best practices across data engineering teams.
- Design and implement CI/CD pipelines for data platforms using GitHub Actions, DevOps, GitLab CI, or Jenkins.
- Establish automated deployment frameworks across Development, Test, Validation, and Production environments.
- Implement unit testing, integration testing, regression testing, and automated quality gates.
- Drive code quality initiatives including
- Peer reviews
- Static code analysis
- Test automation
- Release management
- Version control strategies
- Standardize engineering practices to improve delivery velocity and platform reliability.
- Promote Infrastructure as Code and automated environment provisioning.
Data Quality, Governance & Compliance:
- Establish enterprise data quality frameworks and monitoring capabilities.
- Implement end-to-end data lineage, metadata management, and observability solutions.
- Collaborate with governance, security, quality, and compliance teams to ensure adherence to corporate standards.
- Support implementation of
- Data governance policies
- Role-based access controls
- Data retention policies
- Auditability requirements
- Ensure compliance with
- GxP requirements
- HIPAA
- Promote secure handling of sensitive healthcare and research data.
Performance Optimization & Reliability Engineering:
- Establish platform observability using monitoring, logging, and alerting solutions.
- Define service-level objectives and operational metrics for critical data platforms.
- Lead root-cause analysis and resolution of complex production issues.
- Implement resiliency, disaster recovery, and business continuity strategies.
- Continuously improve platform performance, stability, and operational efficiency.
Leadership, Mentoring & Cross-Functional Collaboration:
- Provide technical leadership and guidance to data engineers across multiple programs and delivery teams.
- Lead architectural reviews, design discussions,
and engineering governance forums.
- Mentor engineers in Databricks, Spark, AWS, DataOps, and software engineering best practices.
- Partner with
- Product Owners
- Business Stakeholders
- Data Scientists
- Cloud Engineering Teams
- Security Teams
- Quality and Compliance Teams
- Influence strategic data platform decisions and long-term technology roadmaps.
- Drive delivery excellence through collaboration, coaching, and continuous improvement.
Required Qualifications:
- Bachelor’s or master’s degree in computer science, Engineering, Information Systems, or related field.
- 12+ years of experience in Data Engineering, Big Data, Data Platforms, or Cloud Engineering.
- 4 to 5 years of experience architecting enterprise-scale data solutions on Databricks / AWS.
- Deep expertise in
- Databricks
- Delta Lake
- Unity Catalog
- Spark
- PySpark
- SQL
- Strong AWS experience across compute, storage, security, and monitoring services.
- Advanced Python development experience and software engineering practices.
- Proven experience building large-scale enterprise data pipelines and Lakehouse architectures.
- Expertise in CI/CD implementation and Git-based development practices.
- Experience implementing automated unit testing and quality assurance frameworks.
- Strong understanding of distributed systems and performance optimization.
- Experience with Infrastructure as Code using Terraform or CloudFormation, or AWS CDK.
- Excellent communication, stakeholder management, and leadership skills.
Preferred Qualifications:
- Life Sciences, Pharmaceutical, Healthcare, or Clinical data domain experience.
- Experience supporting GxP-regulated environments.
- Knowledge of
- Clinical Trial Data
- Regulatory Data
- Pharmacovigilance
- Real World Data (RWD/RWE)
- Omics and Research Data
- Experience with ML/AI data platforms and MLOps foundations.
- Exposure to streaming technologies such as Kafka, Kinesis, or Event Hubs.
- Databricks Certified Data Engineer Professional.
- AWS Solutions Architect Professional.
- AWS Data Analytics Specialty or equivalent certifications.
What Success Looks Like (First 12 Months):
- Data delivery timelines are significantly reduced through reusable frameworks, automation, and CI/CD.
- Data pipelines achieve high reliability, observability, and operational excellence.
- Engineering teams consistently follow software engineering, testing, and deployment best practices.
- Cloud infrastructure and Databricks environments are optimized for performance, scalability, and cost.
- Data products meet quality, governance, and compliance expectations across R&D; and enterprise functions.
- Multiple teams successfully adopt reusable engineering accelerators developed under your leadership.
- Stakeholders recognize the data platform as a strategic enabler for innovation and business outcomes.
Locations: IND - Bengaluru
Worker Type: Employee
Worker Sub-Type: Regular
Time Type: Full time
📌 Principal Data Engineer (Bengaluru)
🏢 Takeda
📍 Bengaluru