20 Sep
|
Velocitai Digital
|
Pune
20 Sep
Velocitai Digital
Pune
Key Responsibilities / Essential Duties
- Design and build scalable Extract, Transform, Load (ETL) pipelines to process large-scale structured
and unstructured data, ensuring alignment with enterprise data architecture standards and
pharmaceutical distribution data models.
- Develop and manage data architectures, including data lakes, data warehouses, and large-scale
processing systems that support advanced analytics capabilities across the organization.
- Optimize data retrieval processes, troubleshoot pipeline bottlenecks, and fine-tune database
performance to maximize efficiency and reduce processing costs.
- Interpret data sets with particular attention to trends and patterns valuable for diagnostic and
predictive analytics efforts, creating visual representations of quantitative and qualitative data to
support business decision-making.
- Implement data quality and governance processes, including validation checks, error handling, and
quarantine procedures to ensure data accuracy, consistency, and compliance with applicable
regulatory requirements.
Page 2 of 4
- Support the delivery of Business Intelligence solutions to internal and external customers, utilizing
business intelligence tools across departments and functional areas to provide detailed reporting
and analysis.
- Assess analytics and reporting needs, provide recommendations for enhanced solutions and
specifications for business intelligence tools, and prioritize reporting requirements across process
areas to align with deployment strategies.
- Monitor pipeline health, system integrity, and automated alerts proactively, diagnosing and
resolving incidents involving Extract, Transform, Load scripts, database connectivity, or
infrastructure issues to minimize impact on downstream consumers.
- Collaborate with data scientists, analysts, software engineers, and operations teams on system
upgrades, migrations, and integration projects, onboarding recent data sources from various
applications and interfaces into central repositories.
- Maintain data management standards and processes to effectively govern data through the
reporting life cycle, proactively identifying opportunities for continuous improvement and
operational excellence.
- Validate data for accuracy, process audit reports, propose solutions to remedy deficiencies, and
confirm results in accordance with applicable data compliance requirements.
- Communicate data engineering requirements, architecture decisions, and technical findings to
cross-functional partners across multiple time zones,
providing expertise and guidance on data
infrastructure and pipeline design.
- Document processes, standard operating procedures, source-to-target mappings, and knowledge
artifacts to ensure operational continuity during transitions, enabling seamless integration of
evolving team structures.
Qualifications
Education
- Bachelors degree in Statistics, Computer Science, Information Technology, or equivalent related
experience (required).
- Masters degree in Computer Science, Data Engineering, Information Systems, or a related technical
discipline (preferred).
Experience
- Minimum professional experience as per the Level in data engineering, with a focus on designing,
building, and maintaining scalable data pipelines and data architectures in enterprise environments.
- Minimum of hands-on experience as per the Level developing and deploying Extract, Transform,
Load (ETL) pipelines, data lakes, data warehouses, and large-scale data processing systems in
production environments.
- Demonstrated experience with end-to-end data pipeline development — including data ingestion,
transformation, quality validation, and delivery — aligned with enterprise data architecture
standards and pharmaceutical distribution data models.
Page 3 of 4
- Experience collaborating with data scientists, analysts, software engineers, and operations teams on
system upgrades, migrations, and integration projects, onboarding new data sources from various
applications and interfaces into central repositories.
- Demonstrated experience maintaining data management standards and governance processes
through the reporting life cycle, including data quality validation, audit compliance, and error
handling procedures.
- Experience in a regulated industry (pharmaceutical, healthcare, or life sciences) is a plus.
Certifications (Required / Preferred)
- Microsoft Certified: Azure Data Engineer Associate (DP-203) — Required.
- Databricks Certified Data Engineer Associate — Required.
- Microsoft Certified: Power BI Data Analyst Associate — Preferred.
- Snowflake SnowPro Core Certification — Preferred.
Knowledge, Skills & Abilities
- Expert-level proficiency in designing and building scalable Extract, Transform, Load (ETL)
pipelines to
process large-scale structured and unstructured data primarily using Databricks platform, ensuring
alignment with enterprise data architecture standards and pharmaceutical distribution data models.
- Strong experience with data architecture design and management, including data lakes, data
warehouses, and large-scale processing systems that support advanced analytics capabilities across
the organization.
- Proficiency in Structured Query Language (SQL) for complex query development, database
performance tuning, and data manipulation across relational database management systems
(RDBMS), with the ability to optimize data retrieval processes and troubleshoot pipeline
bottlenecks.
- Hands-on experience with data integration and ETL tools such as Informatica and Alteryx for
designing end-to-end data ingestion, transformation, and orchestration workflows across onpremises and cloud environments.
- Solid understanding of data quality and governance processes, including validation checks, error
handling, quarantine procedures, and data management standards to ensure data accuracy,
consistency, and compliance with applicable regulatory requirements.
- Proficiency in business intelligence reporting tools such as Microsoft Power BI, Tableau, and Qlik
Sense for supporting the delivery of Business Intelligence solutions and creating visual
representations of quantitative and qualitative data to drive business decision-making.
- Experience with pipeline health monitoring, automated alerting, and incident resolution — including
diagnosing and resolving issues involving ETL scripts, database connectivity, and infrastructure
components — to minimize impact on downstream consumers.
- Proficiency in data analysis and synthesis techniques with particular attention to trends and patterns
valuable for diagnostic and predictive analytics efforts, including supply chain performance
indicators across pharmaceutical distribution operations.
Page 4 of 4
- Strong analytical and problem-solving abilities with a detail-oriented mindset and commitment to
data accuracy, audit compliance, source-to-target mapping integrity, and continuous improvement
throughout the reporting life cycle.
- Excellent communication skills with the ability to convey complex data engineering concepts,
architecture decisions, and technical findings to diverse technical and non-technical stakeholders
across multiple time zones, and to collaborate effectively with data scientists, anal
📌 Azure Data Engineer (Pune)
🏢 Velocitai Digital
📍 Pune