Key ResponsibilitiesDevelop, maintain, and optimize SQL-based ETL processes, stored procedures, and data transformations across DB2 and SQL Server environmentsSupport and monitor ~225 scheduled jobs and ~40 job streams using IBM Workload Scheduler, ensuring timely execution and prompt failure resolutionBuild and maintain data pipelines using IBM DataStage and UNIX scripting for enterprise data integration workflowsSupport Oracle GoldenGate for real-time data replication and change data capture (CDC) across source and target systemsDevelop and maintain data integration workflows from source systems to analytics platforms, including validation and reconciliation logicProvide operational support for DB2 and SQL Server environments encompassing ~160 schemas, ~20TB active storage, ~4,000 tables, and ~3,300 viewsMonitor pipeline health proactively, detect anomalies, and resolve data quality and availability issues within defined SLAsSupport Dev/QA/Prod workplace management including release coordination and production readiness validationAssist with AWS stabilization activities for analytics data layers post migration from on-premises infrastructureTrack and manage all work through ServiceNow, ensuring accurate classification, status updates, and SLA complianceCollaborate with Tableau and BusinessObjects developers to ensure data availability and pipeline reliability for reportingParticipate in L1/L2 triage for pipeline incidents, data quality failures,
and integration issuesContribute to runbook documentation and standard operating procedures for supported pipelines and jobsRequired QualificationsMinimum Degree Required:
Bachelor’s Degree in Engineering, Statistics, Mathematics, Computer Science, Data Science, Economics, or a related quantitative field2-5 years of experience in data engineering, ETL development, or data integration rolesStrong SQL proficiency — query optimization, stored procedures, and database objects in DB2 and/or SQL ServerHands-on experience with at least one enterprise ETL or orchestration tool - IBM DataStage, NetezzaExperience with job scheduling tools (IBM Workload Scheduler, Control-M, or equivalent)Familiarity with UNIX/Linux scripting for pipeline automation and file-based integrationsUnderstanding data warehouse concepts — schemas, star/snowflake models, dimensions, and factsAbility to troubleshoot data pipeline failures end-to-end and communicate resolution steps clearlyExperience working in regulated, compliance-aware environments (healthcare preferred)Preferred QualificationsHealthcare data experience — claims, clinical, EMR, pharmacy, or population health datasetsAWS experience — S3, Glue, RDS, Redshift, or Lambda in a data engineering contextExperience with Oracle GoldenGate or similar CDC/replication toolsFamiliarity with Tableau or BusinessObjects as downstream consumers of engineered dataKnowledge of HIPAA data handling requirements and PHI access controlsExposure to ServiceNow for ITSM-based delivery tracking.
📌 Data Engineer (Mumbai)
🏢 PwC
📍 Mumbai