Key Responsibilities
Develop, maintain, and optimize SQL-based ETL processes, stored procedures, and data transformations across DB2 and SQL Server environments
Support and monitor ~225 scheduled jobs and ~40 job streams using IBM Workload Scheduler, ensuring timely execution and prompt failure resolution
Build and maintain data pipelines using IBM DataStage and UNIX scripting for enterprise data integration workflows
Support Oracle GoldenGate for real-time data replication and change data capture (CDC) across source and target systems
Develop and maintain data integration workflows from source systems to analytics platforms, including validation and reconciliation logic
Provide operational support for DB2 and SQL Server environments encompassing ~160 schemas, ~20TB active storage, ~4,000 tables, and ~3,300 views
Monitor pipeline health proactively, detect anomalies, and resolve data quality and availability issues within defined SLAs
Support Dev/QA/Prod environment management including release coordination and production readiness validation
Assist with AWS stabilization activities for analytics data layers post migration from on-premises infrastructure
Track and manage all work through ServiceNow, ensuring accurate classification, status updates, and SLA compliance
Collaborate with Tableau and BusinessObjects developers to ensure data availability and pipeline reliability for reporting
Participate in L1/L2 triage for pipeline incidents, data quality failures, and integration issues
Contribute to runbook documentation and standard operating procedures for supported pipelines and jobs
Required Qualifications
Minimum Degree Required: Bachelor’s Degree in Engineering, Statistics, Mathematics, Computer Science, Data Science, Economics, or a related quantitative field
2-5 years of experience in data engineering, ETL development, or data integration roles
Robust SQL proficiency — query optimization, stored procedures, and database objects in DB2 and/or SQL Server
Hands-on experience with at least one enterprise ETL or orchestration tool - IBM DataStage, Netezza
Experience with job scheduling tools (IBM Workload Scheduler, Control-M, or equivalent)
Familiarity with UNIX/Linux scripting for pipeline automation and file-based integrations
Understanding data warehouse concepts — schemas, star/snowflake models, dimensions, and facts
Ability to troubleshoot data pipeline failures end-to-end and communicate resolution steps clearly
Experience working in regulated, compliance-aware environments (healthcare preferred)
Preferred Qualifications
Healthcare data experience — claims, clinical, EMR, pharmacy, or population health datasets
AWS experience — S3, Glue, RDS, Redshift, or Lambda in a data engineering context
Experience with Oracle GoldenGate or similar CDC/replication tools
Familiarity with Tableau or BusinessObjects as downstream consumers of engineered data
Knowledge of HIPAA data handling requirements and PHI access controls
Exposure to ServiceNow for ITSM-based delivery tracking.
📌 Data Engineer (Bengaluru)
🏢 PwC
📍 Bengaluru