07 Oct
|
HCA Healthcare - India
|
Hyderabad
07 Oct
HCA Healthcare - India
Hyderabad
General Position Information
Reports Directly To (Title)
Director
Matrix Reports To (Title)
As applicable based on functional alignment
Direct Reports
Individual contributor; no direct reports
Position Summary
Data Engineer to support the HIM Datamart used by Parallon HIM Operations. This critical role is responsible for creating, maintaining, monitoring, and supporting data pipelines that pull data from Teradata, multiple SQL Server environments, various GCP Datamarts, and other enterprise data sources into the HIM GCP Datamart.
The successful candidate will design and support ETL and ELT processes using Airflow, DAGs, Google Cloud Dataproc, Python, PowerShell, SSIS, SQL, and other data integration technologies. The ideal candidate must be comfortable supporting production data pipelines, troubleshooting complex data issues, improving data quality, and ensuring timely and accurate data availability for HIM operations.
This position is best suited for a data engineering professional who can balance development, production support, automation, data quality, cloud data engineering, documentation, and operational reliability in a critical healthcare operations environment.
Responsibilities
- Design, develop, maintain, and support ETL/ELT data pipelines for the HIM GCP Datamart. Extract, transform, and load data from Teradata, SQL Server, GCP Datamarts, and other enterprise data sources into HIM data platforms.
- Develop and support data workflows using Airflow, DAGs, Dataproc, Python, PowerShell, SSIS, SQL, and related data integration tools.
- Create, schedule, monitor, and troubleshoot Airflow DAGs for recurring data ingestion, transformation, validation, reconciliation, and delivery processes.
- Develop and optimize SQL queries, stored procedures, views, functions, and transformation logic across SQL Server, Teradata, PostgreSQL, BigQuery, and other relational platforms.
- Support on-premise to GCP data integration using Google Cloud Dataproc and related GCP services for large-scale data processing and transformation.
- Monitor data pipelines for failures, delays, performance issues, data quality problems, schema changes, and source system impacts.
- Troubleshoot production data issues, including missing data, late-arriving files, failed jobs, transformation errors, data mismatches, and performance bottlenecks.
- Perform root cause analysis and implement long-term fixes to improve pipeline reliability, reduce recurring failures, and improve operational stability.
- Implement data validation, reconciliation, and quality checks to ensure HIM data is complete, accurate, consistent, and timely.
- Collaborate with Parallon HIM Operations, source system owners, and technical teams to understand data needs, dependencies, refresh schedules, reporting impacts, and business priorities.
- Create and maintain data pipeline documentation, data flow diagrams, source-to-target mappings, job schedules, support procedures, and operational runbooks.
- Participate in release management, change control, defect resolution, enhancement delivery, incident response, and production support activities.
- Improve automation, monitoring, alerting, logging, and operational visibility across HIM data pipelines.
- Ensure data engineering solutions follow enterprise standards for security, privacy, compliance, data governance, change control, and production readiness.
- Identify opportunities to modernize legacy ETL processes and improve scalability, maintainability, reliability, and cost efficiency.
Education &
- Experience
- Bachelor’s degree in Computer Science, Information Systems, Data Engineering, Engineering, Healthcare Informatics, or a related field is required.
- Five or more years of experience in data engineering, ETL/ELT development, data warehousing, data integration, or supporting enterprise data platforms, data marts, data lakes, or operational data stores.
- Solid hands-on experience designing, developing, maintaining, and troubleshooting production data pipelines using SQL, Python, PowerShell, SSIS, Apache Airflow, DAGs, and other ETL/ELT technologies.
- Strong SQL experience across SQL Server, Teradata, PostgreSQL, BigQuery, or similar platforms, including complex queries, joins, aggregations, stored procedures, views, functions, performance tuning, and source-to-target validation.
- Experience extracting, integrating, and migrating data from Teradata, SQL Server, file-based sources, GCP Datamarts, and other enterprise data sources into GCP-based data platforms.
- Experience with Google Cloud Platform data services, including Dataproc, Cloud Composer, Cloud Storage, BigQuery, Cloud SQL, Cloud Logging, and Cloud Monitoring.
- Experience with batch processing, job scheduling, dependency management, error handling, logging, alerting, data validation, reconciliation, data quality checks, and production support for critical business operations.
- Experience with file-based integrations, including CSV, fixed-width files, Excel, JSON, XML, Parquet, Avro, and delimited files.
- Experience with CI/CD and source control practices using GitHub, Azure DevOps, Git, YAML pipelines, or similar tools.
- Strong analytical, troubleshooting, problem-solving, written communication, and verbal communication skills, with the ability to work independently and collaboratively with technical and business teams.
Must Have Skills
- Experience with Spark, PySpark, Hadoop, or distributed data processing frameworks.
- Experience with Teradata SQL, BTEQ, FastExport, TPT, or other Teradata data extraction patterns.
- Experience designing scalable data ingestion frameworks and reusable pipeline components.
- Experience with incremental loads, full refreshes, CDC patterns, SCD handling, partitioning, and data archival strategies.
- Experience with secure file transfers, file watchers, landing zones, staging tables, and data lake patterns.
- Experience with data modeling, dimensional modeling, star schemas, snowflake schemas, and reporting-friendly data structures.
- Experience with BI/reporting platforms and downstream analytics dependencies.
- Experience with monitoring, alerting, operational dashboards, and data pipeline observability.
- Experience with healthcare data, HIM operations, revenue cycle, coding, clinical documentation, billing, or Parallon operational data is strongly preferred.
- Experience with data privacy, HIPAA, PHI handling, role-based access, encryption, audit logging, and healthcare compliance requirements.
- Experience working in Agile/Scrum or Lean delivery environments.
Cloud/Data Platforms
- Google Cloud Platform, HIM GCP Datamart, GCP Datamarts, BigQuery, Cloud SQL, Cloud Storage, Data Lakes, Data Marts, Data Warehousing
Orchestration &
- Processing
- Apache Airflow, DAGs, Cloud Composer, Google Cloud Dataproc, PySpark, Batch Processing, Job Scheduling
ETL/ELT &
- Integration
- ETL, ELT, SSIS, File-Based Integrations, Source-to-Target Mapping, Data Validation, Data Reconciliation, Data Quality Checks
Databases &
- SQL
- SQL Server 2019, Teradata, Teradata SQL, PostgreSQL, Stored Procedures, Views, Functions, Indexes
Scripting &
- Automation
- Python, PowerShell, YAML
File/Data Formats
- JSON, XML, CSV, Parquet
DevOps &
- Support
- GitHub, Azure DevOps, CI/CD Pipelines, Cloud Logging, Cloud Monitoring, Monitoring and Alerting, Production Support
Nice To Have Skills
- Five or more years of relevant data engineering, ETL development, or data warehousing experience.
- Hands-on experience with SQL, Python, Airflow, DAGs, SQL Server, and ETL development is required.
- Experience with Teradata and SSIS is strongly preferred.
- Experience with Google Cloud Platform, Dataproc, Cloud Composer, Cloud Storage, BigQuery, or related GCP data services is preferred.
- Experience supporting production data pipelines and critical business operations is required.
- Experience with data validation, reconciliation, monitoring, alerting, and production troubleshooting is required.
- Healthcare IT, HIM, revenue cycle, or Parallon operational data experience is preferred.
Licenses, Certifications &
- Training
- N/A
Knowledge, Skills, Abilities, Behaviors
- Strong ownership mindset with the ability to support critical operational data processes.
- Ability to understand business workflows and translate them into reliable data engineering solutions.
- Ability to communicate data issues, risks, dependencies, and resolution plans clearly to technical and non-technical stakeholders.
- Strong attention to detail, especially when working with operational, financial, healthcare, or compliance-sensitive data.
- Ability to troubleshoot complex data issues across source systems, ETL processes, cloud platforms, databases, and downstream consumers.
- Ability to prioritize production support issues based on business impact and operational urgency.
- Ability to create clear documentation, data flow diagrams, source-to-target mappings, and support runbooks.
- Ability to work effectively with business teams, source system owners, cloud teams, database teams, reporting teams, and application support teams.
- Self-motivated learner who stays current with data engineering, cloud data platforms, orchestration tools, and automation practices.
- Ability to work with minimal supervision while managing multiple priorities.
- Collaborative team player with strong analytical and problem-solving skills.
📌 Senior - Data Engineer (Hyderabad)
🏢 HCA Healthcare - India
📍 Hyderabad