Should have strong knowledge in GCP BigQuery and SQL.
Should be proficient with Python or R for data analysis.
Should have Strong knowledge and proven experience of data mining algorithms and statistical analysis.
Should have Ability to interpret complex data and present insights clearly.
Responsibilities
- Design and implement scalable data pipelines using Cloud Composer and Cloud Dataproc to process high volume payer data sets for analytics and reporting needs
- Develop and optimize PySpark jobs that transform complex healthcare claims and member data into accurate and reliable curated data assets for downstream consumers
- Build efficient data models in BigQuery and Cloud SQL that support payer use cases such as cost analysis utilization management and quality of care metrics
- Implement robust Python based utilities and frameworks that standardize data ingestion validation and orchestration workflows across multiple payer data sources
- Optimize query performance in BigQuery and Cloud SQL by refining schemas tuning queries and applying partitioning and clustering strategies to reduce latency and cost
- Collaborate with product owners and business analysts from the payer domain to translate requirements into technical designs and deliver data solutions with clear value
- Ensure data quality by creating validation rules reconciliation checks and automated monitoring for payer specific metrics across all stages of the data pipeline
- Apply secure development practices to protect sensitive payer information through configuration of access controls encryption and privacy aware data handling patterns
- Create clear technical documentation that explains data flows transformation logic dependencies and operational runbooks for all implemented solutions
- Work closely with operations and support teams to troubleshoot production issues implement fixes and enhance resilience of cloud based data workflows
- Contribute to continuous improvement by identifying opportunities to simplify pipelines reduce manual steps and leverage new capabilities in the cloud data ecosystem
- Participate in agile ceremonies and provide accurate effort estimates progress updates and risk identification to enable predictable delivery in the hybrid work model
- Mentor peers through code reviews and knowledge sharing sessions focused on best practices in PySpark development BigQuery optimization and payer domain patterns
Qualifications
- Demonstrate robust proficiency in Python programming with experience building reusable modules error handling patterns and integration with cloud native services
- Exhibit deep hands on experience in PySpark with the ability to design performant transformations manage large datasets and debug complex distributed processing issues
- Show proven capability in BigQuery including data modeling query optimization cost management and development of analytical datasets for business intelligence
- Possess practical expertise in Cloud SQL and Cloud Dataproc configuration job orchestration and environment tuning to support reliable daily production workloads
- Display applied experience with Cloud Composer for workflow orchestration including dependency management scheduling and operational monitoring of data pipelines
- Bring solid understanding of payer domain concepts such as claims processing eligibility benefits provider networks and associated data structures
- Hold experience working in hybrid delivery models and collaborating effectively with cross functional teams using agile methods for planning and execution
📌 GCP,Bigquery + Python (Chennai)
🏢 Cognizant
📍 Chennai
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.