The Developer will design build and optimize PySpark based data solutions that support analytics and content decision making for a global media and entertainment organization. The role involves hybrid work collaborating with cross functional teams to deliver reliable data pipelines improve audience insights and enable data driven strategies that enhance content performance and operational efficiency.
Responsibilities
Design robust PySpark data pipelines that efficiently ingest transform and aggregate large scale media and entertainment datasets to support analytics and reporting needs
Implement optimized PySpark code that enhances performance of batch and near real time data processing for audience metrics and content consumption patterns
Collaborate with data engineers analysts and product teams in a hybrid work model to understand media business requirements and translate them into scalable technical solutions
Develop reusable data frameworks and components that standardize processing of streaming video on demand and advertising data while ensuring consistency and reliability
Apply data quality checks validation routines and monitoring mechanisms within PySpark workflows to maintain accurate and trustworthy data for content and revenue decisions
Integrate data from multiple media platforms and content management systems into unified data models that enable holistic views of audience engagement and catalog performance
Optimize storage formats partitioning strategies and execution configurations in PySpark to reduce processing time and infrastructure costs for large media datasets
Collaborate with stakeholders to deliver explicit and timely data outputs that support programming strategy recommendation engines campaign measurement and operational dashboards
Document technical designs PySpark jobs data flows and configuration details to ensure maintainability and smooth handover across distributed development teams
Ensure compliance with data governance se