Job Summary The Developer will design build and optimize PySpark based data solutions that support analytics and content decision making for a global media and entertainment organization. The role involves hybrid work collaborating with cross functional teams to deliver reliable data pipelines improve audience insights and enable data driven strategies that enhance content performance and operational efficiency.
Responsibilities
- Design robust PySpark data pipelines that efficiently ingest transform and aggregate large scale media and entertainment datasets to support analytics and reporting needs
- Implement optimized PySpark code that enhances performance of batch and near real time data processing for audience metrics and content consumption patterns
- Collaborate with data engineers analysts and product teams in a hybrid work model to understand media business requirements and translate them into scalable technical solutions
- Develop reusable data frameworks and components that standardize processing of streaming video on demand and advertising data while ensuring consistency and reliability
- Apply data quality checks validation routines and monitoring mechanisms within PySpark workflows to maintain accurate and trustworthy data for content and revenue decisions
- Integrate data from multiple media platforms and content management systems into unified data models that enable holistic views of audience engagement and catalog performance
- Optimize storage formats partitioning strategies and execution configurations in PySpark to reduce processing time and infrastructure costs for large media datasets
- Collaborate with stakeholders to deliver clear and timely data outputs that support programming strategy recommendation engines campaign measurement and operational dashboards
- Document technical designs PySpark jobs data flows and configuration details to ensure maintainability and smooth handover across distributed development teams
- Ensure compliance with data governance security and privacy standards relevant to media and entertainment usage while working within day shift schedules
- Troubleshoot production pipeline issues perform root cause analysis and implement corrective actions to minimize disruption to critical reporting and analytics for content operations
- Contribute to continuous improvement initiatives by proposing enhancements to data pipeline architecture coding practices and automation capabilities in the PySpark ecosystem
- Support testing activities by preparing sample datasets validating outputs and collaborating with quality teams to ensure that delivered data solutions align with business expectations
Qualifications
- Demonstrate strong proficiency in PySpark programming and distributed data processing with a track record of building scalable data pipelines for analytical workloads
- Possess solid understanding of media and entertainment domain concepts including audience measurement content metadata streaming events advertising metrics and catalog management
- Bring experience working with big data platforms and data warehousing technologies that commonly integrate with PySpark driven solutions for large enterprises
- Exhibit familiarity with version control collaborative development practices and hybrid work environments that rely on structured communication and documentation
- Show capability to analyze complex data requirements from media stakeholders and convert them into effective data models and transformation logic using PySpark
- Display knowledge of best practices in performance tuning resource optimization and job scheduling for PySpark applications processing high volume media data
- Demonstrate experience with data quality frameworks validation techniques and monitoring tools that help maintain reliable data assets for business critical reporting