13 Aug
|
Calitii
|
Bangalore Metropolitan Area
13 Aug
Calitii
Bangalore Metropolitan Area
Job Summary
Synechron is seeking an experienced PySpark Data Engineer / Data Scientist to lead data pipeline development and advanced analytics initiatives within our financial data and index analytics division. This role plays a crucial part in building scalable data processing solutions, enabling data-driven insights, and supporting machine learning workflows in both batch and streaming environments. The ideal candidate will possess a strong technical foundation in big data processing, analytics, and software engineering, along with leadership capabilities to drive impactful data projects.
Software Requirements
Required Skills:
- Proven expertise in Python programming, emphasizing clean, maintainable, and scalable code
- Hands-on experience with PySpark in both batch and streaming workflows
- Deep knowledge of data manipulation and feature engineering, including Pandas, NumPy, and visualization libraries (matplotlib, seaborn)
- Experience with Spark components like Spark SQL, DataFrames, and Spark MLlib
- Familiarity with data storage solutions: SQL and NoSQL databases (e.g., Hive,
Cassandra)
- Knowledge of ETL tools such as Apache Airflow, Jenkins, or GithHub Actions for scheduling and automation
- Experience working with cloud environments, especially Azure or AWS for big data processing
Preferred Skills
- Hands-on with containerization and orchestration (Docker, Kubernetes)
- Exposure to distributed storage solutions like Hadoop HDFS or Azure Data Lake
Overall Responsibilities
- 5 years of experience in Design, develop, and optimize large-scale data pipelines using PySpark for structured, semi-structured, and unstructured data
- 5 years of experience to Lead the building of ML pipelines for training, validation, and deployment of models in streaming/batch modes
- Write high-quality, productive code that supports data transformation, cleaning, and feature engineering
- Collaborate with data scientists, analysts, and stakeholders to understand data requirements
📌 PySpark Data Engineer | Big Data (Bangalore Metropolitan Area)
🏢 Calitii
📍 Bangalore Metropolitan Area