30 Jul
|
Yotta Infrastructure
|
Delhi
30 Jul
Yotta Infrastructure
Delhi
Job Scope
We are looking for an enthusiastic Junior Data Platform Engineer to support and manage our Apache Spark, Apache Airflow, and JupyterHub environments. This role is ideal for someone with a strong foundation in Python and Linux, who is eager to build a career in big data engineering and data platform administration.
You will work closely with senior engineers to ensure smooth operation, deployment, and optimization of our data processing ecosystem.
Total /Relevant Experience
1+ years experience
Key Responsibilities
- Assist in the setup, monitoring, and maintenance of Apache Spark clusters, Apache Hive, Hadoop and Airflow environments.
- Support the development and scheduling of data pipelines using Airflow DAGs and Python scripts.
- Help manage and configure JupyterHub for multi-user access and integration with Spark.
- Monitor cluster health and performance under guidance and assist in troubleshooting Spark job failures.
- Write and maintain Python automation scripts for data workflows, ETL, and process automation.
- Participate in code reviews, documentation, and deployment activities.
- Learn and follow best practices for distributed data processing, CI/CD, and DevOps workflows.
- Collaborate with senior engineers and data scientists to implement improvements and current features.
Must-have skill
- Basic understanding of Apache Spark, Apache Hive and Hadoop File System.
- Familiarity with Apache Airflow (understanding of DAGs, scheduling, and task dependencies).
- Hands-on experience with Python scripting (data processing, automation, or API interaction).
- Comfortable working in Linux and container environments (command line, system logs, process management).
- Good understanding of data processing concepts, including ETL and distributed computing.
- Basic knowledge of Git and version control.
Good-to-Have Skills
- Exposure to Jupyter / JupyterHub for collaborative notebook environments.
- Knowledge of Docker or Kubernetes.
- Knowledge of Hadoop and Apache Spark Cluster.
- Familiarity with SQL and working with structured/unstructured data.
- Experience with cloud platforms (AWS, GCP, or Azure) is a plus.
- Interest in big data (Hadoop), DevOps, and data pipeline automation.
Qualifications Criteria
Bachelor's or any relevant Degree.
Certification Criteria
NA
Disclaimer : This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Data Platform Engineer (Delhi)
🏢 Yotta Infrastructure
📍 Delhi