Data Engineer (India)

Data Engineer (India)

11 Sep
|
Leader IT
|
India

11 Sep

Leader IT

India

Job Title: Data Engineer

Job Overview

We are looking for a skilled and motivated Data Engineer with 2–3 years of experience to join our team. The ideal candidate should have strong knowledge of SQL, Python, PySpark data pipelines, ETL/ELT processes, databases, and cloud-based data platforms.

The candidate will work closely with data scientists, AI engineers, software developers, analysts, and product teams to build reliable, scalable, and secure data solutions. The role involves collecting data from multiple sources, transforming it into usable formats, maintaining data quality, and making it available for analytics, reporting, AI, and machine learning applications.

Key Responsibilities

- Design, develop, and maintain scalable data pipelines for structured and unstructured data.

- Build and manage ETL/ELT workflows to collect, clean, transform, and load data from multiple sources.

- Integrate data from databases, APIs, files, cloud storage, enterprise applications, and third-party platforms.

- Develop efficient SQL queries, stored procedures, views, and data transformation scripts.

- Work with relational and NoSQL databases to store, process, and retrieve data efficiently.

- Design and maintain data models, data warehouses, data lakes, and analytical datasets.

- Automate data ingestion, transformation, validation, and reporting workflows.

- Monitor data pipelines and troubleshoot failures, performance issues, and data inconsistencies.

- Implement data validation checks to ensure data accuracy, completeness, consistency, and reliability.

- Optimize database queries and data pipelines for performance, scalability, and cost efficiency.

- Develop reusable data-processing components and maintain proper technical documentation.

- Support the preparation of datasets required for dashboards, analytics, AI, and machine learning models.

- Collaborate with data scientists and AI engineers to understand data requirements and provide model-ready datasets.

- Work with developers and product teams to integrate data services into applications and business workflows.

- Follow data security, access control, privacy, governance, and compliance requirements.

- Participate in code reviews, testing, deployment, and production support activities.

- Maintain version-controlled code using Git and follow software-development best practices.

Data Pipeline and Platform Responsibilities

- Build batch and, where required, near-real-time data-processing pipelines.

- Schedule, monitor, and manage data workflows using tools such as Apache Airflow or similar orchestration platforms.

- Process large datasets using Python, SQL, Apache Spark, or equivalent technologies.

- Work with cloud storage, data warehouses, and data-processing services.

- Create data models suitable for reporting, dashboards, operational applications, and AI use cases.

- Maintain logging,



monitoring, alerting, and error-handling mechanisms for production pipelines.

- Identify and resolve data-quality issues at the source, transformation, and consumption levels.

- Assist in migrating data between legacy systems, cloud platforms, databases, and modern data architectures.

Required Skills

- Strong proficiency in SQL, including joins, aggregations, subqueries, window functions, and query optimization.

- Good programming skills in Python for data processing and automation.

- Hands-on experience in developing and maintaining ETL/ELT pipelines.

- Good understanding of relational databases such as PostgreSQL, MySQL, SQL Server, or Oracle.

- Familiarity with NoSQL databases such as MongoDB, Elasticsearch, or similar technologies.

- Understanding of data modelling, database design, normalization, and dimensional modelling.

- Experience working with data formats such as CSV, JSON, XML, Parquet, and Avro.

- Familiarity with REST APIs and integration of data from third-party systems.

- Working knowledge of cloud platforms such as AWS, Microsoft Azure, or Google Cloud Platform.

- Familiarity with cloud data services, data warehouses, or data lakes.

- Basic understanding of Apache Spark, Kafka, Airflow, dbt, or similar data-engineering tools.

- Experience with Git and collaborative software-development workflows.

- Understanding of data quality, security, governance, and access-control principles.

- Robust analytical and problem-solving abilities.

- Knowledge and Experience with PySpark

Preferred Skills

- Experience with cloud data warehouses such as Snowflake, Amazon Redshift, Google BigQuery, Azure Synapse, or Databricks.

- Familiarity with Docker, containerization, and CI/CD pipelines.

- Exposure to real-time data processing and messaging systems such as Apache Kafka.

- Understanding of data orchestration and transformation frameworks such as Airflow and dbt.

- Familiarity with dashboard and reporting tools such as Power BI, Tableau, or Looker.

- Basic understanding of machine learning and AI data requirements.

- Experience working with large-scale datasets or enterprise applications will be an advantage.

AI and Emerging Technologies

Candidates should be interested in supporting modern AI, machine learning, and Generative AI applications through reliable data engineering.

The candidate should be willing to:

- Build and maintain data pipelines for AI and machine learning use cases.

- Prepare clean, structured,



and traceable datasets for model training, evaluation, and inference.

- Work with structured and unstructured data, including text, documents, audio, images, and logs.

- Explore AI-assisted tools for coding, documentation, testing, debugging, and data analysis.

- Support data preparation for vector databases, embeddings, retrieval-augmented generation, and knowledge-base applications.

- Understand the importance of data privacy, governance, security, and responsible AI practices.

- Continuously learn new data-engineering, cloud, AI, and automation technologies.

Familiarity with tools such as GitHub Copilot, ChatGPT, Gemini, or other AI-powered development and productivity tools will be an advantage.

Qualifications

- B.Tech, B.E., M.Tech, M.E., MCA, M.Sc. in Computer Science, Information Technology, Data Science, Software Engineering, or a related discipline.

- 2–3 years of relevant professional experience in data engineering, database development, ETL development, data integration, or a related role.

- Candidates should have hands-on experience in building data pipelines and working with databases, APIs, and data-processing tools.

- Relevant cloud, database, or data-engineering certifications will be an advantage but are not mandatory.

What We’re Looking For

- Strong foundation in SQL, Python, PySpark, databases, and data pipelines.

- Ability to understand business requirements and translate them into reliable data solutions.

- Strong analytical thinking and problem-solving skills.

- Attention to data accuracy, performance, security, and maintainability.

- Ability to troubleshoot pipeline failures and data-quality issues.

- Interest in cloud platforms, AI, analytics, and emerging technologies.

- Willingness to learn and adapt to new tools and frameworks.

- Ability to work independently while collaborating effectively with cross-functional teams.

- A responsible and quality-focused approach to production data systems.

Why Join Us?

- Opportunity to work on real-world data, analytics, AI, and enterprise application projects.

- Exposure to modern cloud-based data platforms and engineering practices.

- Opportunity to build data pipelines supporting AI, machine learning, dashboards, and business applications.

- Collaboration with data scientists, AI engineers, developers, product teams, and domain experts.

- Opportunity to work with structured and unstructured enterprise data.

- Supportive environment for learning new technologies and strengthening technical expertise.

- Growth opportunities for candidates who demonstrate technical capability, ownership, and initiative.

Experience Level

2–3 years of relevant professional experience in Data Engineering, ETL Development, Data Integration, Database Development, or a related field.

📌 Data Engineer (India)
🏢 Leader IT
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: data engineer (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: data engineer (india) / india