11 Sep
|
Infitoo Systems
|
Noida
11 Sep
Infitoo Systems
Noida
Full time opportunity with Innodata : Sr. Data Engineer
Experience:
- 6 to 10 years of hands-on experience in Data Engineering, ETL development, and data platform implementation.
- Proven experience designing, building, and maintaining scalable data pipelines for both batch and real-time data processing.
Technical Skills
- Solid proficiency in SQL and Python.
- Experience with distributed data processing frameworks such as Apache Spark and PySpark.
- Hands-on experience with cloud platforms, including Microsoft Azure, AWS, or Google Cloud Platform (GCP).
- Experience with cloud-native data services such as:
Azure Data Factory
Azure Synapse Analytics
AWS Glue
Amazon Redshift
Google BigQuery
Equivalent cloud-native data services
- Hands-on experience with Databricks for data engineering and analytics workloads.
Strong understanding of:
- Data warehousing concepts
- Dimensional modeling
- Data lake and lakehouse architectures
Experience with workflow orchestration tools such as:
- Apache Airflow
- Azure Data Factory
- Similar scheduling and orchestration platforms
Knowledge of messaging and streaming platforms, including:
- Apache Kafka
- Azure Event Hubs
- Amazon Kinesis
- Experience working with both relational and NoSQL databases.
Data Engineering & Architecture
- Design, develop, optimize, and maintain scalable ETL/ELT pipelines.
- Build reliable, high-performance, and scalable data processing solutions.
- Implement data quality frameworks, validation rules, monitoring, and logging mechanisms.
- Optimize query performance, partitioning strategies, and storage for large-scale datasets.
- Ensure data security, governance, privacy, and compliance with organizational and regulatory standards.
DevOps & Automation: Experience with Git-based version control systems.
Knowledge of CI/CD pipelines using:
- Azure DevOps
- GitHub Actions
- Jenkins
- Similar automation platforms
Familiarity with Infrastructure as Code (IaC) tools such as:
- Terraform
- ARM Templates
- AWS CloudFormation
Experience with containerization and orchestration technologies, including:
Docker
Kubernetes (preferred)
Key Responsibilities:
End-to-End Data Architecture: Design, build, and continuously evolve all layers of the data ecosystem, from multi-source data ingestion and transformation to semantic layers and business-ready data marts.
Scalable Data Pipelines: Develop and optimize high-volume ETL/ELT pipelines capable of processing complex structured, semi-structured, and unstructured datasets.
Data Modeling: Design and maintain robust dimensional models, including star schemas, to support enterprise-wide analytics and reporting.
Cross-Functional Collaboration: Partner with Product Managers, Architects, Delivery Managers, and business stakeholders to translate business requirements into scalable technical solutions.
Workflow Orchestration: Build, schedule, and manage resilient data workflows using Apache Airflow or equivalent orchestration platforms, ensuring high availability and fault tolerance.
Streaming & Real-Time Processing: Design and implement real-time data pipelines using technologies such as Apache Kafka, Amazon Kinesis, or equivalent streaming platforms.
Data Quality & Engineering Excellence: Establish and enforce best practices for code reviews, CI/CD, automated data testing, observability, and data quality to ensure reliable and trusted data assets.
Performance Optimization: Continuously optimize data processing, storage, and query performance to improve scalability and cost efficiency.
Mentorship & Leadership: Mentor junior and mid-level engineers, promote engineering best practices, and contribute to a culture of technical excellence and continuous improvement.
📌 Senior Data Engineer (Noida)
🏢 Infitoo Systems
📍 Noida