11 Sep
|
ALOIS Solutions
|
Bengaluru
11 Sep
ALOIS Solutions
Bengaluru
8 years of relevant experience.
Should be flexible to adapt to project requirements needs.
Key Responsibilities:
- Data Pipeline Development: Design, build, and optimize robust, scalable, and efficient ETL/ELT data pipelines using Python and PySpark, primarily within Azure Databricks and Azure Data Factory.
- Data Ingestion & Processing: Develop and manage processes for ingesting data from various sources (e.g., transactional databases, APIs, streaming sources) and transforming it into clean, usable formats for downstream consumption.
- Data Quality & Monitoring: Implement comprehensive unit and integration test coverage for data pipelines. Establish and maintain monitoring, alerting, and dashboarding solutions (e.g., Grafana) for data quality, pipeline health, and performance.
- Cloud Infrastructure Management (OpenShift/Azure): Contribute to the setup, configuration, and maintenance of data-related infrastructure on OpenShift, ensuring deployment readiness and leveraging tools like HELM for application packaging and deployment.
- CI/CD & Automation: Drive CI/CD best practices using GitHub Actions, ensuring automated testing (unit tests), build, and deployment processes for data solutions to environments like OpenShift.
- SQL & Data Modeling: Develop and optimize complex SQL queries for data extraction, transformation, and loading. Apply strong data modeling principles for efficient data storage and retrieval in SQL Server and other data stores.
- Azure Ecosystem Leverage: Utilize a broad range of Azure data and analytics services, including Azure Data Factory, Azure Databricks, Azure SQL Server, Azure Key Vault, and others to build comprehensive data solutions.
- Performance Optimization: Proactively identify and resolve performance bottlenecks in data pipelines and databases through query optimization, indexing strategies,
and efficient data processing techniques.
- Collaboration & Documentation: Work closely with data scientists, analysts, and other engineering teams to understand data requirements. Create explicit and concise documentation for data pipelines, architecture, and processes.
Required Core Skills & Qualifications:
- Programming & Data Processing: Strong proficiency in Python and PySpark for large-scale data processing and ETL development.
- Data Warehousing & SQL: Expertise in SQL for complex querying, data manipulation, and schema design.(Optional, but highly preferred): Proven experience in SQL optimization and performance tuning.
- ETL Development: Demonstrable experience in designing, building, and maintaining robust ETL/ELT data pipelines.
- Cloud Data Platform (Azure Focus):Hands-on experience with Azure Databrick. Proficiency with core Azure Analytics Services including Azure Data Factory, Azure SQL Server, and Azure Key Vault.
- DevOps & CI/CD: Experience implementing CI/CD pipelines from GitHub (including GitHub Actions) for automated testing (unit tests), build, and deployment processes.
- Containerization & Orchestration: Familiarity and practical experience with OpenShift (setup, deployment-ready configurations, and management).Experience with HELM for deploying applications on Kubernetes/OpenShift.
- Monitoring & Observability: Experience in setting up and configuring Grafana for dashboards to monitor data quality and pipeline health.
Preferred Qualifications: Bachelor's or Master's degree in Computer Science, Engineering, Data Science, or a related quantitative field.
- Relevant Azure certifications (e.g., Azure Data Engineer Associate).
- Experience with real-time data processing frameworks (e.g., Kafka, Azure Event Hubs).
- Understanding of data governance, data security, and compliance best practices.
Qualifications BE
Additional Information
6-8
📌 Azure Data Engineer (Bengaluru)
🏢 ALOIS Solutions
📍 Bengaluru