20 Sep
|
Tech Mahindra
|
Bengaluru
20 Sep
Tech Mahindra
Bengaluru
Key Responsibilities:
n
n
n
- Data Pipeline Development: Design, build, and optimize robust, scalable, and efficient ETL/ELT data pipelines using Python and PySpark, primarily within Azure Databricks and Azure Data Factory.
n
- Data Ingestion &
• Processing: Develop and manage processes for ingesting data from various sources (e.g., transactional databases, APIs, streaming sources) and transforming it into clean, usable formats for downstream consumption.
n
- Data Quality &
• Monitoring: Implement comprehensive unit and integration test coverage for data pipelines. Establish and maintain monitoring, alerting, and dashboarding solutions (e.g., Grafana) for data quality, pipeline health, and performance.
n
- Cloud Infrastructure Management (OpenShift/Azure): Contribute to the setup, configuration, and maintenance of data-related infrastructure on OpenShift, ensuring deployment readiness and leveraging tools like HELM for application packaging and deployment.
n
- CI/CD &
• Automation: Drive CI/CD best practices using GitHub Actions, ensuring automated testing (unit tests), build, and deployment processes for data solutions to environments like OpenShift.
n
- SQL &
• Data Modeling: Develop and optimize complex SQL queries for data extraction, transformation, and loading. Apply solid data modeling principles for efficient data storage and retrieval in SQL Server and other data stores.
n
- Azure Ecosystem Leverage: Utilize a broad range of Azure data and analytics services, including Azure Data Factory, Azure Databricks, Azure SQL Server, Azure Key Vault, and others to build comprehensive data solutions.
n
- Performance Optimization:
Proactively identify and resolve performance bottlenecks in data pipelines and databases through query optimization, indexing strategies, and efficient data processing techniques.
n
- Collaboration &
• Documentation: Work closely with data scientists, analysts, and other engineering teams to understand data requirements. Create clear and concise documentation for data pipelines, architecture, and processes.
n
n
nRequired Core Skills & Qualifications:
n
n
n
- Programming &
• Data Processing: Strong proficiency in Python and PySpark for large-scale data processing and ETL development.
n
- Data Warehousing &
• SQL: Expertise in SQL for complex querying, data manipulation, and schema design.(Optional, but highly preferred): Proven experience in SQL optimization and performance tuning.
n
- ETL Development: Demonstrable experience in designing, building, and maintaining robust ETL/ELT data pipelines.
n
- Cloud Data Platform (Azure Focus):Hands-on experience with Azure Databrick. Proficiency with core Azure Analytics Services including Azure Data Factory, Azure SQL Server, and Azure Key Vault.
n
- DevOps &
• CI/CD: Experience implementing CI/CD pipelines from GitHub (including GitHub Actions) for automated testing (unit tests), build, and deployment processes.
n
- Containerization &
• Orchestration: Familiarity and practical experience with OpenShift (setup, deployment-ready configurations, and management).Experience with HELM for deploying applications on Kubernetes/OpenShift.
n
- Monitoring &
• Observability: Experience in setting up and configuring Grafana for dashboards to monitor data quality and pipeline health.
n
n
📌 Azure Engineer (Bengaluru)
🏢 Tech Mahindra
📍 Bengaluru