04 Sep
|
Virtusa
|
Bengaluru
"Key Responsibilities
Pipeline Architecture: Design, develop, deploy, and maintain robust, scalable ETL/ELT pipelines to ingest, process, and transform both structured and unstructured data from multiple sources.
Batch & Real-time Processing: Build, monitor, and optimize scalable batch and real-time data pipelines to support business analytics, reporting, and downstream data applications.
Data Modeling & Orchestration: Author clean, modular data models using dbt and automate workflows using Apache Airflow / Amazon MWAA.
Cloud Infrastructure & Optimization: Leverage the AWS ecosystem (Glue, S3, Redshift) to build reliable data architectures and optimize queries, storage, and cluster performance for maximum cost efficiency and speed.
Data Governance & Quality: Ensure data accuracy, integrity, and compliance across pipelines by implementing automated testing, data checks, and error handling.
Cross-Functional Collaboration: Partner with data analysts, data scientists,
and product teams to understand data needs, deliver clean datasets, and troubleshoot data-related issues.
Required Qualifications & Skills
Python Proficiency: Strong, production-grade Python skills for data manipulation, script automation, and custom pipeline development.
Core AWS Stack: Hands-on experience with AWS Redshift, AWS Glue, and Amazon S3.
Data Orchestration: Practical experience using Apache Airflow (or AWS MWAA) to manage complex DAGs and pipeline dependencies.
Contemporary Analytics Stack: Proficient with dbt for building maintainable, version-controlled SQL transformations in the data warehouse.
Database & SQL Expertise: Advanced SQL skills with deep knowledge of relational databases, data warehousing techniques, and data modeling (e.g., Star/Snowflake schemas
📌 Senior Consultant (Bengaluru)
🏢 Virtusa
📍 Bengaluru