Key Responsibilities:
- Data Pipeline Development
Design, develop, and maintain scalable batch and streaming data pipelines using Apache Spark (PySpark/Scala) and Databricks. Build end-to-end ETL/ELT workflows for ingesting, transforming, and validating data from diverse source systems while ensuring data accuracy, reliability, and performance.
- Data Modeling & Analytics Enablement
Design and maintain effective data models, schemas, and curated datasets that support business analytics, reporting, and visualization tools. Optimize data structures for performance, scalability, and cost across lakehouse and data warehouse platforms.
- Data Integration
Integrate data from multiple internal and external sources, including relational databases, APIs, flat files, and streaming sources. Ensure seamless and reliable data movement across cloud platforms, data lakes, and analytics systems.
- Performance Optimization
Identify and resolve performance bottlenecks in Spark jobs, Databricks workloads, and data storage layers. Tune Spark configurations, optimize queries,
and improve pipeline efficiency to support large-scale data processing.
- Data Quality & Governance
Implement data quality checks, validation rules, and governance standards to ensure trustworthy data. Monitor data quality metrics and proactively address data issues in collaboration with stakeholders.
- Collaboration & Stakeholder Engagement
Work closely with data analysts, data scientists, and business teams to understand requirements and deliver data solutions aligned with business objectives. Partner with platform and cloud teams to ensure architectural consistency and best practices.
- Documentation & Best Practices
Document data pipelines, data models, and technical designs. Follow best practices for software development, version control, CI/CD, and deployment in distributed data environments.
- Continuous Improvement
Stay current with emerging data engineering technologies, Spark and Databricks enhancements,
📌 Aws Data Engineer (Pune)
🏢 Nam Info
📍 Pune