Senior Data Engineer (AWS | Python | Data Platform)
Location: Pune (5 Days / Week)
Role Summary
We are looking for a Senior Data Engineer to design, develop, and optimize scalable cloud-based data platforms that power analytics, operational intelligence, and enterprise applications. The ideal candidate is passionate about building reliable data pipelines, designing contemporary data architectures, and enabling data-driven decision making through high-quality, secure, and scalable solutions.
Key Responsibilities:
- Data Engineering & Platform Development Design, develop, and maintain scalable batch and streaming data pipelines on AWS.
- Build robust ETL/ELT frameworks for ingesting data from multiple internal and external sources.
- Develop reusable data services that support analytics, reporting, machine learning, and operational workloads.
- Design efficient data models optimized for both analytical and operational use cases. Ensure data quality, consistency, lineage, and governance across the data platform.
- Cloud & Infrastructure Build and optimize cloud-native data infrastructure using AWS services.
- Design highly available, fault-tolerant, and scalable data processing architectures.
- Automate deployment, monitoring, and operational workflows using Infrastructure as Code and CI/CD practices.
- Optimize storage, compute utilization, and overall cloud costs.
- Performance & Reliability Improve pipeline performance, scalability, and reliability.
- Monitor production workloads and proactively resolve bottlenecks.
- Optimize SQL queries, data partitioning, indexing strategies, and processing performance.
- Troubleshoot production issues and implement preventive improvements.
- Collaboration Work closely with Data Scientists,
Software Engineers, Product Owners, and Architects to understand business requirements.
- Support downstream analytics, dashboards, and reporting solutions.
- Participate in architecture discussions, code reviews, and technical design sessions. Mentor junior engineers and promote engineering best practices.
Required Technical Skills:
- Programming Strong Python programming skills.
- Experience building production-grade data engineering solutions.
- Knowledge of object-oriented design, testing, and clean coding practices.
- AWS Hands-on experience with several of the following: S3 Glue Lambda EMR Athena Redshift RDS Step Functions EventBridge CloudWatch IAM ECS/EKS (preferred) Data Engineering ETL/ELT pipeline development Data warehousing concepts Batch and streaming data processing Data modeling Data validation and quality frameworks Metadata and lineage concepts Databases Advanced SQL PostgreSQL MySQL Redshift Performance tuning and query optimization Development Practices Git CI/CD pipelines Unit testing Docker Agile/Scrum development
Preferred Skills:
- Apache Spark or PySpark Kafka or Kinesis Airflow or AWS Managed Workflows Infrastructure as Code (Terraform or CloudFormation)
- Data Lake architecture Lakehouse concepts
- Experience supporting Machine Learning data pipelines
- Experience with observability and monitoring tools
- Desired Experience 610 years of experience in Data Engineering.
- Experience building cloud-native data platforms on AWS.
- Experience handling large-scale structured and semi-structured datasets.
- Strong understanding of distributed data processing and performance optimization.
- Experience working in enterprise production environments with high availability requirements.
📌 Data Engineer (Pune)
🏢 Calsoft
📍 Pune