We’re looking for a Data Engineer / Analytics Engineer who can design, build, and optimize scalable data pipelines on cloud platforms and work closely with business teams to turn raw data into reliable, analytics-ready insights.
? Key Responsibilities
• Design and maintain Spark/PySpark-based ETL pipelines for high-volume batch and incremental data processing
• Build and manage AWS Glue Jobs, Glue Crawlers, and S3-based data lake ingestion
• Develop dimensional and fact-dimension models and maintain curated data marts using Redshift/Athena
• Optimize SQL and Spark SQL queries for better performance and efficiency
• Drive cloud cost optimization across storage, compute, and query usage using AWS Cost Explorer and CUR analysis
• Build and maintain BI dashboards for revenue, engagement, and operational reporting
• Handle ad-hoc data requests and convert business requirements into optimized queries and temporary data models
• Collaborate with Product, Business, and Analytics teams on data structures, reporting, and KPI definitions
? Required Skills
• 2+ years of hands-on experience with Python, SQL, and Spark SQL/PySpark
• Strong experience with AWS S3, Redshift, Athena, Glue, and CloudWatch
• Valuable understanding of Data Warehousing and Dimensional Modeling
• Strong knowledge of Partitioning, Indexing, and Query Optimization
• Experience with BI/Analytics tools such as Metabase, Matomo, or similar platforms
• Proven experience in improving data pipeline performance and/or optimizing cloud costs
⭐ Nice to Have
• Experience in OTT/Streaming Media Analytics or Insurance/BFSI domains
• Experience with incremental data processing and near real-time pipelines
• Publications or applied research in Data Engineering
? If you enjoy building scalable data solutions, solving complex data problems, and working with modern AWS data technologies, we’d love to hear from you
? Interested candidates are welcome to apply or share their updated CV at:
[email protected]
📌 Data Engineer (Noida)
🏢 AppSquadz
📍 Noida