16 Sep
|
Cloudrise Solutions
|
Bengaluru
16 Sep
Cloudrise Solutions
Bengaluru
Role & responsibilities
- Designed, developed, and maintained scalable ETL/ELT data pipelines using AWS services, Python, SQL, and PySpark.
- Developed data ingestion pipelines to collect data from databases, APIs, CSV, JSON, and other sources and load it into Amazon S3.
- Used AWS Glue for data extraction, transformation, and loading between different data sources and AWS data stores.
- Developed and optimized PySpark jobs for processing and transforming large volumes of structured and semi-structured data.
- Created and maintained AWS Glue Crawlers, Jobs, Databases, and Tables for data cataloging and processing.
- Used Amazon S3 as a data lake for storing raw, processed, and curated datasets.
- Wrote complex SQL queries involving joins, subqueries, CTEs, aggregations, and window functions for data transformation and analysis.
- Worked with Amazon Redshift for data warehousing, loading, querying, and analytical workloads.
- Used Amazon Athena to query data stored in Amazon S3 and perform ad-hoc data analysis.
- Developed data validation and quality checks to ensure accuracy, completeness, and consistency of datasets.
- Implemented data cleansing, transformation, filtering, aggregation, and standardization as part of ETL processes.
- Created reusable Python scripts for data processing, automation, validation, and file handling.
- Worked with AWS Lambda to implement event-driven data processing and automation.
- Used AWS CloudWatch for monitoring jobs, reviewing logs, and troubleshooting pipeline failures.
- Implemented appropriate IAM roles and policies to provide secure access to AWS resources.
- Scheduled and orchestrated data workflows using AWS Step Functions / Apache Airflow.
- Worked with different file formats including CSV, JSON, Parquet, and Avro.
- Implemented partitioning and optimized data storage using suitable formats such as Parquet to improve query performance.
- Identified and resolved data pipeline failures, performance issues, and data quality problems.
- Optimized SQL queries and Spark jobs to improve data processing performance and resource utilization.
- Used Git/GitHub for source-code management, version control, and collaborative development.
- Followed coding standards, documentation practices, and best practices for developing maintainable data pipelines.
- Collaborated with developers, analysts, and business teams to understand data requirements and deliver reliable datasets.
- Monitored scheduled data pipelines and ensured successful completion of daily/periodic data processing jobs.
- Participated in testing, debugging, deployment, and production support activities.
- Documented data pipelines, transformations, data sources, and technical processes for future maintenance.
Preferred candidate profile
- Candidate with strong knowledge of AWS Data Engineering, Python, SQL, and ETL/ELT concepts.
- Hands-on knowledge of Amazon S3, AWS Glue, Amazon Redshift, Amazon Athena, AWS Lambda, and CloudWatch.
- Good understanding of data pipelines, data transformation, data integration, and data warehousing.
- Knowledge of PySpark and Apache Spark for processing large datasets.
- Strong understanding of SQL, including joins, subqueries, CTEs, aggregations, and window functions.
- Familiarity with data lake and data warehouse architecture.
- Understanding of data quality, validation, cleansing, and performance optimization.
- Knowledge of AWS IAM, security, monitoring, and troubleshooting.
- Familiarity with Git/GitHub, Linux, and shell scripting.
- Positive analytical and problem-solving skills with the ability to troubleshoot data pipeline issues.
- Ability to work collaboratively with technical and business teams.
- Willingness to learn new technologies and adapt to evolving cloud and data engineering environments.
- Strong communication skills and ability to understand business requirements and translate them into technical solutions.
Perks and benefits Competitive Salary | Health Insurance | PF | Paid Leave | Flexible Working Hours | Hybrid/Remote Work | Learning & Development | AWS Certification Support | Career Growth | Performance Incentives | Employee Recognition | Work-Life Balance
📌 Aws Data Engineer (Bengaluru)
🏢 Cloudrise Solutions
📍 Bengaluru