28 Sep
|
Infosys
|
Bengaluru
- Primary skills :AWS Glue-Technology->Cloud Platform->Amazon Webservices DevOps,Technology->Cloud Platform->AWS Data Analytics->AWS Glue DataBrew,Technology->OpenSystem->Python - OpenSystem
Key Responsibilities
- Design, build, and maintain ETL/ELT pipelines using AWS Glue and Python for batch and incremental data processing.
- Develop and optimize Glue Jobs (PySpark/Python) including job parameters, bookmarks, retries, and performance tuning.
- Implement data ingestion, transformation, and validation logic to ensure accuracy, completeness, and consistency of datasets.
- Integrate pipelines with AWS services (e.g., S3, IAM, CloudWatch) to enable secure, observable, and scalable workflows.
- Troubleshoot job failures, analyze logs/metrics, and implement fixes to improve stability and runtime efficiency.
- Collaborate with cross-functional teams to gather requirements, define data mappings, and deliver well-documented solutions.
- Follow engineering best practices including code reviews, version control, and reusable modular coding patterns.
Minimum
Qualifications:
- Bachelor’s degree or equivalent (e.g., BE/BTech/MSc/MCA/MTech).
- 3–5 years of experience in data engineering, ETL development, or data integration roles.
- Strong hands-on experience with AWS Glue and Python for building production-grade data pipelines.
- Working knowledge of core AWS concepts including security basics (IAM), storage patterns, and monitoring.
- Ability to debug data pipeline issues and deliver reliable solutions with clear documentation.
Preferred
Qualifications:
- Experience with PySpark and distributed data processing patterns within AWS Glue.
- Strong SQL skills and experience working with structured/semi-structured datasets (CSV/JSON/Parquet).
- Exposure to orchestration and scheduling patterns for ETL workflows and dependency management.
- Familiarity with data quality checks, schema evolution handling, and building resilient pipelines.
- Experience collaborating in Agile teams and contributing to CI/CD or automated deployment practices for data jobs.
Valuable to have skills: PySpark, Amazon S3, AWS IAM, Amazon CloudWatch, SQL
📌 AWS Glue, Python (Bengaluru)
🏢 Infosys
📍 Bengaluru