24 Sep
|
Qentelli
|
Hyderabad
24 Sep
Qentelli
Hyderabad
JOB SUMMARY:
- We are looking for an experienced AWS Data Engineer with hands-on experience in building and maintaining data pipelines, data platforms, and cloud-based data solutions on AWS.
- The candidate should have strong experience in Python, SQL, AWS data services, ETL/ELT pipelines, and data warehousing, along with practical exposure to Generative AI applications.
- The role will involve working with large volumes of structured and unstructured data, building reliable data pipelines, preparing data for AI applications, and supporting solutions that use Generative AI.
- The ideal candidate should be comfortable working with business and technical teams, understanding data requirements, solving data quality issues, and delivering scalable data solutions.
ROLES AND RESPONSIBILITIES:
- Design, develop, and maintain scalable data pipelines on AWS.
- Build ETL/ELT processes to extract data from multiple sources, transform it, and load it into data platforms.
- Develop data solutions using AWS services such as S3, Glue, Lambda, Redshift, Athena, EMR, and Step Functions.
- Develop and optimize SQL queries for data processing and reporting.
- Write Python scripts and applications for data processing and automation.
- Work with structured and unstructured data from databases, APIs, files, and other sources.
- Build data models and prepare datasets for analytics and AI applications.
- Implement data validation and quality checks across data pipelines.
- Monitor data pipelines and troubleshoot failures, performance issues, and data-related problems.
- Improve pipeline performance, reliability, and cost efficiency.
- Work with large datasets and optimize data processing jobs.
- Integrate data from relational databases, cloud storage, APIs, and enterprise applications.
- Build and maintain data warehouses, data lakes, and related data processing solutions.
- Implement appropriate security, access controls,
encryption, and data governance practices on AWS.
- Work with DevOps teams to deploy and manage data solutions across development, testing, and production environments.
- Create technical documentation for data pipelines, data models, and processes.
- Collaborate with data scientists, software engineers, analysts, and business teams to understand requirements.
- Support data preparation for Generative AI applications, including collecting, cleaning, transforming, and organizing relevant data.
- Work with text documents and other unstructured data used by AI applications.
- Build data pipelines that support AI-powered search, question-answering, and documentbased applications.
- Work with teams to store, process, and retrieve information required by Generative AI applications.
- Monitor the quality and accuracy of data used by AI applications.
- Participate in design discussions, code reviews, testing, and production support.
MANDATORY SKILLS:
- 4+ years of experience in Data Engineering.
- Strong hands-on experience with AWS data services.
- Strong experience with AWS S3, Glue, Redshift, Athena, Lambda, and IAM.
- Experience developing ETL/ELT data pipelines.
- Robust knowledge of data warehousing and data lake concepts.
- Experience working with large datasets.
- Good understanding of batch and incremental data processing.
- Experience with data pipeline monitoring and troubleshooting.
Programming & Database
- Strong programming experience in Python.
- Strong SQL skills, including joins,
subqueries, CTEs, window functions, and query optimization.
- Experience with relational databases such as PostgreSQL, MySQL, Oracle, or SQL Server.
- Good understanding of data modeling and database design.
- Experience working with JSON, CSV, Parquet, and other common data formats.
Gen AI Experience
- Practical experience working on applications that use Generative AI.
- Experience preparing and processing documents or other unstructured data for AI applications.
- Understanding of how data is collected, cleaned, processed, stored, and retrieved for Generative AI use cases.
- Experience working with embeddings and vector-based search.
- Experience with at least one vector database or AWS-supported vector search solution.
- Experience integrating data pipelines with AI/Gen AI applications.
- Basic understanding of how document-based question-answering or search applications work.
- Experience working with APIs or cloud services used by Generative AI applications.
Engineering Practices
- Experience with Git and source-code management.
- Experience with CI/CD processes.
- Good understanding of AWS security and IAM.
- Experience with logging, monitoring, error handling, and production support.
- Strong problem-solving and debugging skills.
Preferred Skills
- Experience with AWS AI services.
- Experience with Amazon OpenSearch or other vector search technologies.
- Experience working with document processing pipelines.
- Experience with PDF, Word, HTML, XML, or text document processing.
- Experience with Spark or PySpark.
- Experience with AWS EMR.
- Experience with Apache Airflow or similar workflow orchestration tools.
- Experience with Terraform or CloudFormation.
- Experience with Docker and container-based deployments.
- Experience with REST APIs and API-based data integration
📌 Aws Data Engineer - Gen AI (Hyderabad)
🏢 Qentelli
📍 Hyderabad