16 Sep
|
IT OPENDOORS
|
Hyderabad
16 Sep
IT OPENDOORS
Hyderabad
:
We are currently looking for a AZURE Data Engineer for the client PepsiCo
. This is a full-time and Hybrid position .
Required Technical Skill Set and Tools:
Programming:
1. Programming Languages: Python (OR) Scala (OR) Java
2. Database Languages: SQL (OR) HQL (OR) PL/SQL (OR) Any SQL like (TSQL , Teradata SQL so on)
3. Scripting Languages: Shell Scripting (OR) UNIX Shell Scripting (OR) Power Shell Scripting
4. ETL data pipeline building experience
ETL
1. ETL Tools: Informatica (OR) Infoworks (OR) DataStage (OR) Abinitio
2. ETL Data Pipelines: Build data pipelines Using PySpark (Spark or Scala or Java) for data processing, Cleansing, transformation and loading
Tools
1. Spark (OR) PYSpark (OR) ScalaSpark (OR) Databricks Spark
2. Hive (OR) Any SQL related databases
3. Sqoop
4. Kafka (OR) Ni-Fi (Streaming Technologies)
5. No-SQL Databases like Mongo DB, Cosmo DB so on
6. Informatica,
7. Infoworks
8. Oozie
9. SPARK on GCP is called DATAPROC
Cloud Technologies: AZURE Databricks/Azure HDI (OR) Google Cloud Platform (GCP) (OR) AWS Databricks/AWS EMR
Cloud SQL databases: Snowflake Cloud (OR) Teradata Cloud
Machine Learning (ML)/ Gen AI Job Duties:
1. Performing in development language environments- e.g. Python, Java, Scala, R, SQL, etc. and applying analytical methods to large and complex datasets leveraging one of those languages
2. Experience in machine learning, natural language processing and deep learning.
3. Candidate should have a working exposure to Generative AI based projects that includes designing and implementing solutions based on Langchain framework and designing productive prompt for LLM’s.
4. Candidate should have good experience in pre-training and fine-tuning Large Language Models (LLM’s) on HuggingFace models & other Large Language Models.
5. Proven ability with NLP and text-based extraction techniques.
6. Familiarity with deep learning architectures used for text analysis, computer vision and signal processing.
7.
Understanding of not only how to develop data science analytic models but how to operationalize these models so they can run in an automated context
8. Utilizing and applying knowledge commonly used data science packages including Spark, Pandas, SciPy, and NumPy.
9. Experience manipulating and analyzing complex, high-volume, high-dimensionality data from varying sources
10. Applying techniques such as multivariate regressions, Bayesian probabilities, clustering algorithms, machine learning, dynamic programming, stochastic-processes, queuing theory, algorithmic
knowledge to efficiently research and solve complex development problems and application of engineering methods to define, predict and evaluate the results obtained.
1. Developing and deploying A.I. solutions as part of a larger automation pipeline
2. Demonstrates extensive abilities and/or a proven record of success in the application of statistical modelling, algorithms, data mining and machine learning algorithms problem solving
Cloud Data Engineer Job Duties:
1. Analysis of Business requirement to understand what users need, clarify with users on open questions and resolutions.
2. Design and develop spark Dataframes/RDD, real time streaming application using python and Scala programming languages.
3. Design and develop Hadoop Map Reduce, real time streaming application using Pytho/Scala/Java programing language, Kafka and Flume.
4. Design and Develop Apache Hive tables, hive external tables, partitioning, bucketing, performance tuning.
5. Analyze data volume and design / develop various file formats like parquet, Avro etc.
6. Develop and implemented Databricks,
Hadoop and Big data applications and processes utilizing tools and technologies such as Databricks(Azure/AWS) Streaming, Spark, Scala, Python, HDFS, Hive, Sqoop, Oozie, Kafka, Flume, Zookeeper, etc.
7. Schedule Spark and hive jobs in scheduler using Tidal, Control-M and Autosys.
8. Write SparkSQL queries on source data to analyses data in Google Big Query, Apache Hive and Google Big Table databases.
9. Write Spark Dataframe scripts as required, Issue analysis and providing resolutions by implementing technical solution.
10. Maintain in-depth knowledge of IT industry best practices, technologies, architectures and emerging technologies.
11. Design of Conceptual, Logical and Physical Datamodels for OLTP Databases and OLAP Data Warehouses across various data domains like lifesciences, Retail, Insurance so on.
12. Proficient in Data Modeling techniques using Star Schema, Snowflake Schema, Fact and Dimension tables.
13. Expert in Business Modeling Techniques using Process Flow Modeling and Data Flow Modeling.
14. Extensive experience of using data modelling tools like ER Studio Data Architect and ERwin in Relational and Dimensional Data modeling for creating Logical and Physical Design of Database and ER Diagrams using ER Studio data modeling tool, this includes creating the base tables, associative tables, and views, and choosing and creating all indexes.
15. Work and deliver forward and reverse engineering processes. Create DDL scripts for implementing Data modeling changes. Created ER Studio reports in HTML, RTF format depending upon the requirement, published Data model in model mart, created naming convention files and replicated these data model changes in the database.
16. Work on writing functional specifications, translating business requirements to technical specifications, created/maintained/modified data base design document with detailed description of logical entities and physical tables.
📌 Azure Data Engineer (Hyderabad)
🏢 IT OPENDOORS
📍 Hyderabad