We are looking for part time Freelancer for AWS Data Engineer Position.
Releavant Exp : 9+
Mandatory Tech Stack required:
Source: Salesforce (ECRM & OSC)
Backup: Grax
Cloud: AWS (EC2, S3, CloudWatch, DynamoDB, RDS, Secrets Manager, ALB, ASG)
Infrastructure as Code: Terraform
File Format: Parquet
Target: Enterprise Data Lake (EDL)
Schema: Blue Schema
Database: Blue Database
Query Language: SQL
JD :
1.
Build and maintain AWS data pipelines
* Develop ETL/ELT pipelines using AWS Glue, PySpark, Python, and SQL.
* Ingest data from sources like Salesforce, databases, APIs, and S3.
* Load curated data into Amazon Redshift and data lake storage.
2.
Optimize Athena and data lake performance
* Convert JSON/CSV data into Parquet.
* Use Snappy compression.
* Design proper partitioning strategies.
* Resolve split limit and performance issues.
* Optimize Athena query costs.
3.
Manage contemporary data lake architecture
* Work with Apache Iceberg tables.
* Perform migrations from traditional Parquet tables.
* Support schema evolution, time travel, and ACID transactions.
4.
Production support and troubleshooting
* Investigate Glue jobs that suddenly become slow.
* Debug Lambda timeouts.
* Fix missing records and data quality issues.
* Resolve Redshift performance problems.
* Perform root cause analysis (RCA).
5.
Infrastructure as Code
* Build AWS infrastructure using Terraform.
* Create reusable modules.
* Manage Auto Scaling Groups, ALBs, IAM, Lambda, Secrets Manager, S3, Redshift, and DynamoDB.
* Troubleshoot Terraform state and production deployment issues.
6.
Security
* Manage AWS Secrets Manager.
* Configure Lambda-based secret rotation.
* Ensure Terraform does not overwrite rotated passwords.
* Implement IAM least-privilege access.
7.
Backup and Disaster Recovery
* Work with enterprise backup tools like Rubrik (their environment may use Grax for Salesforce).
* Validate backup jobs.
* Perform restores.
* Support disaster recovery testing.
* Verify restored data.
8.
Data Quality
* Validate source and target record counts.
* Maintain audit/control tables (for example in DynamoDB).
* Check checksums and duplicate records.
* Troubleshoot data discrepancies reported by business users.