19 Sep
|
Tata Consultancy Services
|
Bengaluru Urban
19 Sep
Tata Consultancy Services
Bengaluru Urban
Job Requirements*
We are seeking an experienced Scala / Pyspark associate to design, develop, and maintain scalable data solutions on the Amazon Cloud Platform (AWS). The ideal candidate will have strong expertise in building up-to-date data pipelines, data warehousing, big data processing, and cloud-native analytics solutions. The role requires close collaboration with business stakeholders, data architects, data scientists, and application teams to deliver reliable and high-performance data platforms.
Key Responsibilities*
- Design, develop, and optimize scalable data pipelines using scala and Pyspark in AWS environment.
- Build and maintain batch and real-time data ingestion and processing frameworks.
- Develop enterprise-grade data warehousing solutions using Scala and Pyspark.
- Analyze existing Scala and Spark applications and identify migration requirements.
- Convert Scala-based ETL, batch, and streaming pipelines into PySpark frameworks.
- Optimize PySpark jobs for performance, scalability, and resource utilization.
- Support cloud modernization initiatives on AWS/Databricks/Snowflake platforms.
- Implement ETL/ELT processes for structured and unstructured data.
- Integrate data from multiple sources including databases, APIs, files, and streaming platforms.
- Ensure data quality, governance, security, and compliance across data platforms.
- Automate deployment and operational processes using CI/CD and Infrastructure as Code (IaC).
- Monitor data pipelines and troubleshoot production issues.
- Collaborate with Data Architects, Business Analysts, and Data Scientists to translate business requirements into technical solutions.
- Implement data models, metadata management, and data lineage best practices.
- Support migration of on-premises or multi-cloud data platforms to AWS.
- Lead the migration, modernization, and optimization of Scala and Apache Spark workloads on Cloud environment, including conversion of Scala-based Spark applications to PySpark, performance tuning, cluster optimization, dependency management, and ensuring scalable, cost-effective, and resilient data processing solutions.
Required Technical Skills
- Scala
- PySpark
- AWS Glue
- AWS S3
- Step Functions
- Glue
- Lambda
- Pub/Sub
- Python
- SQL (Advanced)
- Event Bridge
- ECS
- EKS
- Snowflake
Section IV - Job Qualifications & Skills
Section
Details / Example Content
Domain
CMTS
Soft Skills
- Excellent communication
- Team collaboration
- Documentation and knowledge sharing
Education Requirements
Bachelor's/master’s in computer science or equivalent (Preferred)
Certifications
AWS Cloud Data Engineer (Preferred)
Snow pro certifications
📌 Software Associate (Bengaluru Urban)
🏢 Tata Consultancy Services
📍 Bengaluru Urban