30 Sep
|
Tata Consultancy Services
|
Bengaluru
30 Sep
Tata Consultancy Services
Bengaluru
Role & responsibilities
Section
Details / Example Content
Job Requirements*
We are seeking an experienced Scala / Pyspark associate to design, develop, and maintain scalable data solutions on the Amazon Cloud Platform (AWS). The ideal candidate will have strong expertise in building modern data pipelines, data warehousing, big data processing, and cloud-native analytics solutions. The role requires close collaboration with business stakeholders, data architects, data scientists, and application teams to deliver reliable and high-performance data platforms.
Key Responsibilities*
- Design, develop, and optimize scalable data pipelines using scala and Pyspark in AWS environment.
- Build and maintain batch and real-time data ingestion and processing frameworks.
- Develop enterprise-grade data warehousing solutions using Scala and Pyspark.
- Analyze existing Scala and Spark applications and identify migration requirements.
- Convert Scala-based ETL, batch, and streaming pipelines into PySpark frameworks.
- Optimize PySpark jobs for performance, scalability, and resource utilization.
- Support cloud modernization initiatives on AWS/Databricks/Snowflake platforms.
- Implement ETL/ELT processes for structured and unstructured data.
- Integrate data from multiple sources including databases, APIs, files, and streaming platforms.
- Ensure data quality, governance, security, and compliance across data platforms.
- Automate deployment and operational processes using CI/CD and Infrastructure as Code (IaC).
- Monitor data pipelines and troubleshoot production issues.
- Collaborate with Data Architects, Business Analysts, and Data Scientists to translate business requirements into technical solutions.
- Implement data models, metadata management, and data lineage best practices.
- Support migration of on-premises or multi-cloud data platforms to AWS.
- Lead the migration, modernization, and optimization of Scala and Apache Spark workloads on Cloud setting, including conversion of Scala-based Spark applications to PySpark, performance tuning, cluster optimization, dependency management, and ensuring scalable, cost-effective, and resilient data processing solutions.
Required Technical Skills
- Scala
- PySpark
- AWS Glue
- AWS S3
- Step Functions
- Glue
- Lambda
- Pub/Sub
- Python
- SQL (Advanced)
- Event Bridge
- ECS
- EKS
- Snowflake
Section IV - Job Qualifications & Skills
Section
Details / Example Content
Domain
CMTS
Soft Skills
- Excellent communication
- Team collaboration
- Documentation and knowledge sharing
Education Requirements Bachelor's/masters in computer science or equivalent (Preferred)
Certifications
AWS Cloud Data Engineer (Preferred)
Snow pro certifications
Section V - Sample Questions for Training Model
Technical Question
Expected Answer
What to Evaluate
1.How does Scala code run on the JVM, and what JVM issues should a Scala developer understand?
Scala is compiled to JVM bytecode. A senior candidate should understand heap and stack behavior, garbage collection, boxing, class loading, JIT compilation, thread dumps, heap dumps, and Java library interoperability. They should also know that some high-level Scala constructs may create allocations or additional generated classes.
JVM production experience rather than only Scala syntax. Ask which diagnostic tools they used and how they interpreted results.
2.How would you migrate a large application from Scala 2 to Scala 3?
Assess compiler and library compatibility, upgrade to a suitable Scala 2.13 version, remove deprecated features, reduce macro and implicit complexity, validate cross-building support, migrate incrementally, and run complete automated tests and performance benchmarks. Special attention is required for macros, implicits, reflection, syntax changes, and third-party dependencies.
Real migration experience, dependency assessment, phased planning,
backward compatibility, testing, and rollback strategy.
1. How do you determine the appropriate number of partitions?
Consider data volume, file sizes, cluster cores, operation type, shuffle volume, compression, and target output size. There is no universal partition count. The decision should be validated using task duration, partition-size distribution, available parallelism, spill, and Spark UI metrics. Whether the candidate gives an evidence-based answer instead of a fixed number.
1. What is partition pruning?
Partition pruning allows Spark to read only the storage partitions matching filter conditions. It works best when filters directly reference partition columns and when the data source and predicate form allow pruning. Poor filter expressions or inappropriate partition design can cause full scans. Understanding of physical storage, query filters, and scan reduction.
1. Describe a AWS data pipeline you designed that handled both batch and real-time data processing requirements
Candidate should explain: Real-Time Layer
- Transaction events were published to Amazon Kinesis Data Streams.
- Spark Structured Streaming jobs running on Amazon EMR consumed the data.
- Real-time transformations, validations, and enrichment were performed using PySpark.
- Processed data was stored in Amazon S3 and loaded into Amazon Redshift for reporting.
- Critical alerts were sent through AWS Lambda and SNS whenever fraud thresholds were breached.
Batch Layer
- Data from operational databases (PostgreSQL and Oracle) was extracted daily.
- AWS Glue jobs written in PySpark performed ETL transformations.
- Curated data was stored in an S3 Data Lake using Parquet format.
- Aggregated data was loaded into Redshift for BI dashboards.
End-to-end understanding of data pipeline built and service integration lifecycle and readiness of output data to be loaded.
1. How would you optimize a spark sql query that is taking several minutes to run on multi-terabyte datasets?
Candidate should discuss:
- Reviewing execution plan and query stages.
- Using partitioned tables.
- Applying clustering on frequently filtered columns.
- Avoiding SELECT *.
- Reducing unnecessary joins.
- Materialized views for recurring aggregations.
- Incremental processing approach.
Understanding of efficient, production ready optimized process
1. What approach would you follow to build a cost-optimized data platform on AWS?
Candidate should explain:
Storage optimization
- Lifecycle policies
- Coldline/Archive storage
- Compute optimization:
- Autoscaling
- Right-sizing clusters
Spark optimization:
- Partitioning
- Clustering
- Materialized views
- Monitoring:
- Budget alerts
- Cost dashboards
Candidate should understand About Spark Architecture,
Compute Pricing model and
Storage Pricing model.
1. Explain a data migration project from on-premise to AWS that you were involved in.
Candidate should describe:
Migration Approach
- Performed application and dependency assessment.
- Migrated historical data using AWS DMS.
- Enabled CDC (Change Data Capture) for ongoing replication.
- Validated source and target data counts.
- Cut over reporting workloads to AWS after successful testing.
Data Engineering Activities Built PySpark ETL pipelines for:
- Data cleansing
- Deduplication
- Business transformations
- Data quality validations
- Stored processed data in Parquet format on S3.
- Loaded curated datasets into Redshift.
Attention to data accuracy and reconciliation discipline, s services and understanding all the on premises services and
1. What factors would you consider when optimizing Apache Spark jobs running on EMR?
Candidate should explain:
- Store data in Amazon S3 Data Lake.
- Use EMR Managed Scaling.
- Use Glue Catalog for metadata management.
- Enable EMRFS Consistent View when needed.
- Monitor using CloudWatch Metrics and Alarms.
Candidate should understand End to end architecture of
EMR
1. How would you troubleshoot a AWS Glue job that is running slowly or failing intermittently?
Candidate should explain:
- Review job execution history and identify failure patterns.
- Check CloudWatch logs for errors and performance bottlenecks.
- Validate source and target system connectivity.
- Analyze data volume growth and processing trends.
- Verify data quality and schema consistency.
- Review Spark job performance for skew, shuffle, and resource utilization.
- Check partitioning strategy and file formats.
- Assess Glue worker configuration and DPU sizing.
- Investigate database query performance if reading from RDS/PostgreSQL.
- Validate IAM permissions and security-related failures.
- Check for infrastructure or network-related issues.
- Review recent code, configuration, or deployment changes.
- Compare successful versus failed job runs.
- Optimize resource allocation, transformations, and data access patterns.
- Implement monitoring, alerting, and automated retries to improve reliability.
Candidate should explain dataflow architecture and job optimization process
1. Describe a complex orchestration workflow you implemented using Cloud Composer (Airflow).
Candidate should explain:
- AWS Step Functions as the central orchestrator
- AWS Lambda for control and validation tasks
- AWS Glue for ETL processing
- Amazon EMR (PySpark) for large-scale transformations
- Amazon S3 as the Data Lake
- Amazon Redshift as the target warehouse
- Amazon SNS for notifications
- CloudWatch for monitoring and alerts
Workflow Steps
- Trigger workflow based on schedule or file arrival.
- Validate source file availability.
- Perform metadata and schema validation.
- Load raw data into S3.
- Run AWS Glue/PySpark transformation jobs.
- Execute parallel processing for multiple datasets.
- Perform data quality and reconciliation checks.
- Load curated data into Redshift.
- Update audit and control tables.
- Send success/failure notifications.
- Archive processed files.
Candidate should understand Step functions concept very well. Candidate
Should be clear about task, task
Group also makes a worker node.
1. What approach would you follow to build a cost-optimized data platform on AWS?
Candidate should explain:
Storage optimization
- Lifecycle policies
- Coldline/Archive storage
- Compute optimization:
- Autoscaling
- Right-sizing clusters
Spark optimization:
- Partitioning
- Clustering
- Materialized views
- Monitoring:
- Budget alerts
- Cost dashboards
Candidate should understand About Spark Architecture,
Compute Pricing model and
Storage Pricing model.
1. Describe a performance issue you encountered in Pyspark and how you resolved it.
Candidate should discuss:
Investigation
- Analyzed Spark UI and execution plan.
- Identified that one executor was processing significantly more data than others.
- Observed large shuffle reads and writes.
- Found that small dimension tables were not being broadcast.
- Noticed lack of partition pruning.
Resolution
- Converted source data from CSV to Parquet.
- Partitioned data by transaction_date.
- Enabled Adaptive Query Execution (AQE).
- Implemented Broadcast Join for dimension tables.
- Applied salting technique to address data skew.
- Reduced unnecessary columns before joins.
- Tuned spark.sql.shuffle.partitions.
Candidate should understand spark architecture and optimization technique
📌 Scala Data Engineer (Scala with AWS) (Bengaluru)
🏢 Tata Consultancy Services
📍 Bengaluru