Job Overview:
We are seeking a highly skilled and detail-oriented Python Testing with strong expertise in Python, Hadoop, PySpark, R Programming, Unix, Tableau, AWS, Machine Learning, and Data Testing. The ideal candidate will be responsible for developing, testing, validating, and optimizing large-scale data processing pipelines while ensuring data quality, reliability, and performance across cloud and on-premises environments.
The candidate should possess hands-on experience in big data technologies, cloud platforms, data visualization, and testing methodologies to support enterprise-level analytics and business intelligence initiatives.
Key Responsibilities
- Design, develop, test, and maintain scalable data pipelines using Python and PySpark.
- Perform data validation, functional testing, regression testing, and ETL testing for large-scale data applications.
- Develop and optimize data processing workflows using Hadoop ecosystem components.
- Write efficient Python and PySpark scripts for data ingestion, transformation, and processing.
- Utilize R Programming for statistical analysis, predictive modeling, and data exploration.
- Build interactive dashboards and reports using Tableau to support business decision-making.
- Work with AWS services such as S3, EMR, Glue, Lambda, EC2, Athena, and Redshift for cloud-based data solutions.
- Execute SQL queries and perform data reconciliation between multiple data sources.
- Analyze data quality issues and implement automated validation frameworks.
- Collaborate with developers, business analysts, data scientists, and QA teams to ensure data accuracy and completeness.
- Develop and execute test cases, test plans,
and automation scripts for data validation.
- Monitor production jobs and troubleshoot failures in distributed data processing environments.
- Work in Unix/Linux environments to execute shell scripts, schedule jobs, and manage system operations.
- Participate in code reviews and implement best practices for performance optimization and maintainability.
- Support Machine Learning initiatives by preparing, validating, and transforming datasets for model development.
- Document technical designs, testing procedures, and implementation processes.
Required Skills
Mandatory Skills
- Solid experience in Python programming.
- Hands-on experience with Hadoop ecosystem.
- Expertise in PySpark for distributed data processing.
- Good knowledge of R Programming.
- Strong Unix/Linux command-line skills.
- Experience in Tableau dashboard development and reporting.
- Hands-on experience with AWS cloud services.
- Understanding of Machine Learning concepts and data preparation.
- Strong knowledge of Data Testing, ETL Testing and Data Validation.
- Experience writing complex SQL queries.
- Excellent debugging and analytical skills.
Preferred Skills
- Experience with Apache Spark optimization.
- Knowledge of CI/CD pipelines.
- Experience with Git or other version control systems.
- Familiarity with Agile/Scrum methodologies.
- Experience with workflow orchestration tools such as Apache Airflow.
- Knowledge of data warehousing concepts.
- Exposure to DevOps practices.
Qualifications
- Bachelors or masters degree in computer science, Information Technology, Data Science, Engineering, or a related field.
- Relevant experience in Data Engineering, Big Data Development, Data Analytics, or Data Testing.
📌 Python Testing (Telangana)
🏢 CGI
📍 Telangana