06 Aug
|
Sanskruti Solutions
|
Mumbai
06 Aug
Sanskruti Solutions
Mumbai
- (Important Notes: (1) It is a Desk Job (Onsite Job), operating for 6 days / week working pattern
- 5 days working from an office setting and Saturday currently allowed with remote working and (2) Immediate Joiners Preferred)
Job Summary:
- We are looking for a highly skilled Data Engineer with strong programming and workflow automation skills to support and enhance our production bioinformatics pipelines.
- You do not need a biology background
- but you must be comfortable working with biological data formats, large datasets, and cloud-based workflows.
- In this role, you will develop, maintain, and optimize our data processing pipelines used for NGS (Next-Generation Sequencing) analysis, clinical workflows, and research innovation.
- You will work closely with bioinformaticians, software engineers, and data scientists to build scalable, reliable, and efficient systems.
Responsibilities:
- Pipeline & Data Engineering:
- Develop and maintain scalable data pipelines for genomic and clinical datasets.
- Build workflow automation using Python, Shell, Docker, and workflow managers (Nextflow, Snakemake, Airflow, etc.).
- Optimize existing pipelines for performance, resource usage, and reliability.
- Handle large biological datasets (FASTQ, BAM, VCF, CSV/TSV, metadata).
- Software Engineering:
- Write clean, modular, production-level code in Python and Shell.
- Implement CI/CD processes for pipeline deployment.
- Maintain code repositories (Git) and ensure high-quality documentation.
- Cloud & Infrastructure:
- Work with AWS/GCP/Azure services for scalable pipeline execution.
- Manage container-based deployments using Docker.
- Monitor job performance, logs, and system behavior.
- Data Management:
- Maintain data integrity, versioning, and audit trails.
- Develop automated QC checks and validation workflows.
- Support data ingestion, transformation, and ETL processes.
- Collaboration:
- Work with bioinformaticians to translate analysis logic into scalable workflows.
- Collaborate with clinical and operations teams to ensure pipeline readiness.
- Troubleshoot pipeline failures and optimize workflows in production.
Required Skills: Core Technical Skills:
- Strong proficiency in Python
- Hands-on experience with Shell scripting (bash)
- Strong understanding of Docker / containerization
- Knowledge of Git, CI/CD, and software development best practices
- Experience with workflow orchestration: Nextflow, Snakemake, Airflow, Cromwell, Prefect (any one) Data Engineering Skills:
- Experience with large datasets, ETL pipelines, log processing
- Strong understanding of file formats, data transformation, and automation
- Familiarity with Linux and HPC or distributed systems Bioinformatics Data Handling (Training can be provided):
- Understanding of at least one biological data type: FASTQ, BAM/CRAM, VCF, BED, GTF
- Ability to process unstructured or semi-structured scientific data Good to Have Skills (Will Prefer If Available):
- Experience with R
- Exposure to ML/AI workflows
- Experience with cloud (AWS Batch, Lambda, S3, EC2)
- Knowledge of Next Generation Sequencing (NGS) pipelines
- Familiarity with database systems (SQL/NoSQL)
📌 Bioinformatics Data Engineer ( Python / Shell Scripting / Docker ) (Mumbai)
🏢 Sanskruti Solutions
📍 Mumbai