Cloud Bioinformatics & Nextflow Engineer (India)

Cloud Bioinformatics & Nextflow Engineer (India)

10 Sep
|
System Soft Technologies
|
India

10 Sep

System Soft Technologies

India

Job Summary:

- We are seeking a Cloud Bioinformatics & Nextflow Engineer to design, build, and operate production-grade bioinformatics workflows for large-scale spatial multi-omics, single-cell, and high-throughput sequencing datasets. This role combines hands-on Nextflow engineering, Python software development, bioinformatics expertise, and AWS cloud engineering. The successful candidate will create scalable, reproducible, cost-aware pipelines that use AWS Batch and EC2 for elastic compute, Amazon S3 for durable object storage, and Amazon EFS for shared high-performance workflow storage. The engineer will partner with scientists, data engineers, platform teams, and security stakeholders to convert complex analytical requirements into reliable cloud-native solutions.

Responsibilities:

- Nextflow Engineering: Architect, develop, test, and maintain modular Nextflow DSL2 pipelines and reusable components for large-scale spatial transcriptomics, scRNA-Seq, multi-omics, and other sequencing workloads.
- Large-Scale Data Processing: Optimize workflow parallelism, task partitioning, retries, checkpointing, caching, and resource allocation to process terabyte-scale datasets efficiently and reliably.
- AWS Cloud Engineering: Design and operate Nextflow execution patterns on AWS Batch and EC2, including compute environments, job queues, launch templates, instance selection, autoscaling, Spot and On-Demand capacity strategies, and failure recovery.
- Cloud Storage Architecture: Engineer secure and performant data movement and storage patterns across Amazon S3 and Amazon EFS, including staging, lifecycle management, throughput optimization, metadata handling, and management of intermediate workflow files.
- Python Engineering: Develop production-quality Python packages, command-line tools, APIs, validation utilities, and automation used by bioinformatics workflows. Apply unit testing, type hints, structured logging, packaging, and code-quality standards.
- Spatial Transcriptomics Expertise: Develop and optimize analysis workflows for spatial transcriptomics platforms and datasets, including image and molecular data integration, spatial quality control, segmentation, expression quantification, normalization, spatial clustering, cell-type annotation, neighbourhood analysis, and biological interpretation.
- Bioinformatics Integration: Integrate and validate established bioinformatics tools and algorithms for quality control, alignment, quantification,



normalization, cell-level analysis, spatial analysis, and downstream interpretation.
- POC to Production: Receives validated proof-of-concept workflows and modules from Scientific Workflow Developers and is responsible for production deployment, scalability, cloud optimization, operationalization, CI/CD, and long-term maintainability.
- Containers and Reproducibility: Create and maintain Docker or Apptainer/Singularity images with pinned dependencies and reproducible runtime environments. Integrate images with appropriate container registries.
- Performance and Cost Optimization: Profile CPU, memory, disk, network, and I/O usage; tune Nextflow process directives; select appropriate EC2 instance families; and reduce cloud cost without compromising scientific quality or reliability.
- Software Delivery: Implement CI/CD, automated testing, code review, semantic versioning, release management, Git-based development, and infrastructure-as-code practices for workflow and platform components.
- Observability and Operations: Implement logs, metrics, alerts, auditability, run metadata, provenance, and operational dashboards. Troubleshoot failed jobs, storage bottlenecks, dependency issues, and distributed workflow behavior.
- Security and Data Governance: Apply IAM least-privilege access, encryption, network controls, secrets management, data retention, and controlled access patterns appropriate for sensitive scientific data.
- Collaboration: Work closely with computational biologists, bench scientists, data scientists, and cloud platform teams to define requirements, document solutions, support users, and communicate technical tradeoffs.

Experience:

- Bachelor's or master's degree or PhD in Bioinformatics, Computational Biology, Computer Science, Software Engineering, Data Science, or a related discipline, with 3+ years of relevant experience.
- Robust hands-on experience engineering and operating Nextflow pipelines, preferably using DSL2, modules, subworkflows, profiles, configuration management, and cloud executors.




- Demonstrated experience scaling Nextflow or comparable workflow platforms for large datasets and highly parallel workloads.
- Advanced proficiency in Python for bioinformatics automation and production software development, including testing, packaging, error handling, and performance-conscious data processing.
- Practical AWS experience with AWS Batch, EC2, S3, and EFS, including compute, storage, permissions, networking, monitoring, and cost considerations.
- Experience with bioinformatics and high-throughput sequencing data, such as scRNA-Seq, genomics, transcriptomics, or multi-omics.
- Experience with Linux, shell scripting, Git, GitHub or GitLab, and containerized execution using Docker and/or Apptainer/Singularity.
- Understanding reproducibility, data provenance, workflow testing, scientific validation, and software engineering practices.
- Strong troubleshooting, documentation, communication, and cross-functional collaboration skills.

Nice to Have:

- PhD in Bioinformatics, Computational Biology, Computer Science, or related discipline.
- Deep experience with Nextflow on AWS, including AWS Batch queue design, EC2 instance and storage optimization, and large-scale workflow orchestration.
- Demonstrated hands-on experience analyzing spatial transcriptomics data, including spatial quality control, image integration, expression quantification, normalization, cell-type annotation, spatial clustering, and interpretation of tissue-level patterns.
- Experience with Nextflow Tower/Seqera Platform, nf-core development practices, or reusable community pipeline patterns.
- Experience with infrastructure as code and cloud automation, such as Terraform, AWS CloudFormation, or AWS CDK.
- Experience with orchestration, CI/CD, and observability tools, including GitHub Actions, GitLab CI, CloudWatch, or equivalent platforms.
- Experience with spatial transcriptomics technologies and analysis ecosystems, such as 10x Genomics Visium, Xenium, NanoString CosMx, Vizgen MERSCOPE, or equivalent platforms.
- Knowledge of common analysis frameworks and tools such as Scanpy, Squidpy, Seurat, Bioconductor, STAR, Salmon, Cell Ranger, or other spatial-analysis toolsets.
- Experience in immunology, oncology, or another biological domain relevant to translational research.
- Familiarity with regulated or enterprise environments, data governance, security controls, and validated scientific software practices.

📌 Cloud Bioinformatics & Nextflow Engineer (India)
🏢 System Soft Technologies
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: cloud bioinformatics & nextflow engineer (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: cloud bioinformatics & nextflow engineer (india) / india