10 Sep
|
System Soft Technologies
|
India
10 Sep
System Soft Technologies
India
Job Summary:
- We are seeking a Cloud Bioinformatics & Nextflow Engineer to design, build, and operate production-grade bioinformatics workflows for large-scale spatial multi-omics, single-cell, and high-throughput sequencing datasets. This role combines hands-on Nextflow engineering, Python software development, bioinformatics expertise, and AWS cloud engineering. The successful candidate will create scalable, reproducible, cost-aware pipelines that use AWS Batch and EC2 for elastic compute, Amazon S3 for durable object storage, and Amazon EFS for shared high-performance workflow storage. The engineer will partner with scientists, data engineers, platform teams, and security stakeholders to convert complex analytical requirements into reliable cloud-native solutions.
Responsibilities:
- Nextflow Engineering: Architect, develop, test, and maintain modular Nextflow DSL2 pipelines and reusable components for large-scale spatial transcriptomics, scRNA-Seq, multi-omics, and other sequencing workloads.
- Large-Scale Data Processing: Optimize workflow parallelism, task partitioning, retries, checkpointing, caching, and resource allocation to process terabyte-scale datasets efficiently and reliably.
- AWS Cloud Engineering: Design and operate Nextflow execution patterns on AWS Batch and EC2, including compute environments, job queues, launch templates, instance selection, autoscaling, Spot and On-Demand capacity strategies, and failure recovery.
- Cloud Storage Architecture: Engineer secure and performant data movement and storage patterns across Amazon S3 and Amazon EFS, including staging, lifecycle management, throughput optimization, metadata handling, and management of intermediate workflow files.
- Python Engineering: Develop production-quality Python packages, command-line tools, APIs, validation utilities, and automation used by bioinformatics workflows. Apply unit testing, type hints, structured logging, packaging, and code-quality standards.
- Spatial Transcriptomics Expertise: Develop and optimize analysis workflows for spatial transcriptomics platforms and datasets, including image and molecular data integration, spatial quality control, segmentation, expression quantification, normalization, spatial clustering, cell-type annotation, neighbourhood analysis, and biological interpretation.
- Bioinformatics Integration: Integrate and validate established bioinformatics tools and algorithms for quality control, alignment, quantification,
normalization, cell-level analysis, spatial analysis, and downstream interpretation.
- POC to Production: Receives validated proof-of-concept workflows and modules from Scientific Workflow Developers and is responsible for production deployment, scalability, cloud optimization, operationalization, CI/CD, and long-term maintainability.
- Containers and Reproducibility: Create and maintain Docker or Apptainer/Singularity images with pinned dependencies and reproducible runtime environments. Integrate images with appropriate container registries.
- Performance and Cost Optimization: Profile CPU, memory, disk, network, and I/O usage; tune Nextflow process directives; select appropriate EC2 instance families; and reduce cloud cost without compromising scientific quality or reliability.
- Software Delivery: Implement CI/CD, automated testing, code review, semantic versioning, release management, Git-based development, and infrastructure-as-code practices for workflow and platform components.
- Observability and Operations: Implement logs, metrics, alerts, auditability, run metadata, provenance, and operational dashboards. Troubleshoot failed jobs, storage bottlenecks, dependency issues, and distributed workflow behavior.
- Security and Data Governance: Apply IAM least-privilege access, encryption, network controls, secrets management, data retention, and controlled access patterns appropriate for sensitive scientific data.
- Collaboration: Work closely with computational biologists, bench scientists, data scientists, and cloud platform teams to define requirements, document solutions, support users, and communicate technical tradeoffs.
Experience:
- Bachelor's or master's degree or PhD in Bioinformatics, Computational Biology, Computer Science, Software Engineering, Data Science, or a related discipline, with 3+ years of relevant experience.
- Robust hands-on experience engineering and operating Nextflow pipelines, preferably using DSL2, modules, subworkflows, profiles, configuration management, and cloud executors.
- Demonstrated experience scaling Nextflow or comparable workflow platforms for large datasets and highly parallel workloads.
- Advanced proficiency in Python for bioinformatics automation and production software development, including testing, packaging, error handling, and performance-conscious data processing.
- Practical AWS experience with AWS Batch, EC2, S3, and EFS, including compute, storage, permissions, networking, monitoring, and cost considerations.
- Experience with bioinformatics and high-throughput sequencing data, such as scRNA-Seq, genomics, transcriptomics, or multi-omics.
- Experience with Linux, shell scripting, Git, GitHub or GitLab, and containerized execution using Docker and/or Apptainer/Singularity.
- Understanding reproducibility, data provenance, workflow testing, scientific validation, and software engineering practices.
- Strong troubleshooting, documentation, communication, and cross-functional collaboration skills.
Nice to Have:
- PhD in Bioinformatics, Computational Biology, Computer Science, or related discipline.
- Deep experience with Nextflow on AWS, including AWS Batch queue design, EC2 instance and storage optimization, and large-scale workflow orchestration.
- Demonstrated hands-on experience analyzing spatial transcriptomics data, including spatial quality control, image integration, expression quantification, normalization, cell-type annotation, spatial clustering, and interpretation of tissue-level patterns.
- Experience with Nextflow Tower/Seqera Platform, nf-core development practices, or reusable community pipeline patterns.
- Experience with infrastructure as code and cloud automation, such as Terraform, AWS CloudFormation, or AWS CDK.
- Experience with orchestration, CI/CD, and observability tools, including GitHub Actions, GitLab CI, CloudWatch, or equivalent platforms.
- Experience with spatial transcriptomics technologies and analysis ecosystems, such as 10x Genomics Visium, Xenium, NanoString CosMx, Vizgen MERSCOPE, or equivalent platforms.
- Knowledge of common analysis frameworks and tools such as Scanpy, Squidpy, Seurat, Bioconductor, STAR, Salmon, Cell Ranger, or other spatial-analysis toolsets.
- Experience in immunology, oncology, or another biological domain relevant to translational research.
- Familiarity with regulated or enterprise environments, data governance, security controls, and validated scientific software practices.
📌 Cloud Bioinformatics & Nextflow Engineer (India)
🏢 System Soft Technologies
📍 India