Developing a software pipeline for end -to -end ML Model Inference for specific hardware accelerator by achieving maximum performance & accuracy.
• Implementing cutting edge deep learning layers for various model categories like CNN, RNN, LSTM, GANs, etc using customized inference pipeline for NN Processor.
• Performance optimization for inferencing the LLM Models in customized hardware with various layer types including transformer, encoder -decoder, etc based models.
• Hardware architecture aware and computation conscious implementation of solutions in an embedded device and maximize the throughput.
• Develop tools and applications by producing clean, effective code.
• Identify, prioritise and execute tasks based on requirement.