We’re hiring - Model Optimisation Engineer
We’re looking for someone who understands how to make modern models smaller, faster, cheaper and more effective without losing the capabilities that matter.
This is a highly technical, hands-on position.
What you’ll do
Fine-tune language models for specialised applications
Prepare and improve training datasets
Work with SFT, LoRA and QLoRA
Experiment with knowledge distillation
Evaluate models across different tasks
Quantise and compress models
Reduce memory requirements
Optimise inference speed and throughput
Deploy models on smaller GPU configurations
Benchmark models across different hardware
Find ways to maintain quality while reducing compute requirements
Experience we're looking for
PyTorch
Hugging Face Transformers
PEFT
LoRA / QLoRA
Model fine-tuning
Quantisation
GPTQ / AWQ / GGUF
CUDA
vLLM, SGLang or llama.cpp
GPU and inference optimisation
Model evaluation
Experience getting capable models to run on small or resource-constrained hardware is a major advantage.
The person we want
Someone who thinks:
How much capability can we get from the smallest practical model?
You should have practical experience taking models from training → evaluation → optimisation → deployment.
📌 Model Optimisation Engineer (India)
🏢 Cliead
📍 India
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.