APPLY
If you believe you are a strong fit, please complete this application form:
https://docs.google.com/forms/d/e/1FAIpQLSfL7jjqt4DQdm9g948j7wnx6VS1x62v-kNhc4Eb-a8EG6UvdA/viewform - this will ensure your application gets a fair look (rather than just applying in Easy Apply)
ABOUT US
Strontium is one of the fastest-growing AI companies in Canada—fully bootstrapped, with 15+ years in the industry and 100,000+ customers across 60+ countries. We build and sell AI-powered creative tools directly to consumers.
Two of our products
- VideoExpress.ai—our video-generation product, rated 4.8★ on Capterra with 6,000+ reviews and ranked #1 in its category:
https://www.capterra.com/p/10019949/VideoExpress/reviews/
- Artistly.ai—our image-generation product, rated 4.8★ on Trustpilot with 1,000+ reviews:
https://ca.trustpilot.com/review/artistly.ai
Strontium and IntellifAI Labs are collaborating to bring state-of-the-art AI technologies to everyday users. We ship rapid, iterate constantly, and compete with companies 100× our size.
THE ROLE — RESEARCH ENGINEER, INFERENCE OPTIMIZATION
We run image, video, and audio generation models in production at consumer scale. Every second we remove from a generation improves the product and lowers our GPU costs.
Your job is to make our models dramatically faster without sacrificing output quality. This is a research role with a production target.
WHAT YOU’LL DO
- READ
Stay current with the literature on generative-model acceleration. Most published techniques will not work for our models; your job is to identify the 10% that might.
- IMPLEMENT
Take promising ideas from papers and get them running on our models quickly. You will have unlimited access to Claude Code and Codex,
and we expect you to use them aggressively.
- MEASURE
Build and defend benchmarks covering end-to-end latency, throughput under real traffic, cost per generation, and quality regressions. A speedup that quietly degrades output is not a speedup.
- SHIP
Deploy successful optimizations to our own GPU fleet and own them afterward.
OPTIMIZATION AREAS
- Model level: step distillation, few-step samplers, caching and feature reuse, quantization, pruning, and architecture surgery. (40%)
- Kernel level: CUDA/Triton kernels, fused attention, torch.compile, and TensorRT. (40%)
- Serving level: batching, scheduling, and autoscaling. (20%)
ROLE DETAILS
- Remote
- Flexible hours with reasonable overlap with Eastern Time
- Competitive compensation
MINIMUM QUALIFICATIONS
- Bachelor’s or Master’s degree in Computer Science, Electronics, Electrical Engineering, or a related discipline, with a strong coding background.
- Hands-on experience training or running inference for image, video, or audio generation models in PyTorch.
- Strong PyTorch fundamentals: you can profile a model, read a trace, and distinguish a kernel-bound bottleneck from a memory-bound one.
NICE TO HAVE (GENUINELY OPTIONAL)
- Publications or open-source contributions in efficient inference.
- CUDA or Triton kernel-authoring experience.
- Experience operating serving stacks such as vLLM, ComfyUI, or custom schedulers under real traffic.
If you have spent the last year reading papers because you could not put them down, and keep thinking about these ideas all the time this is your job—whether that work happened during a Master’s program, an internship, or on nights and weekends.
Questions? Contact
[email protected].
📌 AI Researcher, Inference Optimization (India)
🏢 IntellifAI Labs
📍 India