We’re a founder-led AI company building an autonomous growth platform that decides, for every customer, which offer, message, channel and timing will work best, and learns from the results. We’re looking for a hands-on senior engineer who has built these decision systems in production.
Location: Remote (3–4 hrs daily overlap with US Pacific time)
Experience: 6+ years of software or ML engineering
Reports to: Founder & CEO
You must have:
• Built and run a production system that chooses between options for each user and learns from outcomes, using multi-armed bandits or reinforcement learning on live traffic (e.g., recommendations, ranking, ad or offer selection, pricing, or send-time/channel optimization)
• 3+ years shipping LLM features or AI agents to real users (not proofs of concept)
• 6+ years of hands-on engineering, excluding internships, teaching and study periods
• Writing code most of the week today, with experience as the technical lead on a product
• Robust Python (TypeScript is a plus)
This role is not a fit if your RL experience is only:
RLHF/DPO fine-tuning, A/B testing of prompts or models, building RL environments or training data for AI labs, or coursework/research projects.
You’ll likely be a great fit if you’ve worked as:
Applied Scientist, ML Engineer (Personalization, Recommendations, Ranking, Ads) or Decision Scientist at a consumer product company in e-commerce, food delivery, ride-hailing, fintech, gaming, adtech or martech.
What you’ll do:
• Own the bandit/RL decision loops at the core of our platform, including reward design and evaluation
• Build and ship AI agents that plan and optimize campaigns
• Write most of the critical code yourself while leading architecture
Why join us: Small team, direct work with our CEO, and real ownership of the product.
📌 Senior AI/ML Engineer – Personalization, Bandits (India)
🏢 OWOW
📍 India