MLOps Engineer (Bengaluru)

MLOps Engineer (Bengaluru)

04 Sep
|
Antrino Labs
|
Bengaluru

04 Sep

Antrino Labs

Bengaluru

We are looking for an excellent MLOps Engineer to join one of Sweden's fastest growing start-ups which is backed by Nvidia, Microsoft and AWS. The company also operates under the Inspection for Strategic Products.

Location: Fully remote within India. From 2027 you join our Bangalore office as part of the founding local team

Employment: Full time, permanent

Start: As soon as you are available

About the job

Antrino Labs is building the platform that makes every square meter intelligent. The physical world runs on processes nobody can actually see. Goods move, people queue, machines idle, space goes unused, and the decisions made about all of it are based on samples, guesses, and reports written after the fact. Software solved this for the digital world twenty years ago. Everything online is measured, understood, and acted on in real time. The physical world is still dark.

We are building the execution layer that closes that gap. Antrino lets any organization deploy vision intelligence into a physical space and get back what is actually happening there, as structured, queryable, real-time information rather than footage. Not a research project, not a custom integration, not a team of ML engineers. A platform, where you describe what matters to you and deploy it.

What people build on it is broader than what we designed for. Operators use it to understand processes, find where time and space are being wasted, measure flow and utilization, catch problems while they are still happening, and turn all of it into data good enough to act on.

Everything is built on Privacy by Design. Data protection and GDPR compliance are part of the architecture, not a layer added afterwards.

We recently closed our Seed round, announced together with Dagens Industri, and we are scaling to meet demand. Antrino Labs is also registered with Inspektionen för strategiska produkter (ISP), the Swedish authority for strategic and dual-use products.

The role

We are hiring an experienced MLOps Engineer to own how models actually run in production.

Read that carefully, because it defines the job. We are not a training shop. We do not fine-tune our way out of problems, and we are not asking you to build a training platform nobody needs yet. Our intelligence comes from composing models, our own detectors and vision-language models alongside frontier APIs — and routing work between them. That makes inference operations the discipline that decides whether this company has a margin, and it is currently the least built-out part of our stack.

Concretely, the platform runs continuous inference pipelines against live streams. Every one of them consumes GPU time or API spend every second it is deployed, across a growing number of sites. The core economic fact of our business is that frame rate and duty cycle drive cost, and the affordable architecture is gated: a cheap model decides when an expensive one is allowed to look. Making that gating reliable, measurable, and automatic across a fleet is the work.

The second half of the job is knowing whether any of it is right.



We deploy into environments where we cannot simply collect and keep data to label, privacy is architectural here, not negotiable, so evaluation has to be built deliberately: curated and synthetically generated scenario sets, regression suites per capability, and honest precision and recall numbers we can show a customer. An operator who stops trusting the output has churned, whatever the contract says.

This is a remote-first role today, working closely with the engineering team in Stockholm. From 2027 you will be part of the founding group in our Bangalore office. If you want to be an early face of a team as it forms rather than an employee number in one that already exists, that is the opportunity.

Your main areas of responsibility will include:

- Owning the inference platform: how models are served, scheduled, batched, scaled, and retired across live and archive workloads
- Building the routing and gating layer — cheap detectors gating expensive vision-language calls, with policy that adapts to load, latency budgets, and cost caps rather than being hand-tuned per deployment
- Driving down cost per stream-hour and latency to result, and making both observable per pipeline, per project, and per customer
- Building the evaluation stack: scenario and regression suites per capability, synthetic scenario generation, and clear precision and recall reporting that survives contact with a customer conversation
- Setting up the model registry, versioning, staged rollout, shadow deployment, and rollback so a model change is a routine operation rather than an event
- Monitoring for drift, degradation, and silent failure — a pipeline that quietly stops producing results is worse than one that crashes, and catching that class of failure is on you
- Managing our use of frontier model APIs: quotas, rate limits, fallbacks, cost controls, and evaluating new models as they ship
- Getting models running within the constraints of on-site compute, where GPU, memory, thermal budget, and network are all limited, and shipping updates to that hardware safely
- Building the GPU and inference infrastructure on Azure, with sane autoscaling and capacity planning
- Working with the product engineers so that model output arrives in the product as something an operator can act on, not as a raw score
- Writing the tooling and internal docs that let the rest of the team ship model changes without going through you

Our Stack

You do not need experience with all of this, but you should recognize most of it and be able to argue about it.

- Serving and inference: Python, FastAPI, PyTorch, ONNX/TensorRT, GPU inference, batching and duty-cycle scheduling
- Models: first-party detectors and vision-language models, alongside Anthropic and Google model APIs




- Streaming: RTSP ingestion, RTSP→HLS conversion, live and archive pipelines, with the architecture built to take more than one class of input
- Infrastructure: Azure, GPU VMs and containers, object storage, GitHub Actions, in office compute reached over Tailscale
- Data and product: PostgreSQL (Supabase, with row-level security), TypeScript/Next.js product surfaces consuming your outputs
- Evaluation: curated scenario sets plus our own synthetic scenario generation pipeline, used where real footage cannot be retained

Who we are looking for

- Substantial experience running machine learning in production, models you owned operationally, not models you handed to someone else to deploy
- Strong Python and real depth in inference serving: GPU utilization, batching, quantization, latency budgets, and where the milliseconds and the rupees actually go
- Experience with computer vision or video workloads, or with any high-throughput real-time inference system where cost per unit of work mattered
- You have built evaluation that people trusted — datasets, regression suites, metrics that meant something — and you are opinionated about what a good number looks like
- Comfort with cloud infrastructure, containers, CI/CD, and infrastructure-as-code. Azure experience is a plus; the reasoning transfers
- Experience with model registries, versioning, staged rollout, and monitoring in production
- Deploying to constrained or edge hardware is a strong advantage
- You are pragmatic about what to build. Early-stage MLOps is choosing the three things that matter and refusing the other twenty
- You work well asynchronously across time zones, write clearly, and raise problems early
- Excellent written and spoken English

What we offer

- Competitive salary, benchmarked to the top of the Indian market for this level and discussed openly early in the process
- Fully remote within India, with genuine ownership rather than delegated tickets, and a seat in the founding group of our Bangalore office from 2027
- Hardware of your choice and a budget for whatever else you need to work well
- Direct access to founders, to the Stockholm engineering team, and to customers
- Travel to Stockholm to work with the team in person

Process

We review applications continuously and contact candidates we think are a good match.

1. An introductory conversation about your background and the role
2. A technical conversation, we discuss systems you have run in production, what broke, and what you did about it. No whiteboard algorithm puzzles
3. A working session on a real problem from our stack, close to what you would actually be doing
4. A conversation with the founders about ownership, options, and terms

References may be requested in the final stage. We aim to keep the process short and to respect your time.

Apply with your CV and, if you have one, a link to something you have built or written. A repository, a post about a system you operated, or a short note on a production failure you learned from tells us more than a cover letter.

📌 MLOps Engineer (Bengaluru)
🏢 Antrino Labs
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: mlops engineer (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: mlops engineer (bengaluru) / bengaluru