13 Aug
|
Zeqo Systems
|
India
13 Aug
Zeqo Systems
India
We're looking for someone genuinely solid on
Google Cloud Platform
— not necessarily years of work experience, but real depth: you've built things on GCP, you understand its IAM, quota, and networking model cold, and you can reason about Cloud Run, Vertex AI, and multi-project setups from first principles. You'll design and build Zeqo's
multi-provider LLM router
— the internal system that decides which model/provider handles each request (Google Gemini, OpenAI, Anthropic, and self-hosted open-weight models), optimizes for cost/latency/quality, and aggressively reduces redundant token spend through prompt caching and smart request management. This is a foundational infrastructure role: get it right and every downstream product (MockLab, Mophy, OCR grading) gets faster and cheaper automatically.
This is intentionally open to early-career engineers — we care far more about depth of GCP knowledge and evidence of things you've actually built than a resume with years on it.
*What you will do:*
- Design and build a unified LLM router
that abstracts Google, OpenAI, Anthropic,
and self-hosted model endpoints behind a single internal API, so product teams never hardcode a provider.
- Implement intelligent routing logic
— route by task type, cost ceiling, latency budget, context length, or model capability (e.g., OCR evaluation vs. conversational tutoring vs. embedding generation).
- Build fallback and failover chains
so a rate limit, timeout, or outage on one provider automatically retries on another without breaking the user experience.
- Design prompt/response caching
(semantic and exact-match) to cut redundant token usage and reduce rate-limit pressure — particularly for repeated question-bank queries, common doubt patterns in Mophy, and OCR grading rubrics.
- Own rate-limit and quota management
across providers and GCP projects, including request queuing, backpressure, and graceful degradation under load.
- Build cost observability
: per-request cost attribution, per-p
📌 Cloud Infrastructure Engineer — LLM Routing (India)
🏢 Zeqo Systems
📍 India