28 Sep
|
UnderAI
|
Gurugram
About UnderAI
UnderAI is building AI-powered software for insurance brokers and enterprises. Our platform processes large documents, runs long-running AI workflows, executes background jobs, and supports critical insurance workflows such as policy audits, RFQ creation, risk analysis, and policy placement.
As our usage grows, reliability and infrastructure become critical. We are looking for an engineer who enjoys building backend systems that remain stable even when individual components fail.
What You’ll Work On
You’ll own and improve the infrastructure and backend reliability of UnderAI.
Your work will include
- Designing fault-tolerant backend systems.
- Building reliable background job and queue architectures.
- Handling retries, timeouts, idempotency, and partial failures.
- Designing systems that recover gracefully when APIs, models, databases, or external services fail.
- Improving deployment pipelines and production infrastructure.
- Building monitoring, logging, tracing, and alerting.
- Investigating production incidents and preventing recurrence.
- Designing scalable APIs and backend services.
- Improving database performance, indexing, migrations, and reliability.
- Implementing caching and rate-limiting strategies.
- Designing systems for long-running AI and document-processing workloads.
- Managing Dockerized services and cloud deployments.
- Improving CI/CD pipelines.
- Building backup and disaster-recovery strategies.
What We’re Looking For Strong understanding of backend and distributed systems fundamentals, including:
- Fault tolerance.
- Retry strategies and exponential backoff.
- Idempotency.
- Distributed systems failure modes.
- Message queues and background jobs.
- Database transactions.
- Concurrency.
- Caching.
- Rate limiting.
- Horizontal scaling.
- Load balancing.
- Observability.
- Graceful degradation.
Engineering Skills Strong experience with several of the following:
- Node.js / TypeScript.
- PostgreSQL.
- Docker.
- Linux.
- REST APIs.
- Cloud infrastructure.
- CI/CD pipelines.
- Git.
Experience with the following is a solid plus:
- Redis.
- Kafka / RabbitMQ / SQS or similar queues.
- Kubernetes.
- AWS / GCP / Azure.
- Railway / Vercel.
- Terraform or Infrastructure as Code.
- Prometheus / Grafana.
- OpenTelemetry.
- Sentry.
- PostgreSQL performance tuning.
- Distributed systems.
- High-availability architecture.
What Makes You Stand Out We are especially interested if you have previously:
- Designed systems that process long-running background jobs.
- Built retry and recovery mechanisms for unreliable external APIs.
- Designed fault-tolerant distributed systems.
- Improved uptime or reliability of a production application.
- Built observability systems using logs, metrics, and traces.
- Diagnosed difficult production incidents.
- Worked on systems where data consistency and failure recovery mattered.
We care significantly more about strong engineering fundamentals and production experience than knowledge of a specific cloud platform. Please include GitHub repositories, architecture write-ups, or examples of backend/infrastructure systems you have built.
Pay: ₹600,000.00 - ₹1,000,000.00 per year
Work Location: In person
📌 Backend / DevOps & Reliability Engineer (Gurugram)
🏢 UnderAI
📍 Gurugram