Site Reliability Engineer (Bengaluru)

Site Reliability Engineer (Bengaluru)

04 Sep
|
Otto AI
|
Bengaluru

04 Sep

Otto AI

Bengaluru

AWS · TypeScript · Reliability · Security · Scale

Reliability, scale, security, and cost are all yours.Otto builds an AI computer.

Our hardware sits on a customer's desk running embedded Linux, maintains a live connection to our cloud, and receives signed over-the-air updates from us.

Every part of that path runs on AWS. Our infrastructure is infrastructure-as-code, and our stack is heavily TypeScript.

You will design and operate the platform behind Otto: cloud infrastructure, CI/CD, observability, fleet reliability, security, capacity planning, and cost.

You will also write code, review PRs, deploy to production, respond to incidents, and participate in on-call.

This is a small team. There is no infrastructure layer between you and the system.

Infrastructure as Code

- AWS CDK v2
- TypeScript
- VPC
- IAM

Compute

- AWS Lambda
- ECS Fargate
- API Gateway
- CloudFront

Data

- RDS PostgreSQL
- DynamoDB
- ElastiCache / Redis
- S3

Delivery

- GitHub Actions
- Docker
- ECR
- Canary, rolling, and blue/green deployments
- Automated rollback

Observability

- CloudWatch
- OpenTelemetry
- Grafana
- SLIs / SLOs
- Alerting and incident response

Nearby Stack

- Node.js
- Next.js
- Vercel

Some of our web surfaces run on Vercel. Experience with it is helpful, but not required. What You'll OwnProduction Infrastructure on AWSDesign, build, and operate our production infrastructure across Lambda, Fargate, RDS, ElastiCache, DynamoDB, S3, CloudFront, API Gateway, VPC, and IAM.

You should be comfortable owning the system end to end rather than relying on a separate platform team.

Infrastructure as CodeEvery resource lives in our CDK tree.

Development and production environments should be reproducible from the same code, with infrastructure changes reviewed through the same PR process as application code.

CI/CDBuild and maintain GitHub Actions pipelines for:

- Application deployments
- Container deployments
- Infrastructure deployments
- Database migrations
- Canary releases
- Rolling deployments
- Blue/green deployments
- Automated rollback

You should know when each deployment strategy is appropriate and why. ObservabilityOwn metrics, logs, traces, dashboards, and alerts across the platform.

The goal is straightforward: page a human only when a human is actually needed.

SLIs, SLOs, and Error BudgetsDefine reliability targets, track them, and help the product and engineering teams make informed decisions when reliability budgets are being consumed.

Incident ResponseLead incidents when they happen.

Run blameless postmortems and turn findings into real reliability work rather than documentation that gets forgotten.





OTA Updates to the Otto FleetOur releases eventually reach physical devices sitting in customers' homes and offices.

You will help own

- Canary rollouts
- Rollback detection
- Fleet health monitoring
- Release safety
- Failure recovery

A bad release can reach hardware we cannot physically access, so deployment discipline matters. AWS CostOwn infrastructure efficiency and visibility.

This includes

- Rightsizing
- Spot and reserved capacity
- Tagging discipline
- Budget alerts
- Cost attribution
- Identifying which services and workloads are actually driving spend

SecurityHelp establish and maintain a strong production security posture, including:
- IAM least privilege
- Network segmentation
- Secrets management
- Patching
- Access controls
- Auditability
- Foundations for SOC 2

Data InfrastructureOperate PostgreSQL, Redis, and DynamoDB in production.

That includes

- Backups
- Tested restores
- Capacity planning
- Upgrades
- Performance
- Availability
- Failure recovery

Killing ToilIf something repetitive can be automated, automate it. Build internal tooling in TypeScript or Bash that allows engineers to self-serve instead of relying on manual infrastructure work.

MentoringHelp raise the operational and reliability bar across the entire engineering team.

Must HaveThese are not "familiarity with" requirements.

We are looking for someone who has owned these systems in production and understands where they fail.

AWS CDKDeep, current experience with AWS CDK v2 in TypeScript, including:

- Constructs
- Stacks
- Cross-stack references
- Cross-region architecture
- Custom resources
- Deployment troubleshooting
- Managing infrastructure changes safely in production

You should know what to do when the synth is clean but the deployment still isn't. AWS in ProductionStrong hands-on experience with:
- VPC architecture and networking
- IAM policy design
- CloudFront
- API Gateway
- S3
- Production security and access controls

We are looking for direct ownership, not experience where another platform team handled the difficult parts.

AWS LambdaYou should understand

- Cold starts
- Concurrency limits
- VPC-attached functions
- Bundle size
- Scaling behavior
- Timeouts
- Connection management
- The problems Lambda creates when talking to PostgreSQL

RDS PostgreSQLExperience actually operating PostgreSQL in production, including:
- Query plans
- Indexing
- Vacuum behavior
- Connection limits
- RDS Proxy




- Major-version upgrades
- Point-in-time recovery
- Backup and restore procedures you have actually tested

DynamoDBStrong understanding of DynamoDB data modeling, including:
- Designing around access patterns
- Partition keys
- GSIs
- Hot partitions
- Conditional writes
- On-demand vs. provisioned capacity

TypeScript and Node.jsYou should write real production TypeScript, not just infrastructure glue. Our infrastructure, backend, internal tooling, and device runtime are all heavily Node.js and TypeScript.

You will regularly read and review code outside of a traditional infrastructure role.

Linux and NetworkingStrong understanding of:

- TCP/IP
- DNS
- TLS
- Load balancing
- WebSockets
- Linux systems

You should be comfortable debugging problems like a persistent socket that only drops for customers on one ISP. CI/CD and Incident ResponseYou should have experience:
- Building deployment pipelines
- Choosing deployment strategies
- Running production incidents
- Writing postmortems
- Establishing SLOs
- Turning incidents into reliability improvements

BonusThese are not required, but they count for a lot.
- Next.js, including App Router, Server Components, and build pipelines
- Vercel, including projects, preview environments, edge configuration, and hosted infrastructure
- Edge, IoT, or embedded Linux fleets with OTA updates
- Redis pub/sub
- Large-scale WebSocket services
- AWS Organizations and multi-account architectures
- OIDC deployment roles
- Ephemeral per-developer environments
- LLM gateways, proxies, or AI workloads on AWS
- SOC 2 or ISO 27001 experience
- FinOps experience
- Self-hosted CI runners
- Container build caching
- Kubernetes experience — useful context, although we do not currently run Kubernetes
- AWS certifications such as Solutions Architect, DevOps Engineer, or SysOps Administrator

What This Is Really LikeOtto is a small team with a short path from decision to production. You will work in the same repositories as the rest of the engineering organization and regularly read code outside your lane.

Our systems span cloud infrastructure, APIs, data infrastructure, web applications, and physical devices running in the field.

We value clear technical writing.

Design documents, postmortems, architecture decisions, and the sentence in a PR explaining why something is changing all matter.

If a deployment changes something persisted on an Otto device that has been running in someone's home for six months, we want that risk understood and documented before the release goes out.

On-call is shared and real.

In return, fixing the system that woke you up is treated as the work — not as an interruption from the work.

📌 Site Reliability Engineer (Bengaluru)
🏢 Otto AI
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (bengaluru) / bengaluru