AI Forward - Relability & Platform Engineer / SRE (Fully Remote) (India)

AI Forward - Relability & Platform Engineer / SRE (Fully Remote) (India)

22 Aug
|
Sceneplay
|
India

22 Aug

Sceneplay

India

We are hiring a senior scalability engineer to help prepare a modern microservices backend for high-concurrency production usage. This role is focused on backend performance, reliability, capacity planning, async workflow design, database scalability, and observability.

This is not a generic DevOps-only role. We are looking for someone who can reason deeply about how backend systems behave under load: database pressure, queue depth, retries, rate limits, stuck jobs, async workflows, API latency, and failure recovery.

Responsibilities

- Analyze backend architecture and identify scalability bottlenecks across APIs, services, database, cache, queues, storage, and third-party integrations.
- Design and execute load tests for high-concurrency user flows.
- Tune Postgres queries, indexes, connection pools, pagination, and transaction patterns.
- Design reliable queue/background-job systems with retries, idempotency, dead-letter queues, and reconciliation jobs.
- Define observability for scale: metrics, dashboards, request IDs, error tracking, alerts, and stuck-state monitoring.
- Improve reliability of webhook-driven and async processing workflows.
- Partner with backend engineers to implement code-level and architecture-level scalability improvements.
- Recommend infrastructure and autoscaling changes based on measured bottlenecks.

Required Skills

- 4+ years of backend, platform, SRE, performance engineering, or scalability engineering experience.
- In-depth system design knowledge is of paramount importance.




- Robust understanding of high-concurrency backend systems and distributed systems failure modes.
- Deep Postgres experience: query plans, indexing, connection pools, slow queries, transaction design, and pagination.
- Experience with Kafka, Redis, queues, background workers, retry/backoff, DLQs, and idempotent processing.
- Hands-on load testing experience with k6, Artillery, Locust, JMeter, or similar.
- Strong observability experience with logs, metrics, dashboards, alerting, and error tracking.
- Cloud/container familiarity with AWS, Docker, ECS/Fargate, Kubernetes, or similar.
- Ability to work with backend application code and collaborate with engineering teams.

Nice To Have

- Production Node.js or NestJS experience.
- TypeORM or ORM performance tuning experience.
- DevOps/platform experience deploying and operating applications at scale.
- Experience with object storage, upload-heavy systems, media processing, or webhook-heavy workflows.
- OpenTelemetry or distributed tracing experience.
- Terraform, Pulumi, AWS CDK, or infrastructure-as-code experience.

What Success Looks Like

- We know what breaks first at 10x and 100x traffic.
- Critical backend flows have load-test baselines and dashboards.
- Database and queue bottlenecks are identified and mitigated.
- Async workflows are retryable, idempotent, and observable.
- The team has clear production scalability priorities before launch.

Employment Details

- Role: Senior Scalability / Platform Engineer
- Location: Remote only
- Seniority: Senior / Lead-level preferred

📌 AI Forward - Relability & Platform Engineer / SRE (Fully Remote) (India)
🏢 Sceneplay
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: ai forward - relability & platform engineer / sre (fully remote) (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: ai forward - relability & platform engineer / sre (fully remote) (india) / india