DevOps / Platform & Reliability Engineer (Chennai)

DevOps / Platform & Reliability Engineer (Chennai)

01 Oct
|
Nenu AI
|
Chennai

01 Oct

Nenu AI

Chennai

DevOps / Platform & Reliability Engineer****About Nenu

Nenu is a fast-moving, Silicon Valley-based, VC-backed startup building ambitious, user-facing products at the intersection of modern web applications and AI-powered systems.

From the heart of Chennai, we're building for global scale, focused on speed, ownership, and creating high-impact products from the ground up.

We're at a stage where the foundations are being built and scaled, which means the people who join now will directly shape how the product, platform, and engineering practices evolve.

We are an AI-first company, using AI not just in our products, but across how we build, test, deploy, and operate.

About the Role

We're looking for a DevOps / Platform & Reliability Engineer with 6–10 years of experience who thinks beyond traditional infrastructure administration and wants to define how software is built, shipped, and run in an AI-first engineering organization.

This role goes deep. You'll build on what already exists and take our infrastructure, delivery, and reliability practice to the level our scale now demands, shaped by how our actual product, systems, and engineering workflows work. You'll use AI throughout the lifecycle to automate, diagnose, validate, and continuously improve how we build and operate software.

In straightforward terms, you are the person who makes sure what we build stays up, ships fast, stays secure, and doesn't cost more than it should.

Today our AI workloads run on managed infrastructure. As we scale we expect to bring more of that capacity in-house, so experience running GPU infrastructure is a real advantage here even though it is not day-one work.

Core Competencies: What You Will Do

- Own the platform end-to-end: Take responsibility for infrastructure, delivery, and production operations across the whole system, not just individual servers, pipelines, or tickets.

- Harden and scale the platform: Evolve our infrastructure into something reproducible and scalable, designed around our actual product, systems, workflows, and failure modes rather than a generic reference architecture.

- Own reliability: Define service level objectives with engineering and product, instrument the systems to measure them, use error budgets to decide when to ship and when to fix, and build the guardrails that catch failure modes before customers do.

- Own production incidents: Lead incidents to resolution, run blameless postmortems, build the on-call rotation and runbooks, and make sure the actions that come out of an incident actually get closed.

- Own the delivery pipeline: Design CI/CD that is fast, safe, and trusted, with automated quality and security gates, environment promotion, progressive rollout, health checks after release, and a rollback that works under pressure.

- Make systems observable: Own metrics, logs, traces, and dashboards so any engineer can answer what broke and where without guessing. An alert nobody acts on is a bug.

- Own infrastructure security: Secrets management, least-privilege access and identity, network boundaries, image and dependency hygiene, and scanning built into the pipeline rather than bolted on afterwards.

- Own cloud cost: Give the company visibility into what each environment and service costs, rightsize continuously, and treat cost as an engineering design constraint rather than a finance report.

- AI-driven operations: Use AI throughout the infrastructure lifecycle,



from writing and reviewing infrastructure code to diagnosing failures, analyzing incidents, spotting configuration drift, and automating operational work.

- Improve developer experience: Reduce the time between an engineer writing code and seeing it safely in production. Build paved paths and self-service environments so the safe path is also the easy one.

- Raise the bar across engineering: Bring in new tools and practices, challenge how infrastructure and operations are normally done, and make reliability, security, and cost shared engineering values through clarity and evidence rather than formal authority.

Skills & Capability Stack: What You Need to Succeed

- 6–10 years of experience in DevOps, platform engineering, site reliability engineering, infrastructure engineering, or a closely related engineering role

- Strong understanding of how production systems fail, and the ability to identify failure modes, blast radius, and single points of failure before they bite
- Hands-on experience building or owning infrastructure and delivery platforms end-to-end, not just operating someone else's
- Strong hands-on experience with containers and orchestration in production, and with infrastructure as code.

- Strong Linux and networking fundamentals: DNS, TLS, load balancing, routing, and firewalls, enough to debug from first principles
- Ability to write real code, not only scripts, in Python, Go, or a comparable language
- Practical experience with observability: metrics, logging, tracing, and instrumentation

- Experience carrying production on-call, leading incidents, and writing postmortems that changed something
- Working knowledge of infrastructure security: secrets management, identity and access design, and hardening practices
- Deeply hands-on with AI engineering tools such as Claude Code, Grok, and Codex, used for infrastructure code, debugging, log and incident analysis, and automation

- Strong debugging skills and the ability to trace an issue across every layer of a live system, from application code to the network
- Ability to work across multiple clouds, providers, frameworks, and tools based on what the product requires
- Ability to balance speed, reliability, security, cost, and maintainability in a fast-moving environment

Added advantage

- Experience operating GPU infrastructure: fleet provisioning and health, GPU scheduling, driver and firmware lifecycle, or multi-tenant GPU isolation. Not required today, valuable as we bring our own capacity in-house.

- Experience keeping model-serving or inference endpoints reliable at scale: autoscaling, routing, rollout, and protecting p99 latency
- Capacity planning, scale testing, and disaster recovery exercises run as a discipline rather than an annual event
- Experience in a managed-services or SaaS operations environment, with incident, problem, and change management run to an agreed standard

Tech Environment: What You'll Work With

- A mixed, multi-provider cloud and hosting environment across our India and US entities not a single cloud console
- Kubernetes and containerised workloads

- Infrastructure as code
- CI/CD and release automation




- Observability: metrics, logs, traces, and dashboards
- Secrets, identity, and access management
- AI and LLM workloads running in production today on managed infrastructure, moving toward our own GPU capacity as we scale

- Claude Code, Grok, Codex, and other AI-assisted engineering and operations tooling

You're not expected to come in with a fixed toolbox. The expectation is that you can learn, adapt, experiment, and use the right tools to solve the infrastructure and reliability challenges we face.

What Makes You a Great Fit

- You don't think of yourself as a traditional sysadmin or release manager; you think like an engineer who owns how software runs.
- You naturally think about what could fail, not just whether the happy path works.
- You are calm when production is not, and your instinct in an incident is to gather evidence rather than guess.
- You are deeply curious about AI and already use AI tools as part of how you work.
- You automate your way out of repeat work, because doing the same manual fix twice bothers you.
- You write things down. Runbooks, postmortems, and decision records, because a platform that only works when you are awake is not a platform.
- You treat security and cost as engineering problems, not as someone else's compliance checklist.
- You care about building straightforward, reliable, scalable systems rather than creating process for the sake of process.
- You thrive in a high-ownership, high-intensity environment where speed and reliability both matter.

What Success Looks Like: First 90–180 Days

- Every environment is defined in code and can be rebuilt from scratch without tribal knowledge.
- CI/CD is fast, trusted, carries automated quality and security gates, and has a rollback that has been tested.
- Service level objectives exist for the critical services, are instrumented, and are actually used in shipping decisions.
- On-call, runbooks, escalation, and the postmortem loop are running, and incident actions are getting closed.
- Metrics, logs, and traces are in one place, and alert noise has gone down rather than up.
- Access, secrets, and infrastructure hardening are in a state that stands up to a customer security review.
- Cloud cost is visible per environment and service, with a monthly view leadership can act on and a measurable reduction behind it.
- Engineers ship more often, with less help, and with more confidence in what production will do.
- You become a trusted technical partner to Engineering, Product, and the QA and AI teams.

Why Join Nenu We believe great teams build great products, which is why we invest deeply in our people:

- Competitive salary and meaningful ESOPs
- Comprehensive healthcare and insurance coverage
- Unlimited leave policy built on trust and ownership
- Real growth in scope, ownership, and technical influence as the company scales

You'll be joining a team building what the future of AI could look like, from Chennai, today. We're solving problems and building products that are years ahead of where the industry is, giving you the opportunity to work on things that don't have a playbook yet. If you want to be part of a company thinking beyond today, building at the edge of AI, and creating something genuinely transformative, this is the right place to be.

If you're someone who takes ownership, thinks ahead, and wants to define how an AI-first company builds, ships, and runs software, we'd love to hear from you.

📌 DevOps / Platform & Reliability Engineer (Chennai)
🏢 Nenu AI
📍 Chennai

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: devops / platform & reliability engineer (chennai) / chennai

Subscribe to this job alert:

Get the latest job offers by email for: devops / platform & reliability engineer (chennai) / chennai