Staff Software Engineer (Bengaluru)

Staff Software Engineer (Bengaluru)

10 Aug
|
GreyOrange
|
Bengaluru

10 Aug

GreyOrange

Bengaluru

This role is responsible for two things that are more connected than they first appear: (1) making production deployments safe enough that rollback is the default response to a problem, and (2) making configuration changes safe, audited, and AI-supervised so that bad configuration is caught before it reaches production. Both are fundamentally release-safety problems.

GreyOrange operates a platform that multiple enterprise customers (DHL, XPO, Continental, others) run in production warehouses. A bad deployment or a bad configuration change is a customer-visible incident. The infrastructure you build is what stands between a code change and a customer outage.

The company's ability to keep all customers on the same major platform version — rather than maintaining separate branches per customer — depends entirely on whether rollback is fast, trustworthy, and routine.

After establishing the deployment and configuration management foundations, you expand into data and messaging engineering — specifically Kafka schema contracts and the Redis claim-check pattern for payload management. The overlap between distributed delivery systems (Kubernetes + Kafka both require schema and contract thinking) makes this a natural progression.

What You Will Do First Six Months - Deployment Safety & Configuration

- Take ownership of the deploy and release engineering domain. Audit the existing Spinnaker pipelines and ArgoCD application definitions to understand the starting point.
- Conduct a Consul configuration inventory — catalog every human-edited key, identify overlaps with the existing configuration management system, and publish a written boundary document clarifying which configs live where.
- Design and ship blue-green progressive delivery for the top three critical stateless services using ArgoCD Rollouts. Write a shared Rollout template that all subsequent services follow.
- Standardise Spinnaker pipeline templates for VM-deployed services.



Define smoke-test requirements as a mandatory promotion gate.
- Ship the Java/Spring feature-flag SDK, piloted against one service. Then Erlang and React SDKs.
- Implement Consul-as-backend: bring human-edited Consul keys behind the configuration management UI. Remove direct write access for humans.

Next Six Months - Configuration Intelligence & Messaging

- Ship AI-supervised configuration: train or fine-tune a model on historical incident data, surface risk warnings in the configuration UI. First version can be rule-based with an LLM layer on top.
- Wire the configuration audit log into the compliance evidence pipeline so configuration changes become part of the joined audit trail alongside access events and incident records.
- Own Kafka schema contract rollout: Schema Registry adoption and Avro migration for the highest-traffic topics. Partner with service teams through the migration.
- Build and ship the Redis claim-check SDK. Monitor payload-size distribution and consumer lag. Demonstrate measurable reduction in Kafka payload sizes on the pilot topics.

Ongoing

- Run rollback drills on a regular cadence — each customer environment should execute at least one controlled rollback per half so the process is practised before it is needed under pressure.
- Maintain the deployment and configuration systems as production services — they are part of the platform team's on-call rotation.
- Contribute to cross-team architecture reviews with a deployment safety and release-engineering lens.




- Write and maintain design documents and runbooks for everything you build.

Requirements

- 8+ years of backend or platform engineering with strong Kubernetes and GitOps depth
- Production experience with ArgoCD, ArgoCD Rollouts, Flagger, or comparable progressive-delivery tooling — has configured traffic shifting, canary analysis, and automated rollback in a real production setting
- Has built or substantially extended a feature-flag or runtime configuration management system — not just used a commercial tool, but understood or built the machinery underneath
- Familiarity with Consul, Vault, or the broader HashiCorp ecosystem at an engineering level
- Experience with Spinnaker or a comparable continuous delivery platform for VM-based deployments
- Uses deployment frequency, change failure rate, and mean time to restore as engineering metrics — not just awareness of what DORA is, but a track record of using it to prioritise engineering investment
- Comfortable with heterogeneous deployment environments - both Kubernetes and traditional VMs as first-class production targets in the same organisation

Nice to Have

- Experience with Apache Camel or Apache Flink - relevant for future work in integration and streaming engineering
- Background with Kafka Connect, schema registries, or MQTT integration - the platform bridges Kafka and MQTT for robot communication
- Multi-tenant SaaS background where customer-segmented feature rollouts are standard practice - GreyOrange deploys to multiple enterprise customers with different risk tolerances and release windows
- Familiarity with LLM APIs (Vertex AI, OpenAI) and retrieval-augmented generation - the AI configuration risk-detection work uses an LLM layer on top of historical incident data
- Erlang/OTP experience or genuine curiosity about it - the Erlang feature-flag SDK requires understanding how OTP applications load and manage configuration at runtime

📌 Staff Software Engineer (Bengaluru)
🏢 GreyOrange
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: staff software engineer (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: staff software engineer (bengaluru) / bengaluru