02 Sep
|
AppleTech
|
Vadodara
02 Sep
AppleTech
Vadodara
You'll own the forensic layer of our platform proving root cause definitively instead of guessing, instrumenting our systems so failures announce themselves before they happen, and retiring the brute-force reboots the team currently survives on.
You'll be the person who can say, with evidence, whether an outage was the kernel, the storage, the network, or the code and then fix the underlying condition so it doesn't recur.This is a hands-on engineering role, not a ticket support one.
What you'll actually do
- Root-cause forensics. When Tomcat or Spring Boot stalls, you'll pull heap and thread dumps, trace database locks, and monitor sockets at the OS level to prove not speculate whether threads are starved by the kernel, blocked on slow storage IOPS, or waiting on bad code. Ambiguous incidents end with a definitive answer.
- Platform instrumentation. You'll deploy lightweight, safe telemetry (node_exporter, jmx_exporter) across the DMS and CI/CD fleet and wire it into our Command Console, giving the 247 team a live "crash-cart" view — the historical context of the moments before a failure, not just the wreckage after.
- Network triangulation. You'll give the CI/CD and infrastructure teams mathematical proof of where things break, using OS-level network tracing to isolate OpenStack API timeouts, DNS degradation, and TCP packet loss. The goal is to officially end the network blame game.
- Technical-debt reduction.
You'll tune JVM garbage collection and Hibernate connection pools to stabilize the legacy DMS application, design a safe data-lifecycle and pruning strategy for MySQL to prevent disk-full outages, and write the operational runbooks that move the team from manual firefighting to repeatable process.
What we're looking for
- Deep, hands-on JVM performance and troubleshooting experience — heap/thread-dump analysis, GC tuning, and diagnosing thread starvation and connection-pool contention under real production load.
- Solid Linux systems fundamentals — you're comfortable at the OS level with sockets, storage/IOPS, process and kernel behavior, and network tracing tools (tcpdump, ss, strace, and similar).
- Practical experience running Java web applications in production: Tomcat, Spring Boot, and Hibernate.
- Solid MySQL operational skills — locking, query behavior, and data-lifecycle/pruning strategy, not just schema design.
- Experience building metrics-based observability with the Prometheus ecosystem (node_exporter, jmx_exporter, Grafana or equivalent dashboards).
- A forensic mindset: you're not satisfied until you can prove the cause, and you write down what you found so the next person doesn't have to rediscover it.
Nice to have
- On-prem / private-cloud infrastructure experience, especially OpenStack.
- Background supporting a 247 operations function and building tooling that makes on-call less painful.
- Experience turning tribal knowledge into runbooks and repeatable process.
📌 Sr. Systems Reliability Engineer (DevOps support) (Vadodara)
🏢 AppleTech
📍 Vadodara