17 Aug
|
Teleport
|
Gurugram
Description
About the Role
We are looking for an engineer who enjoys understanding how systems behave in real production, not just writing features. This role is responsible for maintaining reliability, stability, and smooth functioning of our live platform running on Google Cloud.
You will act as the first technical owner of production systems — monitoring services, investigating alerts, resolving issues, and performing controlled configuration and operational changes. This role works closely with backend developers, QA, and infrastructure teams to prevent incidents and reduce downtime.
This is not a call-center support role and not a pure development role — it is a hands-on technical position focused on debugging, incident handling, and system operations.
Tech Stack
Google Cloud Platform (Compute, Logging, Monitoring)
Java (Spring Boot based microservices)
MongoDB
Apache Kafka (event-driven architecture)
Redis cache
Linux servers
Key Responsibilities
Production Monitoring & Alert Handling
Monitor application health, latency, errors, consumer lag,
database connections, and resource utilization
Acknowledge and investigate monitoring alerts
Perform first-level troubleshooting and stabilize services
Identify whether issue is infra, application, database, or messaging related
Incident Response
Participate in on-call rotation
Diagnose production incidents and restore services with minimal downtime
Safely restart services, scale instances, or rollback deployments when required
Communicate incident status to stakeholders
Technical Support & Operational Changes
Handle technical support tickets requiring engineering understanding
Update configurations and feature flags
Manage scheduled jobs / cron triggers
Trigger or replay events in Kafka
Assist in minor Java configuration/code fixes when needed
Coordinate production releases
Database & Messaging Operations
Investigate MongoDB performance issues and slow queries
Monitor and
📌 System Reliability Engineer Gurugram
🏢 Teleport
📍 Gurugram