15 Sep
|
Technogen India
|
Hyderabad
15 Sep
Technogen India
Hyderabad
Job Summary
Observability Automation and Service Enablement Engineer Suggested internal level: F / Specialist, Technology Engineering About the role Nationwides Cyber Security Operations - Innovation and Integration team is hiring a hands-on engineer to operate and continuously improve the observability services behind key systems the team builds and runs. This role blends responsive service delivery with automation engineering: youll own requests coming through ServiceNow, Microsoft Teams, email, and other channels while evolving the alerting and monitoring around those systems. Youll manage New Relic alerts and dashboards at scale (1k+) using Python and Terraform, and bring transferable experience with Grafanas LGTM stack - Loki, Grafana, Tempo, and Mimir.
This is not a traditional ticket-administration job: success requires a service mindset, sound operational judgment, and the engineering ability to replace repetitive manual work with governed, supportable automation. AI-assisted development experience is a plus - theres room to build agentic workflows that triage and contextualize issues.
Responsibilities
- Own intake across ServiceNow, Teams, email, and other channels - triage, clarify requirements, set expectations, prioritize, and keep status and fulfillment records accurate.
- Develop and maintain Python and Terraform automation that generates and deploys New Relic artifacts (dashboards, alert policies, conditions, notification workflows) at scale via APIs, NerdGraph, and NRQL.
- Partner with application, infrastructure, cloud, security, and platform teams to define helpful telemetry, actionable alerts, ownership, routing,
severity, and SLIs.
- Reduce alert noise and configuration drift by tuning thresholds, routing, enrichment, and runbook context, and by reconciling duplicate, stale, or unowned assets.
- Turn recurring request patterns into governed self-service or automated workflows, and maintain runbooks, standards, and customer-facing guidance.
- Grow into applying the same engineering patterns to Grafana and the LGTM stack.
Must-have qualifications
- Near-native spoken and written English in IT technical contexts, with strong customer engagement - requirements clarification, expectation management, and follow-through.
- Hands-on New Relic experience - dashboards, NRQL, alert policies and conditions, notification workflows, APIs, and troubleshooting.
- Production-quality Python automation - API integrations, structured configuration, error handling, logging, testing, and maintainable packaging.
- Hands-on Terraform - reusable modules, remote state, provider configuration, imports, lifecycle management, validation, and CI/CD integration.
- Git-based configuration workflows - pull requests, peer review, branching, and controlled promotion across environments.
- Ability to troubleshoot integrations and data flows across APIs, networks,
cloud services, agents, collectors, and telemetry back ends.
- Solid alert-design fundamentals - actionable signals, ownership, severity, suppression, deduplication, routing, and noise reduction.
- Experience within IT service management, change control, incident management, and production support.
- Ability to balance interrupt-driven operational work with planned engineering improvements.
Nice to have
- Working knowledge of Grafana and the LGTM stack (Loki, Grafana, Tempo, Mimir or Prometheus-compatible systems); Grafana Enterprise/Cloud, Alertmanager, or OpenTelemetry Collector experience.
- New Relics Terraform provider, NerdGraph, dashboard-as-code, or large-scale alert migration and standardization.
- Distributed tracing and OpenTelemetry; debugging telemetry flows across syslog, HEC, Kafka, raw HTTP, or raw TCP/binary protocols; Splunk SPL and forwarder configuration.
- AWS, Kubernetes, containers, Linux, and cloud-native architectures; CI/CD platforms with automated policy, security, lint, test, plan, and deploy gates.
- Integrating observability platforms with ServiceNow, Teams, webhooks, event management, or on-call tooling.
- SLIs/SLOs, error budgets, golden signals, or RED/USE monitoring patterns; secrets management, least-privilege, and secure auto
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Observability Automation and Service Enablement Engineer (Hyderabad)
🏢 Technogen India
📍 Hyderabad