02 Oct
|
The Standard India
|
Bengaluru
02 Oct
The Standard India
Bengaluru
The next part of your journey is right around the corner — with The Standard India.
A genuine desire to make a difference in the lives of others is the foundation for everything we do. With a customer ‑ first mindset and an intentional focus on building solid teams across the nation, we’ve been able to uphold our legacy of financial stability while investing in innovative technologies that support the needs of our customers. Our high-performance culture thrives thanks to remarkable people united by compassion and a commitment to operational excellence.
You’ll join an engineering organization focused on resilience, reliability, and continuous improvement — where your work ensures the stability and performance of business-critical applications used across the enterprise. Are you ready to make a difference?
Job Summary The Observability Engineer (Datadog), Level 4, owns how observability is done across the estate — the instrumentation standards, tagging taxonomy, SLO framework, cost-governance model, and the migration strategy that others execute. This role sets the standards and reference patterns that the Level 3 Observability Engineers and application teams build against.
This is a hands-on technical leadership role with no direct reports — you lead through architecture, standards, and influence rather than a reporting line. You will own the observability reference architecture, define the SLO and alerting policy adopted platform-wide, own the AppDynamics-to-Datadog migration strategy and sequencing while Level 3 engineers execute the service-by-service work, and get teams that do not report to you — application, platform, security, and finance — to adopt the standards and instrument their own services. Our outcome measure is business impact minutes and the health of the platform as a whole: you are judged not by whether one monitor fired correctly but by whether the estate got more reliable and more cost-efficient because of decisions you made and drove.
This is not a monitoring or command center role.
- Ability to work on‑site in Bengaluru, India is a requirement of the role.
Mode of Working: This role follows a hybrid work model, with a primary base in Bengaluru. While on‑site presence is essential for key engagements, flexibility is offered for remote work based on business needs and team alignment. Reports to: Director of Service Reliability or designated manager.
Key Responsibilities
- Own the observability reference architecture and instrumentation standards the organization builds against — the paved roads, tagging taxonomy, and deployment patterns that Level 3 engineers and application teams adopt, rather than one-off dashboards.
- Own the AppDynamics-to-Datadog migration strategy and sequencing across the estate — deciding order, parity criteria, and decommissioning gates — while Level 3 engineers execute the service-by-service work.
- Define the SLO and error-budget framework and the alerting policy adopted platform-wide, so that signal quality, ownership,
and actionability are consistent across every team rather than tuned service by service.
- Own the telemetry cost and cardinality governance model — the budgets, guardrails, and tradeoff decisions across log volume, indexed spans, and custom-metric cardinality — and the configuration-as-code approach (Terraform and the Datadog API) that keeps the platform reproducible at scale.
- Influence without authority — get application, platform, security, and finance teams that do not report to you to adopt the standards, respect cost budgets, and instrument their own services, turning the reference patterns into org-wide practice.
Skills And Background You’ll Need
Education
Required: High School diploma Preferred: Bachelor’s degree in Computer Science, Engineering, Information Systems, or related field
Experience — Must‑Have Requirements
- Hands-on Datadog experience in production, deep enough to make architecture-level decisions — agent deployment at scale across mixed operating systems, custom and API-based integrations where no native one exists, APM instrumentation owned through a migration, and monitors and dashboards managed as code. You should be able to explain not just a rollout you worked on, but a reference pattern you designed and why.
- Roughly 7 to 10 or more years in observability, site reliability, infrastructure, or platform engineering — but gated on evidence of estate-wide ownership and cross-team adoption, not year count alone.
- Applied competence in Azure Kubernetes Service (AKS) and Linux — comfortable in a containerized environment and able to reason about how instrumentation behaves inside it, including agent deployment as a DaemonSet, auto discovery, and pod and container tagging. Scripting and automation in Python or equivalent.
- Judgment on telemetry cost and cardinality — able to reason about the cost consequence of log volume, indexed spans, and custom metric cardinality, and to catch an unbounded tag or a chatty metric before it reaches the invoice.
- Demonstrated ownership of outcomes: a standard, paved road, or reference pattern you created that other engineers adopted, and a result you moved for an organization — detection coverage, mean time to resolve, or ingest cost trend — not just a system you personally built.
- Cross-functional influence without authority: evidence of getting teams that do not report to you to change behavior — adopt a standard, instrument their own services, or respect a cost budget — supported by clear written communication in English that other engineers act on without you in the room.
Preferred Experience
- Experience migrating from AppDynamics,
Dynatrace, New Relic, or a similar platform onto Datadog.
- Infrastructure as code with Terraform, a directional goal for how we manage the platform.
- Open Telemetry instrumentation experience.
- ServiceNow ITSM integration work in a production incident context.
- Experience in a large, regulated enterprise estate or a global capability center operating model.
Desirable Certifications
- Datadog Certifications (Observability, APM, Logs)
- Google Professional Cloud DevOps Engineer or equivalent SRE coursework (concept-level: SLOs, error budgets, reliability practice)
- Certifications are desirable only and never a substitute for demonstrated estate-wide impact and cross-team adoption.
Additional Requirements
- Ability to work on‑site in Bengaluru, India is a requirement of the role.
- - Hybrid role with on‑site presence required for key meetings, collaboration, and incident readiness exercises.
- Standard India business hours expected, with participation in on‑call rotations based on team schedules.
- India business hours with a minimum three-hour daily overlap with United States Pacific Time.
What Positive Looks Like In Your First Year
- Services migrated off AppDynamics onto Datadog APM against a migration sequence you own, with Level 3 engineers executing to your parity and decommissioning gates.
- An SLO and alerting policy adopted platform-wide, not just applied to services you touched personally.
- A telemetry cost and cardinality governance model in place, with budgets and guardrails other teams work within.
- The observability reference architecture and instrumentation standards published and being built against across the estate.
- Application and platform teams that do not report to you self-serving against your patterns and instrumenting their own services.
Key Behaviors of a Successful Candidate
- Adaptability: Adjusts quickly to evolving priorities and incident situations; remains calm, analytical, and collaborative under pressure.
- Independence: Takes ownership of decisions and technical troubleshooting, proactively identifying reliability risks before they become incidents.
- Willingness to Learn: Continuously grows technical competence in reliability engineering, automation, incident response, observability, and cloud‑native practices.
Why Join The Standard India? You Can Expect
- A rich benefits package supporting health, wellbeing, and long‑term financial goals
- Competitive compensation aligned with Indian market benchmarks
- Annual incentive bonus tied to individual and organizational performance
- A hybrid working environment grounded in trust and collaboration
- Strong mentorship, growth pathways, and support for professional certifications
- Purpose‑driven engineering work that meaningfully impacts customers and communities
Incentive program eligibility is subject to program rules and performance. StanCorp Global Services India Private Limited is an equal opportunity employer.
📌 Observability Engineer (Datadog) IV (Bengaluru)
🏢 The Standard India
📍 Bengaluru