01 Oct
|
AssetMark Global Wealth
|
Hyderabad
01 Oct
AssetMark Global Wealth
Hyderabad
About AssetMark
AssetMark is a leading wealth management platform dedicated to empowering independent financial advisors. AssetMark's mission is to enable financial advisors to make a profound difference in the lives of their clients. Over 10,000 advisors partner with AssetMark for our investment offerings, innovative technology, advanced services, and expertise, which they use to delight their clients and grow their businesses.
We are an integrated team of technologists, investment professionals, and operations experts working together to advance wealth management.
Position Summary The Vice President of Service Resilience is the senior leader responsible for establishing and advancing Site Reliability Engineering and Observability across all AssetMark systems.
This role will define the strategy, operating model, technical standards, and roadmap required to make AssetMark's applications, infrastructure, cloud platforms, data services, networks, and critical integrations more reliable, resilient, observable, performant, and supportable.
Based in Hyderabad with a global mandate, the VP will build and direct a high-performing team and work closely with Cloud and Infrastructure Engineering, Application Engineering, Data, Security, Identity and Access Management, Service Management, and business leaders across time zones.
This is a highly visible technical leadership position. The successful candidate must combine deep expertise in observability and SRE with strong executive presence, strategic thinking, innovation, and the ability to influence engineering practices across the organization.
Primary Responsibilities
- Service Resilience Strategy: Define and execute a multi-year Service Resilience strategy for all AssetMark systems. Establish service tiers, reliability objectives, production-readiness expectations, dependency standards, and an SRE operating model that makes resilience an engineering responsibility rather than a post-incident activity.
- Strategy, Leadership and Team Development: Build and lead the Service Resilience function, including its strategy, operating model, service catalog, technology roadmap, engineering standards, governance practices, budget, vendors, and delivery partners. Recruit, mentor, and develop SRE, Observability, Performance, and Resilience Engineering talent.
- Observability Platform and OpenTelemetry: Own the enterprise observability strategy and platform architecture. Establish standards for metrics, logs, traces, distributed tracing, profiles, events, dashboards, alerting, and service-health reporting. Lead adoption and governance of OpenTelemetry, including instrumentation standards, semantic conventions, collectors, exporters, telemetry pipelines, context propagation, data quality, retention, and cost management.
- APM and End-to-End Monitoring: Lead the use of Azure Monitor, Application Insights, Log Analytics, KQL, Prometheus, Grafana, and enterprise APM platforms such as Dynatrace, Datadog, New Relic, AppDynamics, Splunk, Elastic, or comparable technologies. Create a consistent view of application, API, database, cloud, network, infrastructure, user-experience, and critical-business-transaction health.
- SRE Engineering and Reliability Standards: Establish and mature practices for Service Level Indicators, Service Level Objectives, Service Level Agreements, error budgets, capacity planning, performance engineering, load testing, availability engineering, dependency management, service reviews, operational readiness, and toil reduction. Coach engineering teams on designing services that are measurable, resilient, recoverable, and supportable.
- Performance, Capacity and Resilience Engineering: Establish performance, capacity, and resilience engineering practices across applications, Azure, infrastructure, networks, databases, and critical integrations. Lead capacity planning, dependency analysis, load testing, failure testing, disaster-recovery exercises, resilience game days, and recovery validation.
- Incident, Problem and Recovery Engineering: Partner with Major Incident and Problem Management teams to improve technical response, diagnosis, recovery, and prevention. Establish runbooks, playbooks, escalation paths, dependency maps, blameless postmortems, corrective-action governance,
and repeatable recovery practices. Provide technical leadership during high-severity incidents and executive escalations.
- Azure and Cloud Service Reliability: Provide reliability leadership across Azure and cloud-hosted services. Partner with Cloud and Infrastructure Engineering to establish standards for monitoring, scaling, availability, resiliency, backup, recovery, and operational supportability across Azure compute, storage, databases, networking, containers, applications, and platform services.
- Service Mapping and Dependency Management: Create an end-to-end view of business services, application dependencies, infrastructure components, data stores, networks, identity services, vendors, and external integrations. Establish service ownership and dependency information that supports faster diagnosis, impact analysis, risk management, and executive communication.
- Automation, AI and Innovation: Promote observability-as-code, reliability-as-code, automated runbooks, self-healing capabilities, and automated remediation. Evaluate responsible uses of AI and AIOps for anomaly detection, event correlation, alert triage, root-cause analysis, knowledge management, predictive capacity planning, and operational productivity, with appropriate security and human oversight.
- Reliability Governance and Executive Reporting: Establish executive dashboards that communicate system health, reliability trends, risk, incident performance, SLO attainment, error-budget consumption, MTTD, MTTR, repeat incidents, change-related failures, toil, and roadmap progress. Use reliability data to inform engineering priorities and investment decisions.
- Security, Risk and Compliance: Embed security, privacy, and compliance requirements into observability and SRE practices. Partner with Cybersecurity, IAM, Cloud, and Application teams to ensure telemetry, access, logging, incident response, data retention, and recovery processes support organizational risk requirements.
- Financial and Vendor Management: Own planning and optimization for observability, APM, monitoring, telemetry, reliability, and related platform costs. Manage strategic vendors, contracts, renewals, service performance, commercial negotiations, and technology investments.
- Cross-Functional Delivery: Partner with Cloud and Infrastructure Engineering, IAM, Cybersecurity, Applications, Data, Architecture, Service Management, vendors, and business leaders to translate reliability goals into secure, resilient, measurable, and adopted engineering practices.
Organizational Scope The VP will lead a Hyderabad-based global capability across the following areas:
- Site Reliability Engineering: SRE, service health, SLOs, error budgets, operational readiness, reliability reviews, and reliability improvement programs.
- Observability Platform Engineering: OpenTelemetry, telemetry pipelines, APM, Azure Monitor, dashboards, alerting, logging, tracing, and monitoring standards.
- Performance and Capacity Engineering: Performance testing, capacity planning, application analysis, dependency management, and platform performance.
- Resilience Engineering: Failure testing, disaster recovery, recovery validation, resilience assessments, game days, and service continuity.
- Reliability Automation: Automated runbooks, self-service, observability-as-code, remediation, reporting, and operational productivity.
Required Qualifications
- Experience: Typically 15 or more years of progressive experience in software engineering, infrastructure, cloud, operations, SRE, Observability, or related disciplines, with significant experience leading enterprise technology functions.
- Leadership: Significant experience leading global technical teams, developing managers and engineers, building new capabilities, and directing delivery through a strategic roadmap.
- Service Resilience: Demonstrated success establishing or maturing SRE, Observability,
APM, reliability, performance, or production-engineering capabilities at enterprise scale.
- OpenTelemetry and APM: Deep experience with OpenTelemetry, instrumentation, collectors, telemetry pipelines, metrics, logs, traces, distributed tracing, dashboards, alerting, APM, and application-performance analysis.
- Azure and Monitoring: Strong hands-on experience with Azure and Azure monitoring capabilities, including Azure Monitor, Application Insights, Log Analytics, KQL, and related cloud services.
- SRE Practices: Strong understanding of SLIs, SLOs, SLAs, error budgets, service tiers, incident response, postmortems, resilience engineering, performance engineering, capacity planning, and toil reduction.
- Distributed Systems: Strong understanding of distributed systems, microservices, APIs, cloud-native architectures, containers, Kubernetes, and complex application dependencies.
- Automation: Experience using scripting, APIs, CI/CD, Infrastructure as Code, observability-as-code, workflow automation, or similar engineering practices to improve reliability.
- Executive Presence: Excellent executive communication, stakeholder management, presentation, analytical, and problem-solving skills, with the ability to move between strategy, business impact, and technical detail.
- Innovation: Demonstrated ability to be innovative and forward thinking, evaluate emerging technology, and turn practical ideas into measurable improvements in reliability and engineering productivity.
Preferred Qualifications
- Observability Platforms: Experience with Dynatrace, Datadog, Current Relic, AppDynamics, Grafana, Prometheus, Splunk, Elastic, Azure Managed Prometheus, Azure Managed Grafana, or comparable platforms.
- Cloud-Native Platforms: Experience with Kubernetes, AKS, service meshes, event-driven systems, distributed platforms, synthetic monitoring, real-user monitoring, and cloud-native architecture.
- Reliability Innovation: Experience with chaos engineering, automated remediation, AIOps, AI-assisted operations, predictive analytics, or reliability intelligence.
- Engineering Tooling: Experience with Terraform, GitOps, policy-as-code, Azure DevOps, GitHub Actions, platform APIs, internal developer platforms, or observability-as-code.
- Industry: Experience within financial services or another highly regulated enterprise environment, including audit, regulatory examination, and remediation activities.
- Certifications: Relevant certifications in Azure, cloud architecture, SRE, ITIL, cybersecurity, or comparable disciplines.
Leadership and Accountability Model The VP serves as the enterprise service owner and engineering authority for Service Resilience, SRE, Observability, APM, performance, capacity, and resilience engineering practices. The role establishes the strategy, standards, roadmaps, technical operating model, and executive reporting. Application, platform, data, and business owners retain accountability for their services, while Cloud and Infrastructure Engineering, IAM, Cybersecurity, and Service Management provide complementary platform, identity, security, and operational capabilities.
What Success Looks Like The successful VP will transform Service Resilience from distributed monitoring and operational activity into an integrated enterprise capability. Critical AssetMark services will have clear ownership, service tiers, SLIs, SLOs, dashboards, alerts, runbooks, dependency maps, production-readiness standards, and measurable reliability roadmaps.
OpenTelemetry and common observability standards will provide consistent visibility across applications, APIs, Azure, infrastructure, networks, databases, user experience, and critical business transactions. Mean Time to Detect, Mean Time to Respond, Mean Time to Restore, repeat incidents, alert noise, change-related failures, and operational toil will improve.
Reliability and error-budget data will influence engineering roadmaps and release decisions. Resilience testing, disaster recovery, performance engineering, and failure analysis will become routine practices. The Hyderabad team will become a global center of excellence for Service Resilience, Observability, and SRE, using automation and responsible AI to improve resilience and operational insight.
📌 Vice President, Service Resilience (Hyderabad)
🏢 AssetMark Global Wealth
📍 Hyderabad