25 Sep
|
Marriott Tech Accelerator
|
Hyderabad
25 Sep
Marriott Tech Accelerator
Hyderabad
Role Title:Senior Engineer II - SRE
Position ID: IHC461
Experience: 6+ Years
Employment Type:Full-time
Position Summary:
The Senior AIOps SRE Engineer II will design, engineer and operate highly reliable, scalable and secure services. The role combines Site Reliability Engineering with AI-assisted operations to improve observability, correlate operational signals, accelerate incident diagnosis and automate safe remediation. This engineer will partner with application, cloud, platform, security and support teams to embed reliability into the service lifecycle. The role will also lead major incident response, improve operational processes, mentor engineers and influence reliability practices across teams.
WhatSuccessLooksLike
Primary Outcomes:
- Reusable automation, secure infrastructure as code, observable deployments and measurable toil reduction.
- Service reliability: establish measurable reliability objectives and strengthen availability, resilience and disaster-recovery readiness.
- Operational intelligence: apply anomaly detection, event correlation and AI-assisted analysis to identify risk and speed diagnosis.
- Automation: reduce repetitive manual work through reusable, governed and auditable automation and self-healing patterns.
- Observability: create actionable telemetry across infrastructure, applications, logs, traces, user journeys and business transactions.
- Engineering leadership: lead incidents and post-incident learning while mentoring engineers and promoting blameless reliability practices.
Key Responsibilities:
- Define, implement and continuously improve Service Level Indicators, Service Level Objectives and error budgets for critical services, aligned to business expectations.
- Engineer for availability, scalability, performance, recoverability and graceful degradation across production services.
- Track and improve operational measures such as mean time to detect, mean time to acknowledge, mean time to restore, incident recurrence and SLO compliance.
Lead reliability reviews, failure-mode analysis, operational readiness reviews, capacity planning and disaster-recovery validation.
AIOps and Intelligent Operations:
- Implement AI-assisted anomaly detection, event correlation, alert enrichment and probable-cause analysis across operational data sources.
- Design workflows that combine telemetry, service context, change data, incident history and runbook knowledge to improve triage quality.
- Evaluate and operationalise AIOps capabilities in platforms such as Dynatrace Davis AI, ServiceNow ITOM/AIOps, BigPanda, Moogsoft or equivalent tools.
- Apply Generative AI responsibly to operational knowledge retrieval, incident summarisation and investigation assistance, with appropriate human review, access controls and auditability.
- Measure the effectiveness of AIOps use cases through signal quality, noise reduction, actionability, resolution outcomes and adoption.
Automation and Self-Healing:
- Develop production-grade automation using Python and Bash and engineer reusable infrastructure patterns using Terraform and Ansible.
- Create event-driven remediation workflows for well-understood failure scenarios with safeguards, validation, rollback and observability.
- Build ChatOps and runbook-automation integrations that enable consistent, traceable operational execution.
- Identify high-toil activities, prioritise automation opportunities and maintain reusable automation as supported engineering products.
Observability Engineering:
- Design and maintain full-stack observability using Dynatrace, Prometheus, Grafana, Splunk, ELK/OpenSearch, Open Telemetry or comparable platforms.
- Implement monitoring around the golden signals of latency, traffic, errors and saturation, supplemented by service-specific health and business indicators.
- Standardise logs, metrics, traces, dashboards, alerting policies and service ownership metadata to improve consistency and actionability.
- Build operational and leadership-level reliability scorecards that connect technical health to service outcomes.
Cloud, Platform and Release Engineering:
- Design scalable AWS architectures, including multi-account and multi-region patterns, secure networking, resilience and disaster recovery.
- Operate containerised workloads and Kubernetes platforms, including workload reliability, scaling, deployment safety and platform observability.
- Design and maintain reliable CI/CD pipelines using Harness, GitHub Actions, Jenkins or comparable platforms.
- Embed automated testing, security controls, policy checks, deployment verification and rollback capabilities into delivery pipelines.
- Manage and tune production SQL and NoSQL data services, including replication, failover, backup, recovery and performance.
Incident Management and Continuous Improvement:
- Lead response for critical production incidents, establish explicit technical coordination and support timely stakeholder communication.
- Conduct blameless post-incident reviews, identify systemic causes and ensure corrective actions are prioritised, owned and verified.
- Improve incident playbooks, escalation paths, on-call readiness, knowledge articles and runbooks.
- Collaborate with application support teams and third-party vendors to resolve issues and meet agreed service commitments.
Security, Governance and Leadership:
- Implement security hardening, least-privilege access, secrets management, vulnerability remediation and compliance controls.
- Ensure automation and AIOps workflows follow change, risk, data-protection and audit requirements.
- Mentor engineers, share technical knowledge and influence reliability standards across engineering and operations teams.
Mandatory Experience:
- 6+ years of experience across Site Reliability Engineering, cloud operations, DevOps, platform engineering, production engineering or a closely related discipline.
- Having hands-on experience with LLMs, RAG , Models, MLOps, AWS/ Azure/GCP, Python.
- Demonstrated experience operating business-critical production services and leading complex incident diagnosis and restoration.
- Strong understanding of distributed systems, reliability patterns, service dependencies, failure modes and production risk.
- Experience working across development, infrastructure, security, service management and application-support teams.
Preferred Qualifications and Role Expectations:
Good to Have:
- Hands-on exposure to ServiceNow ITOM/AIOps, Dynatrace Davis AI, BigPanda, Moogsoft, PagerDuty Operations Cloud or comparable platforms.
- Experience integrating operational workflows with cloud-native AI services or enterprise Generative AI platforms.
- Practical understanding of Large Language Models, Retrieval-Augmented Generation, vector search and AI-agent patterns in an enterprise operations context.
- Experience designing self-healing systems, auto-remediation frameworks, event-driven automation and policy-based operational controls.
- Knowledge of performance modelling, long-term capacity forecasting and cost-aware reliability engineering.
- Experience with multi-account AWS environments, enterprise landing zones and platform governance.
Education and Certifications:
- Undergraduate degree or equivalent professional experience/certification.
- Relevant certifications are advantageous, including AWS Solutions Architect or DevOps Engineer, CKA or CKAD, Terraform Associate, Dynatrace, Splunk, ServiceNow ITOM or SRE Foundation.
- Leadership and Behavioural Expectations
- Communicates clearly during high-severity incidents and translates technical complexity into concise business impact and recovery updates.
- Uses data and engineering evidence to prioritise reliability improvements and challenge low-value operational practices.
- Demonstrates ownership while fostering a blameless, learning-focused culture.
- Builds trusted partnerships across global, cross-functional and vendor teams.
- Mentors engineers and contributes reusable standards, patterns and knowledge.
📌 Senior Engineer II - SRE _ (Hyderabad)
🏢 Marriott Tech Accelerator
📍 Hyderabad