Senior Engineer II - Data Science / MLOps-29625] (Hyderabad)

Senior Engineer II - Data Science / MLOps-29625] (Hyderabad)

24 Sep
|
Marriott Tech Accelerator
|
Hyderabad

24 Sep

Marriott Tech Accelerator

Hyderabad

About Us:

Marriott International Inc., headquartered in Bethesda, Maryland, USA, was founded in May 1927 by J. Willard Marriott and Alice S. Marriott with a modest nine-seat A&W; root beer stand. Guided by the family's leadership and core principles, Marriott International today has grown into a global hospitality giant, operating approximately 10,000 properties and over 30 leading brands in more than 140 countries and territories.

From such humble beginnings to becoming the world’s largest hotel company, Marriott International has never stopped searching for inventive ways to serve its customers, provide opportunities for its associates, and grow their business. At Marriott Tech Accelerator center (MTA), Hyderabad, India, Marriott is exploring the world we live in and all its possibilities.

At Marriott Tech

Accelerator, we are a team of passionate engineering minds dedicated to creating and building cutting-edge solutions that streamline operations and elevate guest experiences.

Marriott Tech Accelerator is an ANSR entity incorporated exclusively for providing services to Marriott International.

Role Title: Senior Engineer II - SRE

Experience: 6+ Years

Employment Type: Full-time

Position Summary:

The Senior AIOps SRE Engineer II will design, engineer and operate highly reliable, scalable and secure services. The role combines Site Reliability Engineering with AI-assisted operations to improve observability, correlate operational signals, accelerate incident diagnosis and automate safe remediation. This engineer will partner with application, cloud, platform, security and support teams to embed reliability into the service lifecycle.

The role will also lead major incident response, improve operational processes, mentor engineers and influence reliability practices across teams.

What Success Looks Like:

Primary Outcomes:

- Reusable automation, secure infrastructure as code, observable deployments and measurable toil reduction.
- Service reliability: establish measurable reliability objectives and strengthen availability, resilience and disaster-recovery readiness.
- Operational intelligence: apply anomaly detection, event correlation and AI-assisted analysis to identify risk and speed diagnosis.
- Automation: reduce repetitive manual work through reusable, governed and auditable automation and self-healing patterns.
- Observability: create actionable telemetry across infrastructure, applications, logs, traces, user journeys and business transactions.
- Engineering leadership: lead incidents and post-incident learning while mentoring engineers and promoting blameless reliability practices.

Key Responsibilities:

- Define, implement and continuously improve Service Level Indicators, Service Level Objectives and error budgets for critical services, aligned to business expectations.
- Engineer for availability, scalability, performance, recoverability and graceful degradation across production services.
- Track and improve operational measures such as mean time to detect, mean time to acknowledge, mean time to restore, incident recurrence and SLO compliance.
- Lead reliability reviews, failure-mode analysis, operational readiness reviews,



capacity planning and disaster-recovery validation.

AIOps and Intelligent Operations:

- Implement AI-assisted anomaly detection, event correlation, alert enrichment and probable-cause analysis across operational data sources.
- Design workflows that combine telemetry, service context, change data, incident history and runbook knowledge to improve triage quality.
- Evaluate and operationalize AIOps capabilities in platforms such as Dynatrace Davis AI, ServiceNow ITOM/AIOps, BigPanda, Moogsoft or equivalent tools.
- Apply Generative AI responsibly to operational knowledge retrieval, incident summarization and investigation assistance, with appropriate human review, access controls and auditability.
- Measure the effectiveness of AIOps use cases through signal quality, noise reduction, actionability, resolution outcomes and adoption.

Automation and Self-Healing:

- Develop production-grade automation using Python and Bash and engineer reusable infrastructure patterns using Terraform and Ansible.
- Create event-driven remediation workflows for well-understood failure scenarios with safeguards, validation, rollback and observability.
- Build ChatOps and runbook-automation integrations that enable consistent, traceable operational execution.
- Identify high-toil activities, priorities automation opportunities and maintain reusable automation as supported engineering products.

Observability Engineering:

- Design and maintain full-stack observability using Dynatrace, Prometheus, Grafana, Splunk, ELK/OpenSearch, Open Telemetry or comparable platforms.
- Implement monitoring around the golden signals of latency, traffic, errors and saturation, supplemented by service-specific health and business indicators.
- Standardize logs, metrics, traces, dashboards, alerting policies and service ownership metadata to improve consistency and actionability.
- Build operational and leadership-level reliability scorecards that connect technical health to service outcomes.

Cloud, Platform and Release Engineering:

- Design scalable AWS architectures, including multi-account and multi-region patterns, secure networking, resilience and disaster recovery.
- Operate containerized workloads and Kubernetes platforms, including workload reliability, scaling, deployment safety and platform observability.
- Design and maintain reliable CI/CD pipelines using Harness, GitHub Actions, Jenkins or comparable platforms.
- Embed automated testing, security controls, policy checks, deployment verification and rollback capabilities into delivery pipelines.
- Manage and tune production SQL and NoSQL data services, including replication, failover, backup, recovery and performance.

Incident Management and Continuous Improvement:

- Lead response for critical production incidents,



establish explicit technical coordination and support timely stakeholder communication.
- Conduct blameless post-incident reviews, identify systemic causes and ensure corrective actions are prioritized, owned and verified.
- Improve incident playbooks, escalation paths, on-call readiness, knowledge articles and runbooks.
- Collaborate with application support teams and third-party vendors to resolve issues and meet agreed service commitments.

Security, Governance and Leadership:

- Implement security hardening, least-privilege access, secrets management, vulnerability remediation and compliance controls.
- Ensure automation and AIOps workflows follow change, risk, data-protection and audit requirements.
- Mentor engineers, share technical knowledge and influence reliability standards across engineering and operations teams.

Mandatory Experience:

- 6+ years of experience across Site Reliability Engineering, cloud operations, DevOps, platform engineering, production engineering or a closely related discipline.
- Having hands-on experience with LLM’s, RAG , Models, MLOps, AWS/ Azure/GCP, Python.
- Demonstrated experience operating business-critical production services and leading complex incident diagnosis and restoration.
- Strong understanding of distributed systems, reliability patterns, service dependencies, failure modes and production risk.
- Experience working across development, infrastructure, security, service management and application-support teams.

Preferred Qualifications and Role Expectations:

Good to Have:

- Hands-on exposure to ServiceNow ITOM/AIOps, Dynatrace Davis AI, BigPanda, Moogsoft, PagerDuty Operations Cloud or comparable platforms.
- Experience integrating operational workflows with cloud-native AI services or enterprise Generative AI platforms.
- Practical understanding of Large Language Models, Retrieval-Augmented Generation, vector search and AI-agent patterns in an enterprise operations context.
- Experience designing self-healing systems, auto-remediation frameworks, event-driven automation and policy-based operational controls.
- Knowledge of performance modelling, long-term capacity forecasting and cost-aware reliability engineering.
- Experience with multi-account AWS environments, enterprise landing zones and platform governance.

Education and Certifications:

- Undergraduate degree or equivalent professional experience/certification.
- Relevant certifications are advantageous, including AWS Solutions Architect or DevOps Engineer, CKA or CKAD, Terraform Associate, Dynatrace, Splunk, ServiceNow ITOM or SRE Foundation.
- Leadership and Behavioural Expectations
- Communicates clearly during high-severity incidents and translates technical complexity into concise business impact and recovery updates.
- Uses data and engineering evidence to prioritise reliability improvements and challenge low-value operational practices.
- Demonstrates ownership while fostering a blameless, learning-focused culture.
- Builds trusted partnerships across global, cross-functional and vendor teams.
- Mentors’ engineers and contributes reusable standards, patterns and knowledge.

📌 Senior Engineer II - Data Science / MLOps-29625] (Hyderabad)
🏢 Marriott Tech Accelerator
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior engineer ii - data science / mlops-29625] (hyderabad) / hyderabad