10 Aug
|
Scientific Games Technologies
|
Bengaluru
10 Aug
Scientific Games Technologies
Bengaluru
Job Summary
Scientific Games is seeking an experienced Senior Data Platform Reliability Engineer to ensure the reliability, resilience, observability, performance, and operational excellence of our enterprise data platform built on Databricks and AWS.
Working closely with the Principal Analytics Engineer, Principal Enterprise Data Architect, and engineering teams, this role is responsible for engineering a highly available, secure, and resilient platform that supports business-critical data services and analytics. The role drives operational excellence through automation, proactive monitoring, performance optimization, incident engineering, and continuous improvement.
This is a highly hands-on engineering role. The successful candidate is expected to design, automate, monitor, troubleshoot, and continuously improve production platform services while leading reliability engineering initiatives across the Data Platform.
Success requires deep expertise in Databricks, AWS, cloud operations, observability, disaster recovery, and platform reliability engineering, along with the ability to influence engineering practices through technical leadership.
Mission
Engineer operational excellence by ensuring the Databricks and AWS Data Platform is reliable, observable, resilient, secure, performant, scalable, and cost-efficient.
Scope
Owns the operational reliability of the enterprise Data Platform built on Databricks and AWS.
Success is measured by:
- Platform availability and SLO/SLA compliance
- MTTR, MTBF, RTO and RPO achievement
- Platform health and reliability
- Operational automation
- Platform cost efficiency
• Reduced incidents and improved developer experience
Job Duties / Key Accountabilities
Reliability Engineering
- Engineer highly available and resilient platform services.
- Define and improve SLOs, SLAs, and error budgets.
- Eliminate single points of failure.
• Continuously improve resilience through automation.
Observability Engineering
- Engineer monitoring, logging, metrics, traces, dashboards, and alerting.
- Develop platform health dashboards and proactive alerting.
• Improve diagnostics and operational visibility.
Incident Engineering
- Lead Data Platform incident response.
- Develop operational runbooks and recovery procedures.
• Perform root cause analysis and automate recovery where practical.
Performance & Capacity Engineering
- Optimize Databricks clusters, Spark workloads, SQL Warehouses, storage, and compute.
- Engineer capacity planning and growth forecasting.
• Continuously optimize performance and cloud cost.
Databricks & AWS Platform Operations
- Operate Databricks for workspaces, clusters, jobs, workflows, SQL Warehouses, and Unity Catalog operational components.
- Operate AWS services including IAM, S3, CloudWatch, CloudTrail, VPC, KMS, Secrets Manager, networking, and storage.
• Engineer automation that reduces manual operational effort.
FinOps
- Optimize cloud consumption and cluster utilization.
- Engineer autoscaling, cluster policies, and cost optimization.
• Build operational cost dashboards and reporting.
Disaster Recovery & Business Continuity
- Design and maintain multi-region disaster recovery capabilities.
- Engineer backup, replication, failover, and recovery automation.
- Define and validate RTO and RPO objectives.
• Conduct regular disaster recovery exercises and maintain recovery runbooks.
Technical Leadership
- Lead reliability engineering reviews.
- Mentor engineers on SRE and operational excellence.
- Partner with the Principal Data Engineer and Principal Enterprise Data Architect.
• Promote automation and continuous improvement.
Qualifications / Skills / Knowledge
Required
- Bachelor's degree in Computer Science, Engineering, Information Systems, or related discipline.
- 8+ years in Site Reliability Engineering, Platform Engineering, Cloud Operations, or DevOps.
- 4+ years operating Databricks platforms in production.
- Strong AWS operational experience including IAM, S3, CloudWatch, CloudTrail, VPC, KMS, Secrets Manager, and networking.
- Solid Python and scripting experience.
- Strong Spark and Databricks operational knowledge.
- Experience with monitoring and observability platforms.
- Experience with Terraform and Infrastructure as Code.
- Experience operating highly available multi-region AWS platforms.
- Experience with disaster recovery, business continuity, backup, failover, and operational automation.
• Strong troubleshooting and incident management skills.
Desired
- Databricks Certification.
- AWS Certified DevOps Engineer or Solutions Architect.
- Kubernetes experience.
• Experience in regulated industries.
Authority / Decision Making
Authority To
- Define operational engineering standards.
- Recommend reliability improvements.
- Define monitoring, alerting, and automation standards.
• Drive platform operational excellence.
Requires Approval For
- Architecture changes outside approved standards. • Technology adoption outside the approved platform strategy.
Key Contacts
- Head of AI, Data & Infrastructure
- Principal Analytics Engineer
- Principal Enterprise Data Architect
- Platform Engineering
- Governance Engineering
- Product Leadership
- Databricks
• AWS
Language Skills
Required: English
Desired: Additional languages considered an asset.
Job Conditions
- Remote position based in India.
- Regular collaboration across India and North America.
- Flexible schedule with recurring North America overlaps.
- Occasional international travel (up to 10%).
• Participation in operational reviews, architecture reviews, disaster recovery testing, incident response, sprint planning, and technical planning.
📌 Senior Data Platform Reliability Engineer (Bengaluru)
🏢 Scientific Games Technologies
📍 Bengaluru