Senior Site Reliability Engineer (E&P) (Bengaluru)

Senior Site Reliability Engineer (E&P) (Bengaluru)

10 Sep
|
Cisco ThousandEyes
|
Bengaluru

10 Sep

Cisco ThousandEyes

Bengaluru

Senior Site Reliability Engineer — Efficiency & Performance / AWS Cost Optimization / FinOps

Meet the Team The ThousandEyes Efficiency & Performance team is responsible for improving the reliability, scalability, performance, and cost efficiency of the ThousandEyes cloud platform. The team partners closely with engineering, platform, finance, product, and leadership stakeholders to drive infrastructure optimization, cloud cost governance, performance improvements, and operational excellence across AWS environments.

As a Senior Site Reliability Engineer in the Efficiency & Performance team, you will play a key role in managing and optimizing ThousandEyes cloud infrastructure cost, improving AWS efficiency, driving FinOps practices, and helping engineering teams make data-driven decisions around cost, performance, and scalability.

Your Impact

- Own and drive AWS cost optimization and efficiency initiatives across ThousandEyes infrastructure.
- Analyze cloud usage across compute, storage, databases, observability, networking, and data platform workloads to identify savings opportunities.
- Partner with engineering teams to improve application and infrastructure performance while reducing cloud wastage.
- Drive FinOps practices including cost visibility, tagging hygiene, showback/chargeback, budgeting, forecasting, anomaly detection, and cost allocation.
- Improve infrastructure efficiency through better resource utilization, autoscaling, right-sizing, and capacity planning.
- Build automation, dashboards, reports, and guardrails to improve cost governance and operational visibility.
- Support ThousandEyes cost reviews, OKR tracking, leadership updates, and stakeholder communications.
- Collaborate with product, finance, engineering, and platform teams to align cost optimization with business priorities.
- Identify and reduce underutilized infrastructure, idle resources, over-provisioned workloads, and inefficient service usage.
- Improve observability, alerting, and monitoring for cost, performance, and platform health indicators.
- Influence engineering teams to adopt cost-aware architecture, reliable design patterns, and performance-efficient implementation practices.
- Provide senior-level technical leadership, mentor engineers, and independently drive cross-functional initiatives to closure.
- Balance reliability, scalability, performance, and cost efficiency in all platform decisions.

Minimum Qualifications

- Bachelor’s degree or higher in Engineering, Computer Science, or equivalent practical experience.
- 8–12 years of relevant experience in Site Reliability Engineering, Cloud Infrastructure, Platform Engineering, DevOps, Production Engineering, or Performance Engineering.
- Strong hands-on experience with AWS services such as EC2, S3, RDS, EMR, Lambda, CloudWatch,



OpenSearch, ElastiCache, IAM, VPC, and related cloud-native services.
- Practical experience in cloud cost optimization, FinOps, AWS billing analysis, cost allocation, tagging, budget tracking, forecasting, and cost governance.
- Experience with performance analysis, capacity planning, infrastructure optimization, and reliability improvements for large-scale cloud platforms.
- Experience with infrastructure-as-code and automation tools such as Terraform, CloudFormation, Puppet, Ansible, or similar.
- Strong scripting or programming skills in Python, Go, Shell, or similar languages for automation, reporting, and operational tooling.
- Experience with observability and monitoring platforms such as ThousandEyes, CloudWatch, Prometheus, Grafana, Splunk, OpenSearch, Datadog, or similar.
- Robust Linux systems knowledge, troubleshooting skills, and understanding of distributed systems.
- Experience in incident management, production support, reliability engineering, and operational excellence.
- Ability to analyze large-scale infrastructure, performance, and cost data and convert findings into actionable engineering recommendations.
- Strong communication skills with the ability to present technical, performance, and cost insights to engineering teams, finance, leadership, and cross-functional stakeholders.
- Experience working in Agile/Scrum environments and managing priorities across multiple teams.

Preferred Qualifications

- Experience supporting or optimizing large-scale SaaS platforms in AWS.
- FinOps certification or equivalent hands-on experience with cloud financial management practices.
- Experience with AWS Savings Plans, Reserved Instances, Spot adoption, Graviton migration, storage lifecycle optimization, workload right-sizing, and commitment planning.
- Experience with cloud cost management tools such as AWS Cost Explorer, AWS CUR, Cloudability, CloudHealth, or similar platforms.
- Experience with performance tuning of cloud infrastructure, distributed systems, databases, data pipelines, and high-scale services.
- Experience with data platforms or large-scale infrastructure components such as EMR, Kafka, OpenSearch, RDS, Airflow, Spark, Redis, Cassandra, or similar.
- Experience driving cross-team cost optimization programs, efficiency OKRs, governance reviews, executive reporting, and measurable savings outcomes.




- Ability to influence engineering teams toward cost-conscious architecture, performance-aware design, and operational best practices.
- Exposure to AI/ML infrastructure cost optimization, capacity planning, GPU/accelerator cost governance, or AI-driven efficiency tooling is a plus.

What Success Looks Like

- Improved AWS cost visibility, governance, and accountability across ThousandEyes engineering teams.
- Measurable reduction in cloud wastage and improved infrastructure utilization.
- Improved infrastructure efficiency, capacity planning, and resource optimization.
- Better performance visibility across critical ThousandEyes services and infrastructure.
- Stronger FinOps practices including forecasting, anomaly detection, tagging hygiene, and cost allocation.
- Clear leadership reporting on cost trends, performance risks, savings opportunities, and execution progress.
- Reliable, scalable, performant, and cost-efficient ThousandEyes cloud infrastructure.

Hello Team,

Sharing the updated interview panel guide for the Senior Site Reliability Engineer – Efficiency & Performance role.

The primary focus areas for this position are AWS cost optimization, FinOps, performance engineering, observability, infrastructure efficiency, SRE practices, and operational excellence.

Interview Round

Panel Member(s)

Focus Area

Format

Hiring Manager

Deepak

Role alignment, ownership, leadership, communication, and cross-functional collaboration

Virtual / In-person

SRE Technical / System Design/ FinOps / Cloud SME

Gururaj — [email protected]

AWS technical depth, FinOps, cost optimization, automation, capacity planning, and cloud governance

In-person

Performance / Observability / Infrastructure

Prathik — [email protected][email protected]

Performance troubleshooting, observability, scalability, capacity planning, and infrastructure efficiency

In-person

Behavioural / Project Retrospective

Diego — [email protected]

Collaboration, stakeholder management, project ownership, conflict handling, and lessons learned

Virtual — UK

SRE Mindset & Operational Excellence

Joao — [email protected]

Incident management, SLOs, reliability practices, toil reduction, automation, and continuous improvement

Virtual — UK

Please assess candidates based on practical experience, technical depth, individual ownership, measurable outcomes, and their ability to balance reliability, performance, scalability, and cost.

As Kubernetes/EKS is not a core requirement in the updated JD, it should not be treated as a mandatory hiring criterion. The relevant round should primarily focus on performance, observability, and infrastructure architecture.

Please share clear strengths, concerns, supporting evidence, and an overall recommendation after each round.

-

📌 Senior Site Reliability Engineer (E&P) (Bengaluru)
🏢 Cisco ThousandEyes
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior site reliability engineer (e&p) (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: senior site reliability engineer (e&p) (bengaluru) / bengaluru