Site Reliability Engineer - Long Term Contract, Hybrid opportunity (Bengaluru)

Site Reliability Engineer - Long Term Contract, Hybrid opportunity (Bengaluru)

17 Sep
|
World Wide Technology
|
Bengaluru

17 Sep

World Wide Technology

Bengaluru

Worldwide Technology (WWT), a 36-year-old global technology solutions provider specializing in systems integration, Infra-Cloud security, application development, AI Services, and supply chain solutions. With a workforce of 10,000+ employees and strategic partnerships with leading OEMs such as Cisco, Dell EMC, Microsoft, and NVIDIA, WWT delivers cutting-edge infrastructure, cloud, security, and custom application services to clients across 35 countries.

Our Advanced Technology

Centers (ATCs)lab setup environments spanning over one million square feet of world-class integration and distribution space enable us to deliver unmatched value and innovation at scale. Recognized as one of the Best Places to Work by Glassdoor and Fortune for 14 consecutive years, WWT is also ranked #6 on Indias Great Place to Work list for 2025.

WorldWide Technology Holding Co, LLC. (WWT) currently has an exciting opportunity available for the role of Site Reliability Engineer (operations)- Hybrid -Long Term Contract Opportunity. If you are interested in this opportunity, please respond with an updated resume to [email protected]

This is a contract Role & we use a global payroll partner to manage payroll for all contract opportunities with us. A proper background verification of past employments, criminal checks would be done for onboarding to any opportunity with WWT.

Role Site Reliability Engineer (operations)

Location Bangalore ( 3 days office, 2 days WFH)

Duration 12 – Months (can be extended)

Work hrs: as below , Resource will work both shifts, when going to work 3 days a week he will login from office in the morning and home in the evening.

- First Shift (Independent) : 9:30 AM – 1:00 PM (IST)
- Second Shift (Overlap) : 7:00 PM – 11:00 PM (IST)

Summary

- The LLM Proxy is WWT's internal AI gateway a centralized platform that provides a single access point for large language models from multiple cloud providers, including Microsoft Azure, AWS, Google Cloud, and OpenAI. It delivers centralized governance, reporting, performance-based routing,



and automatic failover between providers.
- As AI adoption expands across Ciscos Collab business unit, the LLM Proxy has become the mandated interface for collab when building AI-powered features in production. Usage and visibility are growing rapidly, which means more incidents.

Responsibilities Monitoring & Observability

- Own day-to-day platform monitoring using Prometheus and Grafana
- Write and maintain PromQL queries to track service health, latency, throughput, and error rates
- Build and maintain dashboards for engineering teams and leadership
- Configure and tune alerts to reduce noise while catching real incidents early
- Track and report against SLOs and SLAs
- Incident Response
- Act as the primary incident communicator during platform disruptions
- Translate technical findings into clear, concise language for non-technical stakeholders
- Coordinate with upstream cloud providers (Azure, AWS, GCP) during third-party outages
- Maintain incident logs and produce post-incident reports and root cause analyses

Requirements

Experience

- 3+ years in an SRE role
- Proven track record in an incident response
- Experience supporting production services at scale on a major cloud provider (Azure, AWS, or GCP)
- Experience working with API gateway or proxy infrastructure is a plus

Must-Have Skills

- Prometheus & Grafana — writing PromQL queries from scratch, building dashboards, and interpreting results
- Incident communication — structured, calm, and transparent written and verbal updates under pressure
- Observability fundamentals — metrics, logs, and traces; understanding the difference and when to use each
- Cloud fluency — working knowledge of at least one major provider's managed services and common failure modes
- Documentation — ability to write clear runbooks, incident reports, and post-mortems

Common Tools & Technologies

- Observability: Prometheus, Grafana, PromQL
- Cloud: Microsoft Azure (primary), AWS, GCP
- Incident management: PagerDuty or equivalent
- Logging: any major log aggregation platform (Loki, Splunk, ELK, etc.)
- Ticketing & communication: Jira, Confluence, Slack or Teams

📌 Site Reliability Engineer - Long Term Contract, Hybrid opportunity (Bengaluru)
🏢 World Wide Technology
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer - long term contract, hybrid opportunity (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer - long term contract, hybrid opportunity (bengaluru) / bengaluru