Site Reliability Engineer (SRE) (Hyderabad)

Site Reliability Engineer (SRE) (Hyderabad)

15 Sep
|
Bahwan Cybertek
|
Hyderabad

15 Sep

Bahwan Cybertek

Hyderabad

Bachelor s degree in Computer Science, Engineering, or equivalent practical experience (preferred).

5+ years experience in reliability engineering, DevOps, production operations, or software engineering with on-call responsibilities (preferred).

Demonstrated experience improving production reliability through automation, monitoring, and incident/problem management.

Required

Strong grounding in SRE/DevOps practices: incident management, blameless postmortems, SLOs/SLIs, error budgets, production readiness.

Experience building/operating monitoring and ing, and using logs/metrics to diagnose issues.

Automation/scripting skills (e.g., Python, PowerShell, Bash) and ability to reduce manual operational work.

Strong understanding of cloud-based platforms such as Azure DataBricks + Unity Catalog, AWS S3 and RDS.

Solid experience in ETL / ELT work.

Understanding of CI/CD concepts, safe deployment patterns, rollback strategies, and change risk controls.

Preferred

Experience with cloud environments and infrastructure-as-code.

Experience with large datasets (Multi-million row datasets).

Experience with container orchestration and modern runtime platforms (where applicable).

Experience building dashboards and reliability reporting for executives and delivery teams.

The Site Reliability Engineer (SRE) is responsible for improving the reliability, availability, performance, and operability of PAH-supported software systems. This role combines software engineering and IT operations to automate operational work, monitor system performance,



and reduce toil. The SRE establishes and manages monitoring, ing, incident response, and problem management practices to ensure applications remain available and performant during updates and failures.

The role partners with engineering, architecture, and product teams to define reliability standards and production readiness requirements. SRE is a practical implementation of DevOps focused on maintaining software quality in fast-paced development environments.

Define, implement, and maintain observability (monitoring, logging, tracing) and actionable ing aligned to service health.

Drive incident management: on-call readiness, triage, incident command support, communications, and post-incident reviews (RCA)

Reduce operational toil through automation (runbooks-to-automation, self-healing, deployment/rollback automation).

Establish reliability standards: SLOs/SLIs, error budgets, production readiness reviews, and release risk controls

Performance and reliability engineering: capacity planning, load/performance analysis, resilience testing, and failure-mode mitigation

Partner with engineering teams to improve operational hygiene (deployability, rollback strategy, configuration, secrets, dependency management)

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

📌 Site Reliability Engineer (SRE) (Hyderabad)
🏢 Bahwan Cybertek
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (sre) (hyderabad) / hyderabad