18 Sep
|
Bounteous
|
Bengaluru
18 Sep
Bounteous
Bengaluru
About the Role
We are seeking an experienced Site Reliability Engineer to join our infrastructure team. This role combines core SRE responsibilities reliability, observability, and incident management — with hands-on ownership of our Grafana-based monitoring stack and quality assurance automation using Selenium or Playwright. The ideal candidate bridges the gap between operational excellence and product quality, ensuring our systems are both reliable and thoroughly tested end-to-end.
Key Responsibilities
Site Reliability & Operations
- Design, build, and maintain highly available, scalable, and fault-tolerant production systems
- Define and track SLIs, SLOs, and error budgets across services
- Lead incident response, root cause analysis, and post-mortems for production issues
- Automate operational tasks through infrastructure as code (Terraform, Ansible, or similar)
- Participate in an on-call rotation and drive continuous improvement in mean-time-to-detect (MTTD) and mean-time-to-resolve (MTTR)
- Collaborate with development teams to improve deployment pipelines, capacity planning, and system performance
Observability & Grafana
- Design, implement, and maintain Grafana dashboards, alerting rules, and visualizations across metrics, logs, and traces
- Integrate Grafana with data sources such as Prometheus, Loki, Tempo, InfluxDB, or Elasticsearch
- Build meaningful, actionable alerts that reduce noise and improve signal for on-call engineers
- Establish observability standards and best practices across teams
- Continuously tune dashboards and alert thresholds based on production learnings
QA Automation
- Design and maintain automated test suites using Selenium or Playwright for functional, regression, and end-to-end testing
- Integrate automated tests into CI/CD pipelines to enable fast, reliable release cycles
- Collaborate with QA and development teams to identify gaps in test coverage
- Build reusable, maintainable test frameworks and contribute to test strategy
- Analyze test results and flaky test patterns to improve overall test suite reliability
Required Qualifications
- 5+ years of experience in a Site Reliability Engineering, DevOps, or Infrastructure Engineering role
- Hands-on experience designing and maintaining Grafana dashboards and alerting (Prometheus, Loki, or similar observability stack)
- Practical experience with test automation using Selenium or Playwright
- Strong proficiency with Linux systems administration and networking fundamentals
- Experience with cloud platforms (AWS, GCP, or Azure)
- Proficiency in at least one scripting/programming language (Python, Go, JavaScript/TypeScript, or similar)
- Experience with containerization and orchestration (Docker, Kubernetes)
- Solid understanding of CI/CD pipelines and tools (Jenkins, GitLab CI, GitHub Actions, or similar)
- Familiarity with infrastructure as code tools (Terraform, Ansible, CloudFormation)
- Strong troubleshooting skills and experience with incident management processes
Preferred Qualifications
- Experience with log aggregation and tracing tools (Loki, Tempo, Jaeger, ELK stack)
- Exposure to chaos engineering or reliability testing practices
- Experience writing custom Grafana plugins or advanced PromQL queries
- Familiarity with API testing tools (Postman, REST Assured)
- Relevant certifications (AWS/GCP/Azure, CKA, or similar)
- Experience mentoring junior engineers or leading small technical initiatives
Soft Skills
- Strong communication skills to work cross-functionally with development, product, and QA teams
- Ability to balance reliability priorities with delivery timelines
- Comfortable working independently and taking ownership of complex systems
- Analytical mindset with a proactive approach to identifying and preventing issues
What We Offer
- Market-competitive salary and performance bonuses
- Comprehensive health, dental, and vision benefits
- Flexible/remote work options
- Professional development budget and certification support
- Collaborative, engineering-driven culture
📌 QA + SRE (Bengaluru)
🏢 Bounteous
📍 Bengaluru