Position Overview
- We are seeking a highly technical Site Reliability Engineer with deep expertise in Splunk and Grafana to own and evolve our observability ecosystem. In this role, you will move beyond easy monitoring to architect a comprehensive, scalable telemetry platform. You will be our subject-matter expert in Splunk optimisation, ensuring our logging architecture is performant, cost-effective, and deeply integrated with our automated workflows.
- You will treat infrastructure as codeutilising Terraform and strong coding proficiency in Go, Python, or Rubyto automate the deployment of agents and collectors across complex distributed systems.
Key Responsibilities
- Splunk Architecture Optimisation: Lead the design and tuning of Splunk environments. Optimise indexer performance, search efficiency, and data models to ensure rapid troubleshooting and cost-efficiency.
- Advanced Visualisation: Architect and maintain sophisticated Grafana dashboards that correlate disparate data sources into a single pane of glass for real-time system health.
- Automated Infrastructure: Design,
build, and maintain scalable observability infrastructure using tools like Terraform.
- Pipeline Engineering: Optimise the collection, processing, and storage of telemetry data (Metrics, Logs, Traces) to ensure high reliability and low latency.
- Workflow Automation: Develop custom Splunk workflows and integrations that trigger automated responses to system events, reducing Mean Time to Resolution (MTTR).
- Incident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements through "observability-driven development."
Required Skills Experience (The Essentials)
- Splunk Mastery: Deep, hands-on experience with Splunk administration, search optimisation (SPL), and architecting complex data pipelines. You know how to make Splunk "hum" at scale.
- Grafana Expertise: Proven ability to build actionable, intuitive dashboards in Grafana that go beyond