13 Sep
|
Techdome
|
Hyderabad
13 Sep
Techdome
Hyderabad
About The Role Techdome runs live infrastructure for Healthcare, FinTech, AI, and SaaS products — environments where downtime isn't an inconvenience, it's a compliance incident or a lost transaction. We're hiring a Senior SRE who treats uptime as a personal metric, not a team KPI, and who wants direct, high-leverage ownership over production systems handling real money and real patient data — not staging environments.
You'll report close to founders and senior engineering leadership, and your infrastructure decisions will ship the same week you make them.
Key Responsibilities
- Define SLIs/SLOs, own the error budget, and decide when to slow down shipping to protect it
- Build zero-downtime CI/CD pipelines supporting Blue-Green, Canary, and Rolling releases
- Manage all environments through Terraform and Ansible — no manual console changes
- Instrument systems with Prometheus, Grafana, ELK, Datadog, and OpenTelemetry so alerts are actionable, not noisy
- Lead incident response, drive root cause analysis, and ensure postmortem action items are closed
- Right-size infrastructure and forecast capacity proactively, rather than reacting to billing
- Apply AI to alert triage, incident summarization, and automated runbooks
- Participate in a shared on-call rotation as a dependable, trusted responder
Required Qualifications
- 2+ years of production ownership experience as an SRE, DevOps, Platform,
or Cloud Engineer
- Hands-on, production-grade experience with AWS, Azure, or GCP
- Real-world Docker/Kubernetes experience under production load
- Daily use of Terraform and Ansible (or equivalent IaC tools)
- Experience building at least one CI/CD pipeline from scratch (Jenkins, GitHub Actions, GitLab CI, or similar)
- Solid fundamentals in Linux, networking, and distributed systems
- Proficiency in Python, Go, or Bash for automation and scripting
- Proven experience shipping Blue-Green, Canary, and Rolling deployments in production
- Prior work in FinTech, Payments, Healthcare, or another high-availability domain
- Practical, everyday use of AI tools such as Copilot, Claude, Cursor, or ChatGPT
Preferred Qualifications
- Experience building AI-powered operations tooling — triage bots, incident auto-summarization, or reliable anomaly detection
- Fluency in SLOs, error budgets, and chaos engineering as core practice, not theory
Why Join Techdome?
- Direct reporting line to founders and senior engineering leadership
- Infrastructure decisions ship the same week they're made — no bureaucratic delay
- You'll protect systems handling real healthcare data and real financial transactions
- A team that shares on-call and builds a culture where the SRE on call at 2am is someone people trust
📌 Site Reliability Engineer (SRE) (Hyderabad)
🏢 Techdome
📍 Hyderabad