02 Sep
|
Infinx
|
Bengaluru
INFINX
Senior Site Reliability Engineer
Hybrid Cloud: AWS / IBM Cloud • Experience: 4–6 years • Bengaluru (Hybrid) • Full-time
Company Overview
Infinx Healthcare is a leading healthcare payment solutions provider specializing in next-generation, cloud-based SaaS products. Our mission is to maximize and preserve revenue across the US healthcare revenue cycle by combining human expertise with artificial intelligence.
Our platform carries critical clinical transactions for providers across the United States. Reliability engineering sits close to the centre of our engineering organisation rather than at its edge.
To learn more, visit www.infinx.com.
About the Role
We are looking for a Senior Site Reliability Engineer to own the availability and performance of production systems that span AWS, IBM Cloud, etc.. This is a hands-on engineering role rather than a monitoring rotation. You will write the automation, define the service levels, and change the architecture that keeps recurring incidents from recurring.
Workloads run on managed Kubernetes in AWS and on Red Hat OpenShift in IBM Cloud. Your job is to make that estate behave like one system: consistent service levels, one observability pane, one incident process, and infrastructure defined as code no matter which provider sits underneath.
What You Will Own
Reliability and Service Levels
- Own SLIs, SLOs, and error budgets for a defined set of critical services, negotiated with product and service owners rather than imposed on them.
- Use error-budget burn to make real prioritisation calls, including slowing a rollout when the budget is spent.
- Run capacity and headroom planning across cloud and on-premises resources, where procurement lead times differ by months.
- Push reliability improvements upstream into application design reviews instead of absorbing them at the infrastructure layer.
Incident Response and Learning
- Act as incident commander for high-severity events, coordinating triage, escalation, customer communication,
and resolution.
- Facilitate blameless postmortems and make sure follow-up actions are assigned, scheduled, and actually closed.
- Reduce time to detection by improving signal quality. Every alert should be actionable, urgent, and owned by someone.
- Publish reliability metrics and incident trends to engineering leadership, with honest commentary rather than green dashboards.
Hybrid Cloud and Platform Operations
- Operate production Kubernetes across Amazon EKS and Red Hat OpenShift on IBM Cloud, including version upgrades, node lifecycle, and cluster hardening.
- Own hybrid connectivity: AWS Direct Connect, IBM Cloud Direct Link, site-to-site VPN, transit routing, splithorizon DNS, and cross-environment identity.
- Maintain Terraform as the single source of truth across providers, covering module design, state isolation, drift detection, and safe promotion paths.
- Design and prove failover and disaster-recovery paths that cross provider boundaries, and validate them through regular game days.
- Advise on workload placement against cost, latency, data residency, and compliance constraints.
Observability and Automation
- Build a federated observability layer that gives one view across both clouds, based on OpenTelemetry using
Prometheus, Thanos, Grafana, etc.
- Standardise instrumentation so services emit comparable telemetry regardless of where they run.
- Write production-grade Python and Bash to remove toil through auto-remediation, self-service tooling, and safe runbook automation.
- Set and defend a toil budget. If the team spends more than its target share of time on manual operations,
reducing that becomes the work.
Security and Compliance
- Operate within HIPAA and SOC 2 expectations: least-privilege access, key management, audit trails, and evidence that holds up under review.
- Contribute to cloud security posture management and vulnerability remediation across both cloud providers.
- Enforce secure-by-default platform patterns including network segmentation, secrets management, admission control, and image provenance.
Technical Leadership
- Mentor engineers on reliability practice and review their designs and change plans.
- Write and maintain the standards others build against, and document decisions so they outlive any one engineer.
- Represent reliability in architecture forums, vendor conversations, and audit discussions.
Must-Have Skills
Experience
- 4–6 years in site reliability engineering, DevOps, or production cloud infrastructure, including at least three years carrying on-call responsibility for systems you helped build.
Hybrid and Multi-Cloud
- Deep, hands-on production experience with AWS and preferably IBM Cloud.
- Practical experience running workloads across more than one environment, including the networking and identity glue between them.
Containers and Platform
- Production ownership of Kubernetes beyond day-to-day kubectl: upgrades, capacity, networking, storage, and debugging failures under load.
- Docker and Helm. OpenShift experience is strongly preferred given our IBM Cloud footprint.
Reliability Engineering
- Demonstrated use of SLIs, SLOs, and error budgets to drive decisions, with specific examples you can walk through.
- Incident command experience on high-severity, customer-visible outages.
Automation and Infrastructure as Code
- Strong Python and Bash for operational automation, at a standard you would trust to run unattended in production.
- Terraform in a real team setting: modules, state management, peer review, and blast-radius control. Ansible is a welcome addition.
Observability
- Hands-on with Prometheus and Grafana, plus at least one of ELK/EFK, OpenTelemetry, Datadog, CloudWatch, or
IBM Cloud Monitoring.
- Ability to design alerting that engineers trust, and the judgement to delete alerts they do not.
Networking
- Solid command of routing, DNS, TLS, load balancing, VPNs, and private interconnects, with the ability to debug a hybrid path methodically rather than by guesswork.
Security and Compliance
- Comfort operating in a regulated environment, with HIPAA awareness or equivalent exposure to SOC 2,
HITRUST, PCI DSS, or ISO 27001.
Ways of Working
- Clear written communication across postmortems, design documents, and live incident updates.
- Sound judgement under pressure, including the discipline to slow down when slowing down is the safer call.
Good-to-Have Skills
- Red Hat OpenShift certification, or IBM Cloud Skilled Architect or SRE certification.
- AWS Professional or Specialty certifications.
- Go for tooling, plus Kubernetes operator or controller development.
- AIOps and automated remediation at scale, including anomaly detection on production telemetry.
- Service mesh such as Istio or Linkerd, and progressive delivery with Argo Rollouts or Flagger.
- GitOps practice with Argo CD or Flux.
- FinOps discipline: cross-cloud cost attribution, egress optimisation, and commitment planning.
- Chaos engineering and structured game-day programmes.
- Experience migrating workloads between clouds, or repatriating them to on-premises.
- Background in healthcare, fintech, or another regulated domain.
Our Tech Stack
- Cloud: AWS (EKS, EC2, RDS, S3, VPC, Transit Gateway, Direct Connect) and IBM Cloud (Red Hat OpenShift on IBM
Cloud, VPC, Direct Link, Cloud Object Storage, Key Protect).
- Containers and platform: Kubernetes, Red Hat OpenShift, Docker, Helm, Argo CD.
- Observability: Prometheus, Thanos, Grafana, OpenTelemetry, ELK/EFK, Datadog, IBM Cloud Monitoring.
- Infrastructure as code: Terraform (primary), Ansible, Git-based workflows.
- Languages: Python, Bash
- Incident and delivery: PagerDuty, Opsgenie, Jira, Jenkins.
- On-premises: VMware vSphere, bare metal, hybrid DNS, and identity federation
📌 Senior Site Reliability Engineer (Bengaluru)
🏢 Infinx
📍 Bengaluru