03 Sep
|
Ironbook AI
|
Bengaluru
03 Sep
Ironbook AI
Bengaluru
We are seeking a high-caliber Site Reliability Engineer (SRE) with deep expertise in Kubernetes and Infrastructure-as-Code (IaC) to own, operationalize, and scale our Crossplane-backed control planes. In this role, you will bridge the gap between application engineering and cloud infrastructure by building self-service API primitives, maintaining state reconciliation, and running Crossplane as a mission-critical, high-availability service in production.
Key Responsibilities1.
Crossplane
Architecture &
• Management
- Deploy & • Operate Crossplane:
Install, configure, upgrade, and maintain core Crossplane controllers, Custom Resource Definitions (XRDs), and Compositions across production Kubernetes clusters.
- Provider Lifecycle Management: Manage and optimize Crossplane Providers (AWS, GCP, Azure, Upbound, etc.), ensuring secure credential rotation, proper scope limiting, and efficient resource usage.
- Composition Engineering: Author and optimize modular infrastructure Compositions and Composition Functions (Go, Python, KCL) to package cloud resources into clean, developer-friendly abstractions.
- State & • Reconciliation Management:
Monitor and maintain the Crossplane reconciliation loop, resolving API rate-limiting issues, drift anomalies, and cascading resource dependency failures.
2. Reliability, Operations & • Observability
- Cluster & • Control Plane Reliability:
Maintain high availability, scalability, and disaster recovery strategies for the Kubernetes clusters hosting Crossplane instances.
- SLO/SLI & • Error Budgets:
Define and enforce Service Level Objectives (SLOs) for infrastructure provisioning velocity, API availability, and reconciliation latency.
- Observability & • Alerting:
Build robust telemetry around Crossplane controllers using Prometheus, Grafana, and OpenTelemetry to track metric indicators (e.g., reconciliation loop duration, provider API failure rates, queue depths).
- Incident Response & • Support:
Participate in an on-call rotation for platform outages, conduct root cause analyses (RCAs), and act as Tier-3 support for developer teams consuming Crossplane claims.
3. GitOps, Security & • Developer Enablement
- GitOps Integration: Drive continuous deployment of Crossplane configurations using GitOps tools (ArgoCD or Flux).
- RBAC & • Security Governance:
Implement multi-tenant RBAC policies separating Platform Builder permissions from Platform Consumer claims, integrating native cloud IAM roles securely.
- Developer Self-Service: Partner with product teams to gather requirements for self-service cloud infrastructure (e.g., databases, queues, networking) and eliminate manual ticketing/toil.
Qualifications & • RequirementsTechnical Skills (Must-Have)
- Kubernetes Expertise (3+ years): Deep operational experience with Kubernetes internals, custom controllers, CRD design, RBAC, API extension mechanisms, and cluster administration.
- Crossplane Production Experience (1–2+ years): Hands-on experience building, deploying, and supporting Crossplane XRDs, Compositions, Compositions Functions,
and managing ProviderConfigs in production environments.
- Cloud Infrastructure (AWS, GCP, or Azure): Robust background in public cloud architectures and cloud provider APIs (S3, RDS, IAM, VPCs, Managed Kubernetes, etc.).
- Programming & • Scripting:
Proficiency in Go (for controller debugging/writing composition functions) and Python/Bash for tooling and automation.
- GitOps & • Infrastructure-as-Code:
Demonstrated experience with GitOps controllers (ArgoCD or Flux) and traditional IaC tools (Terraform/Pulumi) for migration or hybrid strategies.
Professional Experience & • Soft Skills
- SRE / DevOps Background: 4+ years in an SRE, DevOps, or Platform Engineering role owning production-critical infrastructure.
- Observability Fluency: Real-world experience building dashboards and tuning alerts with Prometheus, Grafana, Datadog, or similar.
- Systems Mindset: Strong understanding of distributed systems, operational toil reduction, and disaster recovery principles.
Nice-to-Have Skills
- CNCF Certifications: CKA (Certified Kubernetes Administrator) or CKAD.
- Upbound Universal Crossplane (UXP) enterprise experience.
- Knowledge of dynamic schema languages like KCL or CUE for Crossplane functions.
What Success Looks Like (First 90 Days)
- Day 30: Audit current Crossplane deployments, audit RBAC patterns, and baseline existing reconciliation bottlenecks and alert noise.
- Day 60: Harden Crossplane controller observability with dedicated Prometheus metrics and establish clear SLOs for platform claim fulfillments.
- Day 90: Refactor legacy Compositions into modular, version-controlled composition functions and automate provider credential rotations end-to-end.
📌 Senior Site Reliability Engineer (Platform & Crossplane) (Bengaluru)
🏢 Ironbook AI
📍 Bengaluru