- 5+ years of hands-on Kubernetes platform engineering experience.
- Cluster design, deployment, administration, and troubleshooting.
- Platform upgrades, scaling, and lifecycle management.
- Kubernetes security and multi-tenancy.
- High Availability (HA) and disaster recovery design.
Data Platform Technologies
- Robust experience with Stackable Data Platform or comparable Kubernetes-native data platforms.
- Apache Spark (PySpark and/or Scala).
- Spark SQL and performance tuning.
- Apache Kafka administration and operations.
- Real-time and batch data processing architectures.
Storage
- S3-compatible Object Storage (MinIO, Ceph, FlashBlade).
- Object storage performance optimisation.
- Data lifecycle management, backup, and recovery.
- 6+ years of overall platform engineering experience.
- 5+ years operating Kubernetes in production environments.
- Experience supporting mission-critical enterprise data platforms.
- Experience delivering production services in hybrid cloud environments.
What Success Looks Like
The successful candidate will demonstrate the ability to:
Independently build and operate Kubernetes clusters
Deploy and manage Stackable platform services
Scale Spark and Kafka workloads for enterprise use cases
Design resilient object storage integrations
Automate platform operations using GitOps and Infrastructure as Code
Troubleshoot complex platform incidents and perform effective RCA
Drive platform reliability, security, and operational excellence