What You’ll Work On
• Architect and scale Search & NoSQL platforms including Solr, HBase and Redis
• Build and manage multi-cloud infrastructure across AWS & GCP using Terraform/OpenTofu
• Run production-grade stateful workloads on Kubernetes , including StatefulSets and persistent storage
• Help evolve infrastructure toward AI-powered and Vector Search architectures
• Build observability using Datadog, Prometheus and deep-stack monitoring
• Develop automation and infrastructure tooling using Python, Go or Bash
• Manage complex VPCs, Load Balancing, Service Meshes and caching layers
• Troubleshoot complex distributed-system issues and drive RCA and reliability improvements
• Participate in an on-call rotation and lead resolution of critical production incidents
• Leverage LLMs such as ChatGPT, Claude and GitHub Copilot to accelerate engineering and operational workflows
What We’re Looking For
✅ 7+ years in DevOps, Infrastructure or SRE
✅ Strong experience managing large-scale distributed data systems
✅ Hands-on experience with Solr, HBase and Redis preferred
✅ Strong Linux expertise, including performance tuning, I/O, memory management and JVM tuning
✅ Strong Terraform/OpenTofu experience
✅ Deep experience running Kubernetes in production , particularly stateful workloads
✅ Robust programming/scripting skills in Python or Go
✅ Experience with comparable technologies such as Elasticsearch, OpenSearch, Cassandra or BigTable is also relevant.
✅ Practical experience using AI/LLM tools to improve engineering productivity
✅ Strong troubleshooting, communication and problem-solving skills
Why This Role?
You’ll have the opportunity to work on high-scale cloud infrastructure, distributed search and data platforms, Kubernetes and emerging AI-driven architectures —with significant ownership and autonomy.