01 Aug
|
FlexAI
|
Bengaluru
Role Overview
FlexAI is looking for a Senior DevOps / SRE Engineer to build and operate the infrastructure powering our AI and PaaS platform.
You’ll work closely with developers to ensure our systems are reliable, performant, and scalable, while enabling quick product iteration. This role is hands-on and execution-focused, with opportunities to contribute to system design and reliability practices as we scale.
What You’ll Do
Build & Operate Infrastructure:
Build and maintain infrastructure for our AI and PaaS platform
Deploy and operate Kubernetes clusters and containerized services
Implement Infrastructure as Code using Pulumi (or similar tools)
Reliability & SRE Practices:
Help define and implement SLIs, SLOs, and error budgets
Improve system reliability, availability, and performance
Participate in on-call rotations, incident response, and postmortems
CI/CD & Automation:
Build and improve CI/CD pipelines for reliable and quick releases
Automate operational workflows and reduce manual toil
Contribute to GitOps and platform engineering practices
Observability & Performance:
Implement and maintain observability using VictoriaMetrics, Grafana (metrics, logs, traces)
Monitor systems and troubleshoot performance issues (latency, throughput, cost)
Collaboration:
Work closely with developers, platform, and AI teams to support production systems
Help debug issues across infrastructure and application layers
Contribute to improving engineering productivity and developer experience
What You’ll Need to Be Successful
4+ years of experience in DevOps, SRE, or Infrastructure Engineering
Experience operating production systems at scale
Hands-on experience with:
Kubernetes & containers
Infrastructure as Code (Pulumi, Terraform, etc.)
Cloud or hybrid environments (AWS, GCP, Azure, or on-prem)
Observability tools (Prometheus, Grafana, OpenTelemetry)
Experience with CI/CD systems and automation
Proficiency in Python, Go, or Bash
-
📌 Senior Devops Engineer/sre Bengaluru
🏢 FlexAI
📍 Bengaluru