07 Oct
|
MindBrain
|
India
Senior DevOps Engineer – Production Stability
Experience: 9+ Years
Employment Type: Contract-to-Hire (C2H)
Location: Remote
Job Overview
We are looking for a highly experienced Senior DevOps Engineer with a strong production-operations mindset to support and stabilize business-critical 24x7 systems. The role involves working across hybrid cloud and on-premises environments, legacy and modern technologies, CI/CD pipelines, Kubernetes, networking, API gateways, and observability platforms.
This is a highly hands-on role (~90% individual contributor) requiring strong troubleshooting and incident-response capabilities. The selected candidate will work as part of a 3–4 member contractor engineering pod and participate in an on-call rotation, primarily on weekdays.
Key Responsibilities
Production Operations & Incident Management
- Support and maintain 24x7 production systems and integrations.
- Participate in the on-call rotation and respond to production incidents.
- Troubleshoot issues across CI/CD, Kubernetes, applications, networking, and API gateways.
- Perform incident triage, mitigation, recovery, and root-cause analysis.
- Ensure reliable deployments, rollback procedures, and production stability.
CI/CD & Deployment
- Manage and troubleshoot CI/CD pipelines using GitHub Actions, Azure DevOps, Octopus, or similar tools.
- Diagnose and resolve pipeline and deployment failures.
- Support deployments across DEV, QA, PROD, and special environments.
- Improve deployment reliability, automation, and consistency.
Kubernetes & Infrastructure
- Operate and troubleshoot Kubernetes clusters across cloud and on-premises environments.
- Perform cluster troubleshooting, patching, and basic upgrades.
- Monitor and optimize CPU, memory, and other infrastructure resources.
- Troubleshoot supporting services such as Redis, RabbitMQ, databases, and API gateways.
Cloud, API Gateway & Legacy Systems
- Support AWS API Gateway and related cloud infrastructure.
- Work with existing Terraform modules and configurations.
- Troubleshoot routing, domain, and traffic-related issues involving Cloudflare/APIM.
- Support critical legacy systems, including Windows-based services.
Monitoring & Observability
- Use Prometheus, Grafana, logs, and monitoring platforms for production troubleshooting.
- Analyze system performance and identify potential reliability issues.
- Improve alerts, dashboards, and monitoring where required.
Documentation & Knowledge Sharing
- Create and maintain deployment, rollback, recovery, and troubleshooting documentation.
- Develop and maintain operational runbooks for critical systems.
- Share technical knowledge with the contractor lead and engineering team.
Mandatory Skills
- Robust hands-on Linux experience – Mandatory
- 9+ years of experience in DevOps / SRE / Production Operations
- Strong production troubleshooting and incident-management experience
- Hands-on Kubernetes operations
- Experience with CI/CD tools such as GitHub Actions, Azure DevOps, or Octopus
- Experience with AWS and/or Azure
- Working knowledge of Terraform
- Experience with Prometheus, Grafana, and centralized logging
- Experience supporting hybrid cloud + on-premises environments
- Strong scripting skills using Bash, Python, PowerShell, or similar
Preferred Experience
- AWS API Gateway / Azure API Management
- Cloudflare and networking
- Redis and RabbitMQ
- Database troubleshooting
- Windows-based legacy systems
- Production on-call and 24x7 support environments
- Incident response, RCA, and operational runbooks
What We’re Looking For
- Strong operator/SRE mindset
- Comfortable working in complex and partially documented environments
- Ability to quickly understand existing systems and take ownership
- Calm and effective under production pressure
- Strong troubleshooting and problem-solving skills
- Focused on stability, reliability, and business continuity
- Comfortable with hands-on production ownership and on-call responsibilities
First 90-Day Success Criteria
- Independently support critical production systems.
- Stabilize and improve CI/CD pipeline reliability.
- Handle production incidents efficiently and minimize downtime.
- Document critical deployment, rollback, and recovery workflows.
- Establish effective collaboration with the contractor lead and engineering team.
Not a Good Fit If You
- Prefer only greenfield or highly standardized environments.
- Focus primarily on architecture rather than hands-on operations.
- Prefer to avoid production ownership or on-call responsibilities.
- Require fully documented systems before taking ownership.
- Are uncomfortable troubleshooting under time-sensitive production conditions.
📌 Senior DevOps Engineer (9+ yrs) (India)
🏢 MindBrain
📍 India