29 Aug
|
MindBrain
|
India
Senior DevOps Engineer – Production Stability
Client: Neurealm
Employment Type: C2H (Contract-to-Hire)
Location: Remote
Experience: 7+ Years
Role Overview
We are looking for a Senior DevOps Engineer – Production Stability with strong hands-on experience in Linux, DevOps/SRE, Kubernetes, CI/CD, cloud infrastructure, and production operations .
This role focuses on operating and stabilizing a complex, business-critical DevOps setting supporting 24x7 Unite Services and Integrations . The environment includes a combination of legacy and modern technologies, hybrid cloud/on-prem infrastructure, and partially documented systems.
The role is approximately 90% individual contributor (IC) and requires participation in an on-call rotation, primarily on weekdays .
Key Responsibilities
Production Operations & Incident Response
- Support 24x7 production systems for Unite Services and Integrations.
- Participate in the on-call rotation, primarily on weekdays.
- Troubleshoot incidents across CI/CD, Kubernetes, API Gateway, networking, and applications .
- Perform incident triage, mitigation, recovery, and root-cause investigation.
- Ensure safe production deployments and reliable rollback procedures.
- Maintain high availability and stability of critical production systems.
CI/CD & Deployment Stability
- Operate and maintain CI/CD pipelines using GitHub Actions, Azure DevOps, Octopus , or similar tools.
- Quickly diagnose and resolve pipeline and deployment failures.
- Support deployments across DEV, QA, PROD, and special environments .
- Improve deployment reliability, consistency, and recovery processes.
Kubernetes & Infrastructure Operations
- Support and troubleshoot Kubernetes clusters across cloud and on-prem environments.
- Perform Kubernetes patching, troubleshooting, and basic upgrades.
- Monitor and tune CPU/memory resources and investigate performance issues.
- Support infrastructure components including Redis, RabbitMQ, databases, and API gateways .
API Gateway, Cloud & Legacy Systems
- Support AWS API Gateway and related infrastructure.
- Work with existing Terraform modules and configurations .
- Troubleshoot routing, domain, and traffic-related issues involving Cloudflare and APIM .
- Support critical legacy systems, including Windows-based services .
Observability & Monitoring
- Use Prometheus, Grafana, and centralized logging systems for monitoring and troubleshooting.
- Investigate system and application issues using logs and metrics.
- Improve alerts, dashboards, and monitoring coverage where required.
- Collaborate with support teams on operational issues and escalations.
Documentation & Knowledge Sharing
- Document deployment, rollback, troubleshooting, and recovery procedures.
- Create and maintain operational runbooks for critical systems.
- Share technical knowledge with the contractor lead and team.
- Contribute to continuous improvement of operational processes.
Mandatory Skills & Experience
- 7+ years of hands-on DevOps / SRE experience
- Strong Linux administration and troubleshooting experience – Mandatory
- Strong production operations and incident management experience.
- Hands-on experience with Kubernetes operations .
- Experience with CI/CD tools such as GitHub Actions, Azure DevOps, Octopus, Jenkins, or similar.
- Experience with AWS and/or Azure .
- Working knowledge of Terraform ,
preferably with existing codebases/modules.
- Experience with Prometheus, Grafana, and logging/monitoring tools .
- Experience working in hybrid cloud + on-prem environments .
- Scripting experience in Bash, Python, PowerShell, or similar .
- Strong troubleshooting and problem-solving skills in production environments.
Preferred / Good-to-Have Skills
- AWS API Gateway
- Azure API Management (APIM)
- Cloudflare
- Redis
- RabbitMQ
- Database troubleshooting
- Windows-based legacy services
- Octopus Deploy
- Production incident management / SRE practices
- Root Cause Analysis (RCA)
- On-call support experience
Success in the First 90 Days
- Independently support critical production systems.
- Stabilize and improve CI/CD pipelines.
- Handle production incidents quickly and effectively.
- Establish reliable deployment, rollback, and recovery procedures.
- Document critical workflows and operational runbooks.
- Build strong collaboration with the contractor lead and engineering/support teams.
Who Will Succeed in This Role
- Strong operator mindset with a focus on production stability.
- Comfortable working in complex and imperfect environments.
- Quick learner who can take ownership with minimal supervision.
- Calm and effective during high-pressure incidents.
- Strong troubleshooting and analytical skills.
- Focused on reliability, stability, and continuous improvement .
- Comfortable with hands-on IC responsibilities and on-call support.
This Role Is NOT a Fit If You
- Prefer only greenfield or highly standardized environments.
- Focus primarily on architecture rather than production operations.
- Prefer not to take production ownership or participate in on-call support.
- Require fully documented systems and processes.
- Are uncomfortable troubleshooting legacy and hybrid environments.
📌 Senior DevOps Engineer (7+ yrs) (India)
🏢 MindBrain
📍 India