Overview
We are seeking an exceptional Senior DevOps Engineer with more than 10 years of hands-on experience in automation, cloud infrastructure, and CI/CD pipeline development. This role demands advanced expertise in AWS, Infrastructure as code, and container orchestration. The ideal candidate excels in planning new environments, integrating Docker-based applications, and evolving CI/CD processes to meet business needs.
As a Senior DevOps Engineer, you will work directly with engineering and cross-functional teams, leveraging deep AWS knowledge to deliver scalable, secure, and efficient cloud-native solutions.
Key Responsibilities
DevOps Strategy & Implementation
- Lead the adoption and execution of DevOps practices across multiple projects, promoting automation, standardization, and continuous improvement.
- Drive the evolution of DevOps culture and processes throughout the organization.
AWS Infrastructure Management
- Create and maintain AWS environments using EC2, RDS, S3, IAM, ELB, CloudWatch, Route53, Systems Manager (SSM), and Secrets Manager.
- Ensure infrastructure is scalable, reliable, and aligned with business requirements.
Infrastructure as Code (IaC)
- Build and manage AWS infrastructure using Terraform for reproducible, version-controlled deployments.
- Promote best practices in infrastructure automation and maintainability.
CI/CD Pipeline
- Design and optimize CI/CD pipelines leveraging GitLab, Artifactory, Terraform Cloud and Helm
- Integrate with SAST/DAST tools for security scans and Code Quality
Containerization & Application Integration
- Design and deploy containerized workloads using EKS
GitOps
- Implement GitOps workflows using ArgoCD to drive declarative, Git-based deployment and configuration management for Kubernetes (EKS) workloads.
- Maintain Git as the single source of truth for infrastructure and application manifests, enabling automated sync, drift detection, and rollback.
- Define branching and promotion strategies (dev → staging → production) to support protected, auditable, and repeatable releases.
Cost Efficiency
- Conduct performance analysis and recommend on AWS cost-saving strategies.
Monitoring & Observability
- Implement and maintain monitoring and alerting solutions using Grafana, Prometheus, Loki, New Relic and CloudWatch.
- Setup Alerts using Slack and JSM (on call support)
- Enhance system visibility and maximize uptime across environments.
- Leverage Grafana Cloud (Managed Prometheus, Loki, and Tempo) for unified metrics, logs, and tracing, reducing self-hosted observability overhead.
- Build and maintain Grafana Cloud dashboards and SLO-based alerting to provide real-time visibility into service health and performance.
Security & Access Management
- Create and manage IAM roles and policies, enforcing secure access control and least-privilege compliance.
- Ensure adherence to organizational security standards.
Disaster Recovery & Business Continuity
- Participate in Quarterly Disaster Recovery drills, creating complete infrastructure replicas in alternate AWS regions.
- Validate application recovery procedures and ensure business continuity readiness.
On-Call Support & Production Troubleshooting
- Participate in on-call rotation to provide 24/7 production support, responding promptly to incidents and escalations via Slack and JSM.
- Diagnose and troubleshoot production issues across infrastructure, application, and network layers, using Grafana Cloud, CloudWatch, and logs to identify root cause.
- Drive incident response and resolution,
ensuring timely restoration of service and minimal business impact.
- Conduct post-incident reviews and root cause analysis, documenting findings and implementing preventive measures to reduce recurrence.
Collaboration & Feature Enablement
- Work closely with development teams to plan, deploy, and test new features, while maintaining strong DevOps automation and reliability.
Required Skills & Technologies
- Infrastructure as Code: Terraform
- CI/CD Tools: GitLab Runners, Helm, Artifactory
- Containerization & Orchestration: EKS, Docker
- AWS Services: EKS, EC2, RDS, S3, IAM, ElastiCache, SNS, SQS, Secrets Manager, Systems Manager (SSM), CloudWatch, CloudFront, Route53, ELB, API Gateway
- Secrets & Configuration Management: 1Password, AWS Secrets Manager
- Monitoring & Logging: New Relic, Grafana, Prometheus, Loki
- GitOps: ArgoCD, Flux
- Observability Platform: Grafana Cloud (Managed Prometheus, Loki, Tempo)
Qualifications
- Bachelor’s or Master’s degree in Computer Science, Information Technology, or related field.
- 10+ years of hands-on experience in DevOps, Cloud, or Infrastructure Engineering roles.
- Proven expertise in AWS Cloud, CI/CD pipeline development, infrastructure automation, and container orchestration.
- Strong proficiency with Terraform, GitOps and GitLab CI/CD.
- Demonstrated success in CI/CD migration projects and AWS cost optimization initiatives.
- Excellent analytical, troubleshooting, and collaboration skills.
- Effective communicator and mentor with leadership experience.
Preferred Attributes
- AWS DevOps Engineer/Terraform/Kubernetes certifications
- Experience with multi-region, high-availability infrastructure and disaster recovery planning.
- Proven success leading DevOps transformations and toolchain modernizations on a scale.
If Interested please share your updated resume at
[email protected]
📌 Senior DevOps Engineer (Hyderabad)
🏢 Evoke Technologies
📍 Hyderabad