Overview We are seeking an exceptional Senior DevOps Engineer with more than 10 years of hands-on experience in automation, cloud infrastructure, and CI/CD pipeline development. This role demands advanced expertise in AWS, Infrastructure as code, and container orchestration. The ideal candidate excels in planning new environments, integrating Docker-based applications, and evolving CI/CD processes to meet business needs.
As a Senior DevOps Engineer, you will work directly with engineering and cross-functional teams, leveraging deep AWS knowledge to deliver scalable, secure, and efficient cloud-native solutions.
Key Responsibilities DevOps Strategy & Implementation - Lead the adoption and execution of DevOps practices across multiple projects, promoting automation, standardization, and continuous improvement.
- Drive the evolution of DevOps culture and processes throughout the organization. AWS Infrastructure Management - Create and maintain AWS environments using EC2, RDS, S3, IAM, ELB, CloudWatch, Route53, Systems Manager (SSM), and Secrets Manager.
- Ensure infrastructure is scalable, reliable, and aligned with business requirements. Infrastructure as Code (IaC) - Build and manage AWS infrastructure using Terraform for reproducible, version-controlled deployments.
- Promote best practices in infrastructure automation and maintainability. CI/CD Pipeline - Design and optimize CI/CD pipelines leveraging GitLab, Artifactory, Terraform Cloud and Helm
- Integrate with SAST/DAST tools for security scans and Code Quality Containerization & Application Integration - Design and deploy containerized workloads using EKS GitOps - Implement GitOps workflows using ArgoCD to drive declarative, Git-based deployment and configuration management for Kubernetes (EKS) workloads.
- Maintain Git as the single source of truth for infrastructure and application manifests, enabling automated sync, drift detection, and rollback.
- Define branching and promotion strategies (dev → staging → production) to support safe, auditable, and repeatable releases. Cost Efficiency - Conduct performance analysis and recommend on AWS cost-saving strategies. Monitoring & Observability - Implement and maintain monitoring and alerting solutions using Grafana, Prometheus, Loki, New Relic and CloudWatch.
- Setup Alerts using Slack and JSM (on call support)
- Enhance system visibility and maximize uptime across environments.
- Leverage Grafana Cloud (Managed Prometheus, Loki, and Tempo) for unified metrics, logs, and tracing, reducing self-hosted observability overhead.
- Build and maintain Grafana Cloud dashboards and SLO-based alerting to provide real-time visibility into service health and performance. Security & Access Management - Create and manage IAM roles and policies, enforcing secure access control and least-privilege compliance.
- Ensure adherence to organizational security standards.
Disaster
Recovery & Business Continuity - Participate in Quarterly Disaster Recovery drills, creating complete infrastructure replicas in alternate AWS regions.
- Validate application recovery procedures and ensure business continuity readiness. On-Call Support & Production Troubleshooting - Participate in on-call rotation to provide 24/7 production support, responding promptly to incidents and escalations via Slack and JSM.
- Diagnose and troubleshoot production issues across infrastructure, application, and network layers, using Grafana Cloud, CloudWatch, and logs to identify root cause.
- Drive incident response and resolution,
ensuring timely restoration of service and minimal business impact.
- Conduct post-incident reviews and root cause analysis, documenting findings and implementing preventive measures to reduce recurrence. Collaboration & Feature Enablement - Work closely with development teams to plan, deploy, and test new features, while maintaining solid DevOps automation and reliability.
Required Skills & Technologies - Infrastructure as Code: Terraform
- CI/CD Tools: GitLab Runners, Helm, Artifactory
- Containerization & Orchestration: EKS, Docker
- AWS Services: EKS, EC2, RDS, S3, IAM, ElastiCache, SNS, SQS, Secrets Manager, Systems Manager (SSM), CloudWatch, CloudFront, Route53, ELB, API Gateway
- Secrets & Configuration Management: 1Password, AWS Secrets Manager
- Monitoring & Logging: New Relic, Grafana, Prometheus, Loki
- GitOps: ArgoCD, Flux
- Observability Platform: Grafana Cloud (Managed Prometheus, Loki, Tempo) Qualifications - Bachelor's or Master's degree in Computer Science, Information Technology, or related field.
- 10+ years of hands-on experience in DevOps, Cloud, or Infrastructure Engineering roles.
- Proven expertise in AWS Cloud, CI/CD pipeline development, infrastructure automation, and container orchestration.
- Strong proficiency with Terraform, GitOps and GitLab CI/CD.
- Demonstrated success in CI/CD migration projects and AWS cost optimization initiatives.
- Excellent analytical, troubleshooting, and collaboration skills.
- Effective communicator and mentor with leadership experience.
Preferred
Attributes - AWS DevOps Engineer/Terraform/Kubernetes certifications
- Experience with multi-region, high-availability infrastructure and disaster recovery planning.
- Proven success leading DevOps transformations and toolchain modernizations on a scale. If Interested please share your updated resume at
[email protected]
📌 Senior DevOps Engineer (Hyderabad)
🏢 Evoke Technologies
📍 Hyderabad