07 Oct
|
RChilli
|
Sahibzada Ajit Singh Nagar
07 Oct
RChilli
Sahibzada Ajit Singh Nagar
Role: DevOps Engineer
Exp 2-4 Years
Work mode: Work from office, Mohali, Punjab shift time: 10am to 8pm IST (10hrs shift - 9 hrs working and 1 hr break)
RChilli is looking for a DevOps Engineer to manage, automate, secure, and continuously improve the cloud infrastructure supporting our SaaS products and services.
The role involves working across multi-cloud infrastructure, Kubernetes, containerized applications, CI/CD, infrastructure automation, monitoring, databases, security, disaster recovery, and production operations. The DevOps Engineer will work closely with Development, QA, AI/ML, Security, and IT teams to deliver highly available, scalable, secure, and cost-efficient infrastructure.
The ideal candidate should have robust hands-on troubleshooting skills and be comfortable managing business-critical production environments across multiple cloud platforms and geographic regions.
Key Responsibilities
1. Cloud Infrastructure Management
Design, deploy, manage, and optimize cloud infrastructure across Google Cloud Platform (GCP), AWS, Oracle Cloud Infrastructure (OCI), and other required cloud platforms, including compute, networking, storage, IAM, load balancing, DNS, security, and managed cloud services.
2. Kubernetes Administration
Manage Kubernetes environments across Development, QA, Performance, Production, and Disaster Recovery, including clusters, nodes, node pools, Deployments, StatefulSets, Services, Ingress, ConfigMaps, Secrets, Jobs, CronJobs, resource requests/limits, autoscaling, scheduling, and Kubernetes RBAC.
3. CI/CD & Release Automation
Build, maintain, and improve CI/CD pipelines using technologies such as Jenkins, Bitbucket Pipelines, GitHub, Docker, Kubernetes, YAML, and Shell scripting to automate build, testing, security scanning, and application deployments.
4. Docker & Container Management
Build, optimize, secure, version, scan, and maintain Docker images and containerized applications. Manage images through Docker Hub, cloud artifact registries, and private container registries while following container security and image lifecycle best practices.
5. Infrastructure & Configuration Automation
Automate infrastructure provisioning, configuration, deployment, patching, migration, and repetitive operational activities using Ansible, Shell/Bash scripting, Kubernetes manifests, CI/CD pipelines, and Infrastructure-as-Code/automation practices.
6. AI-Assisted DevOps Automation
Leverage modern AI and automation tools to improve DevOps operations, including log analysis, troubleshooting, configuration generation, incident investigation, infrastructure analysis,
documentation, security reviews, and repetitive operational workflows.
7. Monitoring, Observability & Alerting
Implement and manage infrastructure and application monitoring using Prometheus, Grafana, ELK/centralized logging, Site24x7, cloud-native monitoring services, and related observability tools. Proactively investigate alerts, performance degradation, capacity issues, and service failures.
8. Production Reliability & High Availability
Maintain high availability and reliability of production services through proactive monitoring, capacity planning, resource optimization, autoscaling, redundancy, performance tuning, preventive maintenance, and effective incident response.
9. Database Operations
Support MySQL, Google Cloud SQL, and other SQL/database services, including backup and restore, replication/synchronization, migration, performance monitoring, access control, availability, maintenance, and disaster recovery activities.
10. Web & Application Server Management
Configure, maintain, secure, and troubleshoot NGINX, Apache HTTP Server, Apache Tomcat, reverse proxies, load balancers, SSL/TLS certificates, and application hosting environments.
11. Security & Infrastructure Hardening
Implement infrastructure security best practices, including IAM/RBAC, least-privilege access, server hardening, vulnerability remediation, patch management, network restrictions, firewall/WAF controls, secrets management, SSL/TLS management, container security, and secure configuration management.
12. Vulnerability & Patch Management
Perform regular OS, application, container, Kubernetes, and infrastructure patching. Coordinate remediation of identified vulnerabilities and support vulnerability scanning, penetration testing remediation, and security validation activities.
13. Disaster Recovery & Business Continuity
Maintain production and DR readiness through infrastructure replication, database synchronization, backup validation, restoration testing, failover/failback procedures, DR drills, and periodic verification of recovery procedures.
14. Incident Management & Troubleshooting
Troubleshoot infrastructure, networking, Kubernetes, application, database, deployment,
and performance issues across production and non-production environments. Participate in incident investigation, root-cause analysis, corrective actions, and preventive improvements.
15. Collaboration with Engineering Teams
Work closely with Development, QA, AI/ML, Security, and IT teams on application releases, infrastructure requirements, performance optimization, troubleshooting, architecture improvements, and production readiness.
16. Documentation & Standardization
Create and maintain SOPs, architecture documentation, deployment procedures, troubleshooting guides, DR procedures, operational runbooks, infrastructure inventories, and automation workflows to standardize DevOps operations.
Experience
Minimum 1 year and preferably 2–4 years.
(Freshers are not eligible for this position)
Required Skills & Experience The candidate should have strong hands-on experience with Linux administration, Kubernetes, Docker, cloud infrastructure, CI/CD, networking, troubleshooting, and automation.
Hands-on knowledge should include GCP and/or AWS, with exposure to multi-cloud environments preferred
• Kubernetes platforms such as GKE/OKE/EKS
• Docker and container registries
• Jenkins/Bitbucket/GitHub-based CI/CD
• Ansible and Shell/Bash scripting
• NGINX, Apache and Tomcat
• MySQL/Cloud SQL
• Prometheus/Grafana/ELK or similar monitoring stacks
• SSL/TLS, DNS, networking, firewalls and load balancing
• IAM/RBAC and infrastructure security; backup, restore and DR processes; and Git-based configuration management.
Experience supporting high-availability SaaS production environments and troubleshooting complex infrastructure/application issues is strongly preferred.
Good to Have
Knowledge or practical experience with Terraform or other Infrastructure-as-Code technologies, Helm, Cloudflare/WAF, Kubernetes security and governance, vulnerability scanning tools, centralized security monitoring/SIEM, and AI-assisted DevOps automation would be an advantage.
Core Technology Stack
Cloud: GCP, AWS, Oracle Cloud
Containers: Kubernetes, GKE/OKE, Docker
CI/CD: Jenkins, Bitbucket Pipelines, GitHub, YAML
Automation: Ansible, Bash/Shell, Infrastructure-as-Code
Monitoring: Grafana, Prometheus, ELK, Site24x7
Web/Application: NGINX, Apache Tomcat
Database: MySQL, Cloud SQL
Security: IAM, RBAC, WAF, SSL/TLS, vulnerability management, server/container hardening
Source Control: Bitbucket, GitHub
Emerging Technologies: AI-assisted DevOps, DevSecOps, FinOps and intelligent infrastructure automation
📌 DevOps Engineer (Sahibzada Ajit Singh Nagar)
🏢 RChilli
📍 Sahibzada Ajit Singh Nagar