16 Sep
|
Multycomm Interactive Media
|
Delhi
16 Sep
Multycomm Interactive Media
Delhi
We are looking for a DevOps Engineer who will own the deployment, configuration, patching and day-to-day operational health of the MCSSE platform across all customer environments. This is a hands-on infrastructure role. You will work at the Linux shell daily, manage Docker-based services, automate provisioning through Ansible, run the web and proxy tier, and keep multi-node clusters healthy and secure.
Because the platform carries live voice traffic, this role demands discipline around change control, patching windows and rollback planning. Prior telecom or real-time-communications exposure is a strong advantage but not mandatory — we will train the right engineer on the VoIP stack.
Key Responsibilities Linux System AdministrationAdminister Debian and other Linux distributions across production, staging and lab environments.
Work fluently at the command line: process and service management (systemd), file systems and storage, permissions, user and group administration, networking (ip, ss, iptables/nftables, tcpdump), and log analysis (journalctl, grep, awk, sed). Diagnose performance issues using standard tooling (top, htop, iostat, vmstat, netstat, sar). Write and maintain shell scripts for routine automation and operational tasks.
Containerisation and DockerBuild, deploy and maintain services using Docker and Docker Compose . Author and optimise Dockerfiles and multi-service Compose stacks for the MCSSE platform components.
Manage container lifecycle: image builds, registry management, tagging and versioning, volume and network configuration, resource limits, health checks and log drivers. Troubleshoot container-level issues — networking, storage, permissions, host/kernel compatibility and CPU instruction-set constraints on virtualised hosts. Maintain a private image registry and enforce image hygiene and vulnerability scanning.
Configuration
Management and AutomationWrite, test and maintain Ansible playbooks and roles for repeatable deployment of platform components. Manage inventories, variables, templates (Jinja2), handlers and vaults for multi-tenant and multi-site deployments. Convert manual runbooks into idempotent automation; reduce deployment time and eliminate configuration drift.
Maintain infrastructure and configuration code in Git with proper branching, review and release tagging. Web Services, Proxying and Load BalancingDeploy, configure and tune Nginx , HAProxy and Apache HTTP Server .
Configure
Nginx as a reverse proxy — upstream definitions, proxy headers, TLS termination, SNI, WebSocket proxying, caching, rate limiting, timeouts and buffering. Configure HAProxy for high-availability load balancing — frontends and backends, health checks, stick tables, TCP and HTTP modes, failover behaviour.
Manage TLS/SSL certificates end to end: issuance, renewal, automation, cipher-suite hardening and certificate inventory. Troubleshoot proxy-layer issues affecting WebSocket dashboards, API gateways and web-based softphone clients.
Patch
Management and Security HardeningPlan and execute Linux OS patching — kernel, security and package updates — across all environments with defined maintenance windows and rollback plans. Patch and upgrade application and middleware software: Docker engine, Nginx, HAProxy, Apache, CouchDB, RabbitMQ and platform services. Track CVEs relevant to the stack; assess applicability, prioritise by severity and drive remediation to closure.
Apply and maintain system hardening baselines: SSH hardening, firewall rules, fail2ban, service minimisation, file integrity and audit logging. Support VAPT cycles by remediating findings and providing evidence of closure. Node and Cluster ManagementManage multi-node clusters supporting the MCSSE platform, including CouchDB clusters, RabbitMQ clusters and the Kamailio/FreeSWITCH media and signalling tier.
Handle node provisioning, joining and removal, quorum and replication health, failover testing and capacity planning. Maintain high-availability configurations (VRRP/Keepalived, active–active and active–passive patterns) and validate failover through scheduled drills. Own backup,
snapshot and restore procedures; test restores rather than assuming they work.
Monitoring, Observability and Incident ResponseDeploy and maintain monitoring and alerting (Prometheus/Grafana, Zabbix, Nagios or equivalent) with meaningful thresholds and low alert noise. Maintain centralised logging and make logs searchable for engineering and support teams. Participate in an on-call rotation; perform first-line triage, escalate appropriately, restore service and complete root cause analysis.
Produce post-incident reports and drive corrective actions to completion. Documentation and ProcessMaintain deployment runbooks, network and architecture diagrams, and configuration baselines to Multycomm's documentation standards. Record all production changes through the change-management process with rollback steps defined in advance.
Support customer acceptance activities (UAT) and contribute to SLA compliance reporting.
Required Skills and Experience Mandatory 3–6 years in a DevOps, SRE, systems engineering or infrastructure operations role. Solid hands-on Linux administration, preferably Debian or Ubuntu. Production experience with Docker and Docker Compose. Demonstrated ability to write and maintain Ansible playbooks and roles. Working configuration experience with Nginx (including reverse proxy), HAProxy and Apache. Practical experience with OS and software patch management in production.
Experience operating multi-node and clustered services. Competence with Git and a scripting language (Bash required; Python preferred).
Solid networking fundamentals: TCP/IP, DNS, NAT, routing, firewalls, VPN, TLS.
Preferred
Exposure to VoIP/SIP infrastructure — FreeSWITCH, Kamailio, Asterisk, Kazoo or similar.
Experience with RabbitMQ and CouchDB or other NoSQL/clustered datastores.
Experience running services in virtualised environments (KVM/QEMU, VMware, Proxmox) and on public cloud (AWS, Azure or GCP). Familiarity with CI/CD tooling (GitLab CI, Jenkins, GitHub Actions). Exposure to Kubernetes or container orchestration beyond Compose.
Experience supporting audited or certified environments (ISO 27001, SOC 2) and VAPT remediation.
Relevant certifications: RHCE, LFCS/LFCE, CKA, or cloud associate-level certifications.
📌 DevOps Engineer (Delhi)
🏢 Multycomm Interactive Media
📍 Delhi