17 Sep
|
Numerator
|
India
We're reinventing the market research industry. Let's reinvent it together.
At Numerator, we believe tomorrow's success starts with today's market intelligence. We empower the world's leading brands and retailers with unmatched insights into consumer behavior and the influencers that drive it.
We are seeking an experienced (5+ years) Cloud Infrastructure Engineer to manage the reliability of our Azure virtual machine estate. Predominantly Windows today, with a smaller Linux footprint, currently around 60 VMs and growing to 300+ as we bring more business units onto the platform. Your job is to keep it patched, backed up, monitored, and recoverable, and automate yourself out of the repetitive tasks. We expect problems to be solved once, in code, with Terraform and scripting.
Key Responsibilities
. Manage the health and reliability of the Azure VM estate (Windows and Linux), availability, performance, and capacity.
. Run patching and update management with Azure Update Manager: patch compliance, maintenance windows, remediation of failures, and handling applications that need a version held or pinned without falling out of the compliance cycle.
. Manage backup and disaster recovery: Azure Backup, Azure Site Recovery, and regular restore testing. A backup that was never restored doesn't count.
. Build monitoring and alerting with Azure Monitor and KQL (Kusto Query Language) and drive auto-remediation so known issues fix themselves.
. Automate VM lifecycle operations (provisioning, configuration, decommissioning) with Terraform and scripting.
. Respond to incidents affecting the VM estate, drive root cause analysis, and fix the class of problem, not just the instance.
. Diagnose cases where the real root cause is security tooling. Antivirus/EDR flagging and quarantining an application file, for example,
rather than assuming the fault sits in the application or the infrastructure.
. Support security-led investigations on the VM estate: pull logs, process activity, and access history on request, and hold off on remediating or restarting a box until security has cleared it.
. Reduce toil: identify repetitive manual work and eliminate it through automation.
. Document runbooks and operational standards, so the platform is operable by the whole team.
Skills & Requirements
. 5+ years of experience in systems engineering, or infrastructure operations, with solid administration depth in Windows Server, plus good working knowledge of Linux.
. Proven experience running VM estates on Azure: patching (Azure Update Manager), backup and DR (Azure Backup, Site Recovery), and monitoring (Azure Monitor).
. Hands-on experience with KQL (Kusto Query Language) for Log Analytics queries and alert logic. This underpins how monitoring and alerting has been built here.
. Strong automation mindset with proven experience managing infrastructure as code with Terraform.
. Strong scripting skills in PowerShell and Bash.
. Solid incident management and root cause analysis skills.
. Linux package management (apt/yum) including holding or pinning versions where an application requires it and rolling back a change that breaks a workload.
. Comfortable assisting security team investigations, Windows/Linux log and process/user activity review, without needing to be a dedicated security specialist.
. Ability to work autonomously, own problems end to end and communicate clearly with distributed teams.
. Experience working on enterprise-scale projects with distributed teams.
Preferred Qualifications
. Bachelor's degree in Computer Science, Information Technology, or a related field.
. Microsoft certifications related to Azure (e.g., AZ-104 Azure Administrator Associate).
📌 Infrastructure Engineer (Azure) (India)
🏢 Numerator
📍 India