04 Aug
|
Armor Defense
|
Maharashtra
04 Aug
Armor Defense
Maharashtra
SUMMARY
We are looking for a highly skilled Site Reliability Engineer (SRE) to join our infrastructure team with expertise across Cloud Deployments, Microsoft Entra ID (Azure AD), Active Directory, Office 365, Zerto, Rubrik, VMware, and NSX-T. This hands-on, automation-heavy role focuses on system reliability, scalability, and proactive problem prevention. You will be responsible for building and maintaining resilient infrastructure, automating repetitive tasks, monitoring, and improving performance, and driving incident reduction strategies across hybrid cloud environments.
This role operates in a hybrid structure with on-site presence four days a week, specifically Monday, Tuesday, Wednesday, and Thursday, based in Pune, India.
ESSENTIAL DUTIES AND RESPONSIBILITIES (Additional duties may be assigned as required)
- Administer and maintain Microsoft Entra ID and on-prem Active Directory environments.
- Configure conditional access, identity protection, and secure authentication policies.
- Deploy and administer VMware vSphere and NSX-T environments.
- Design and implement scalable, secure virtual network topologies.
- Assist in migrating and managing workloads across VMware, AWS, Azure, and OCI when necessary.
- Manage and optimize Active Directory and Office 365 services, including Exchange Online, SharePoint, Teams, and Intune.
- Write clean, maintainable code (Python, PowerShell, etc.) to automate ops tasks and monitor system health.
- Implement self-healing mechanisms and proactive detection of issues.
- Contribute to infrastructure-as-code efforts using Terraform, with a focus on modularity and DRY principles.
- Define SLIs/SLOs and build custom dashboards and alerts using tools like Datadog, Prometheus, Grafana, Splunk, or equivalent.
- Collaborate with engineering and security teams to reduce toil and improve platform stability.
- Design and implement monitoring, alerting, and incident response workflows for new production deployments.
- Maintain and tune Zerto replication and Rubrik backup infrastructure.
- Ensure DR strategies meet business RTO/RPO and compliance requirements.
- Lead response and resolution for high-impact infrastructure incidents.
- Conduct blameless postmortems and drive continuous improvement.
- Participate in on-call rotations and root cause analysis for incidents impacting production services.
- Lead the automated vulnerability patch management program through automation.
REQUIRED SKILLS
- 8+ years of experience in SRE, DevOps, Systems, or Infrastructure Engineering roles in production environments.
- 8+ years of experience in Windows production environments.
- 3+ years of experience with Linux/Unix production deployments.
- 3+ years of experience with Kubernetes.
- Strong troubleshooting skills, especially in complex or hybrid environments.
- Strong communication skills with a robust command of English.
- Deep hands-on knowledge of VMware technologies (vSphere, ESXi, vCenter, NSX-T).
- Experience with Oracle Cloud Infrastructure (OCI), including compute, networking, and IAM.
- Experience in AWS using Terraform to manage environments.
- Experience in Azure environments, particularly focused on Entra ID.
- A transparent understanding of Secure Landing Zone concepts.
- Expertise in Microsoft Entra ID (Azure AD) and on-prem Active Directory administration, monitoring, and troubleshooting.
- Experience with Zerto and Rubrik platforms.
- Cloud understanding of RTO/RPO and meeting SLAs for DR.
- Experience with Terraform, Ansible, or similar IaC tools.
- Strong scripting skills in Python, PowerShell, Bash, or equivalent.
- CI/CD and GitOps experience using tools like GitLab CI and Jenkins.
- Proficient with version control systems (Git).
- Experience with monitoring and alerting tools like Prometheus, Grafana, ELK, and Datadog.
- Understanding of system-level networking, DNS, firewalls, and load balancing.
- Familiarity with security and compliance frameworks (PCI, HIPAA, ISO, etc.).
WORK ENVIRONMENT
The work environment characteristics described here are representative of those an employee encounters while performing the essential functions of this job. The noise level in the work environment is usually low to moderate. The work environment can be either in an office setting or remotely from anywhere.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Site Reliability Engineer (Maharashtra)
🏢 Armor Defense
📍 Maharashtra