03 Oct
|
Xebia IT Architects
|
Gurugram
03 Oct
Xebia IT Architects
Gurugram
DevOps & Automation Engineer
We have 4 roles
2 for 7 + yrs
4 for 10+ yrs
Openings: Senior DevOps Engineer (7+ years) and Lead / Principal DevOps Engineer (10+ years)
Location: Delhi NCR
Work model: Hybrid, with on-site work at the client office in Gurgaon two days a week
About the Role
We are looking for Cloud agnostic DevOps and automation engineers to integrate PagerDuty of any similar toll with our monitoring and observability ecosystems on Cloud platforms. You will build reliable, automated alerting and incident-response workflows so that the right alerts reach the right people at the right time, with less noise and faster resolution. You will work closely with the client team and be on-site at their Gurgaon office two days a week.
Key Responsibilities
- Design, implement, and maintain PagerDuty integrations with monitoring and observability tools (e.g., Amazon CloudWatch, Datadog, Prometheus/Alertmanager, Grafana, Recent Relic, Splunk, Dynatrace, Nagios/Zabbix).
- Automate PagerDuty / or Similar tool onboarding, configuration management, and incident response workflows using Terraform, PagerDuty APIs, and Rundeck-based operational automation.
- Configure PagerDuty services, escalation policies, schedules, event orchestration, routing rules, and alert grouping and suppression to reduce alert fatigue[Good to have].
- Build event-driven automation on AWS/Azure/GCP (EventBridge, SNS, Lambda, Systems Manager, Step Functions) for auto-remediation and incident enrichment.
- Integrate PagerDuty or similate tool with ITSM and collaboration tools such as ServiceNow, Jira, Slack, monitoring, cloud, identity and data platforms and Microsoft Teams for ticketing and incident communication.
- Build and maintain CI/CD pipelines (Jenkins, GitHub Actions, GitLab CI, or AWS CodePipeline) to deliver monitoring and alerting configuration as code.
- Create runbooks, integration documentation, and standard operating procedures.
- Collaborate with the client's SRE, application, platform,
and security teams to onboard services onto standardised alerting and on-call practices.
Required Skills (Both Levels)
- Hands-on experience with PagerDuty or any alternate similar tool (services, integrations, escalation policies, event orchestration, Events API v2, REST API).
- Strong Cloud infrastructure and DevOps experience.
- Proficiency in Infrastructure as Code, preferably Terraform.
- Requirement for adaptability across platforms and tools, not restricted to a single stack
- Strong automation and scripting skills in Python, Bash, JavaScript(for webhooks, event transformers, and payload handling), including REST API and webhook integrations.
- Should have strong understanding of AIOPs concepts
- Experience with monitoring and observability tools and with building meaningful alerts and SLO-based alerting.
- Working knowledge of CI/CD tooling and Git-based workflows.
- Solid understanding of incident management, on-call practices, and SRE principles.
Level-Specific Expectations
Senior DevOps Engineer (7+ years)
- 7+ years in DevOps, SRE, or cloud automation, with at least 2 years working with PagerDuty.
- Independently implements and supports integrations across multiple monitoring tools.
- Writes reusable Terraform modules and automation scripts.
- Troubleshoots integration and alert-routing issues end to end.
- Contributes to runbooks, standards, and post-incident reviews.
Lead / Principal DevOps Engineer (10+ years)
- 10+ years in DevOps, SRE, or cloud engineering, with 3+ years of deep PagerDuty ownership at enterprise scale.
- Owns the architecture and strategy for incident management and alerting across multiple teams, accounts, and environments.
- Defines alerting standards, on-call policies, and event-orchestration frameworks.
- Designs multi-account AWS monitoring and alert-aggregation patterns.
- Leads the automation roadmap, including auto-remediation and AIOps or event-intelligence capabilities.
- Drives operational-maturity initiatives such as SLO/SLI adoption, blameless post-mortems, and reliability reporting.
📌 DevOps & Automation Engineer (Gurugram)
🏢 Xebia IT Architects
📍 Gurugram