07 Aug
|
Cvent
|
Gurugram
About the Role: Site Reliability is about combining development and operations knowledge and skills to help make the organizationbetter. Whether you have a development background and are interested in learning more about operations and security or have an operations or security background and are interested in developing internal tools and automation, SRE can benefit from your skillsets. Ultimately, we are looking for passionate people who love learning, love technology and always want to make things better.
In This Role, You Will:
- Enlighten, Enable and Empower a fast-growing set of multi-disciplinary teams, across multiple applications and locations.
- Tackle complex development, automation and business process problems. Champion SRE standards and bestpractices.
- Ensure the scalability, performance, and resilience of security related systems and processes.
- Work with product development teams, Information Security, Cloud Automation and other SRE teams to ensurea holistic understanding of security concerns and their effective and efficient identification and resolution.
- Identify recurring problems and anti-patterns in development, operational and security processes.
- Develop build, test and deployment automation that seamlessly targets multiple on-premises and AWS regions.
- Give back by working on and contributing to Open-Source projects.
Heres What You Need:
Must Have Skills:
- 7-10 years of hands-on experience in Site Reliability Engineering with a demonstrated track record of owning reliability, security, and operational excellence at scale in production environments.
- Excellent communication skills and a track record of driving alignment across multi-disciplinary teams.
- A passion for and track record in making things better for your peers.
- Hands-on experience with AWS WAF including rule authoring, rate-based rules,
bot control integration, WAF rule group management, and multi-product WAF sharing strategies (e.g., managing WAF rule limits across applications sharing the same WebACL).
- Experience designing and implementing DDoS protection using AWS Shield Advanced including transitioning endpoints from count to block mode, building observability solutions (Lambda + CloudWatch alarms), and self-service enablement for product teams.
- Experience with bot mitigation strategies including AWS Bot Control, silent challenge / token-based traffic classification (verified humans, verified bots, unknown traffic), JA4+ASN fingerprinting, and evaluation of third-party bot mitigation vendors (e.g., Datadome).
- Experience managing AWS services and operational knowledge of running applications in AWS ideally via automation and Infrastructure as Code (IaC) using CloudFormation or CDK.
- Strong understanding of CI/CD pipelines experience with Jenkins or equivalent, PR-based deployment workflows, build/test/deploy automation, and troubleshooting pipeline failures in distributed environments.
- Incident management experience able to act as IC, write clear incident summaries, drive RCA, and coordinate resolution across teams under pressure.
- Change management discipline ability to communicate changes proactively to stakeholders, document rollout strategies, and manage phased production deployments with rollback plans.
- Fluent in at least one scripting language such as TypeScript, JavaScript, Python, Ruby, or Bash.
- Experience with SDLC methodologies (preferably Agile).
AI & Automation Literacy (Must Have):
Practical understanding and hands-on exposure to AI fundamentals as applied to SRE and operational workflows:
- Prompt Engineering ability to design effective prompts for LLMs to assist with incident analysis, RCA generation, runbook creation, and on-call triage.
- Retrieval-Augmented Generation (RAG) basic understanding of RAG patterns; ability to leverage or contribute to RAG-based internal tools that surface relevant runbooks, past incidents, and knowledgebase articles during operational events.
- AI-assisted Workflow & Process Automation experience using or building AI-powered automations in operational contexts, such as automated incident summarization, alert enrichment, change risk assessment, or post-mortem drafting using LLM integrations (e.g., via MCP tools, Slack bots, or custom pipelines).
Positive to Have Skills:
- Disaster recovery planning and execution experience with multi-region failover, DR runbooks, and recovery time / recovery point objective (RTO/RPO) management.
- Experience managing CloudFront distributions, API Gateways, and ALBs as part of a layered security posture.
- Experience with APM, monitoring and logging tools (Datadog, New Relic, Splunk).
- Familiarity with security assessment tools and methodologies:
- Cloud Security Posture Management (CSPM)
- Infrastructure vulnerability management
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Lead Site Reliability Engineer (Security) (Gurugram)
🏢 Cvent
📍 Gurugram