10 Aug
|
Synthlane Technologies Private
|
Gurugram
10 Aug
Synthlane Technologies Private
Gurugram
We are seeking an accomplished Senior Site Reliability Engineer (SRE) to lead the design, implementation, and evolution of highly available, scalable, and resilient systems across our multi-cloud infrastructure. In this senior role, you will drive architectural decisions, establish reliability standards, and mentor teams while ensuring operational excellence across complex distributed systems. You will partner with engineering leadership, development teams, and product stakeholders to shape infrastructure strategy, implement sophisticated automation, and champion a culture of reliability engineering.
As a Senior SRE, you'll tackle sophisticated, large-scale challenges using cutting-edge technologies across AWS and Azure platforms. You will lead critical initiatives that impact system reliability at scale, architect solutions for complex infrastructure problems,
and guide teams in adopting industry-leading practices that drive meaningful improvements across our entire technology ecosystem.
Requirements ● Architect and implement highly reliable, scalable, and cost-effective infrastructure solutions for mission-critical applications across multi-cloud environments (AWS and
Azure).
● Lead the definition and refinement of service level objectives (SLOs), service level indicators (SLIs), and error budgets, establishing reliability standards across the organization.
● Design and implement sophisticated Infrastructure as Code (IaC) solutions using
Terraform, Ansible, and Azure Resource Manager (ARM) templates or Bicep.
● Drive automation strategies to eliminate toil, improve operational efficiency, and enable self-service capabilities for development teams.
● Lead incident response efforts, conduct thorough post-incident reviews, and implement systemic improvements to prevent recurrence.
● Champion cloud-native architectures and modern reliability practices, serving as a technical advisor for infrastructure and platform decisions.
● Participate in and help optimize the on-call rotation, ensuring sustainable practices and effective escalation procedures.
● Establish and maintain comprehensive documentation standards, runbooks, and knowledge repositories that enable team autonomy and effective incident response.
● Design and implement advanced monitoring, logging, and alerting strategies using observability platforms to enable proactive issue detection and resolution.
● Lead container orchestration initiatives using Kubernetes (AKS, EKS) and implement sophisticated deployment strategies including blue-green, canary, and progressive delivery patterns.
● Ensure security, compliance, and governance standards are embedded throughout the infrastructure lifecycle, implementing security-as-code practices.
● Drive capacity planning, performance optimization, and cost management initiatives across cloud platforms.
● Collaborate with architecture and security teams to establish platform standards,
reference architectures, and best practices. Skills, Knowledge and Expertise ● 5+ years of proven experience as a Site Reliability Engineer or similar role, with demonstrated expertise in designing, implementing, and operating large-scale,
distributed systems.
● Deep expertise in Infrastructure as Code (IaC) with Terraform and Ansible, including module development, state management, and multi-environment orchestration.
● Extensive hands-on experience with both AWS and Azure cloud platforms, including advanced services, networking, and security features in both environments.
● Expert-level knowledge of container orchestration with Kubernetes, including architecture, custom resource definitions (CRDs), operators, service mesh implementations, and production-scale cluster management.
● Advanced proficiency in Linux system administration, performance tuning, and troubleshooting complex system-level issues.
● Proven experience implementing GitOps workflows using ArgoCD, Flux, or similar tools,
including advanced deployment patterns and progressive delivery.
● Deep understanding of observability principles and hands-on experience with tools such as Prometheus, Grafana, Datadog, Azure Monitor,
or the ELK stack.
● Expert knowledge of networking concepts, including load balancing, CDNs, DNS, VPNs,
service mesh architectures, and distributed systems communication patterns.
● Strong programming and scripting capabilities in Python, Bash, Go, or PowerShell, with the ability to develop custom tooling and automation frameworks.
● Extensive experience designing and optimizing CI/CD pipelines using Jenkins, GitLab
CI, Azure DevOps, GitHub Actions, or CircleCI.
● Demonstrated ability to lead incident response, conduct root cause analysis, and drive systemic reliability improvements.
● Excellent communication and leadership skills with proven ability to influence technical decisions and collaborate with stakeholders at all levels.
● Current certification in AWS (Solutions Architect Associate/Professional or equivalent)
and Azure (Azure Administrator or Azure Solutions Architect), with practical experience managing production workloads on both platforms.
Good To
Have ● Experience with hybrid and multi-cloud networking strategies, including ExpressRoute,
Direct Connect, and cloud interconnects.
● Knowledge of serverless architectures on AWS (Lambda) and Azure (Functions, Logic
Apps) and their operational considerations.
● Proven experience with disaster recovery planning, business continuity, and implementing multi-region active-active architectures.
● Understanding of machine learning operations (MLOps), data pipeline orchestration, and supporting ML workloads in production.
● Experience with service mesh technologies such as Istio, Linkerd, or Consul.
● Familiarity with chaos engineering principles and tools like Chaos Monkey or Gremlin.
● Experience with configuration management at scale and policy-as-code tools like Open
Policy Agent (OPA).
● Knowledge of FinOps principles and cloud cost optimization strategies. Benefits
Why Work at Synthlane?
Real production infrastructure exposure (not just toy tasks). Work closely with engineers building real solutions and client deployments. Learn what “high reliability” actually means in production. Fast-paced, high-impact setting where good work is noticed quickly.
📌 Senior Software Engineer - Site Reliability (Gurugram)
🏢 Synthlane Technologies Private
📍 Gurugram