Site Reliability Engineer (SRE) (Bengaluru)

Site Reliability Engineer (SRE) (Bengaluru)

09 Sep
|
VARITE
|
Bengaluru

09 Sep

VARITE

Bengaluru

Company Name: VARITE India Private Limited

About The Client

CDW helps its customers to navigate an increasingly complex IT market and maximize return on their technology investments.

About The Job:

- As part of our Managed Services organization, you’ll play a critical role in ensuring the reliability, scalability, and performance of our systems, bridging the gap between software engineering and infrastructure operations.

- You will be responsible for maintaining the operational excellence and reliability of Managed Services software and infrastructure—on-premises and in the cloud.

- You’ll drive improvements through automation, monitoring, and deep technical troubleshooting, while mentoring less experienced staff and providing escalation paths for critical issues.

- You’ll enforce best practices in system architecture, reliability, and service performance.

Essential Job Functions:

- Work with a variety of tools and technologies to ensure the reliability and performance of the Managed Services organization.

- Drive initiatives to optimize and secure infrastructure operations while contributing to scalability and continuous improvement efforts.

- Maintain and ensure operational excellence of systems and applications supporting Managed Services infrastructure.

- Troubleshoot and resolve issues in a fast-paced, distributed environment, focusing on root cause analysis and post-incident reviews.

- Build and maintain observability frameworks using tools like Prometheus, OpenTelemetry, and Dynatrace to ensure reliable monitoring of systems and applications.

- Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and manage Error Budgets to balance reliability and innovation.

- Use automation tools like Ansible and scripting languages like Python to reduce operational toil and optimize processes.

- Manage and troubleshoot Kubernetes clusters and containerized environments,



ensuring smooth service-to-service communication.

- Diagnose and resolve networking issues across OSI layers 1-3 on systems using including packet capture analysis.

- Oversee PKI certificate management and utilize tools like HashiCorp Vault for secrets management.

- Collaborate with cross-functional teams to identify system improvements and resolve critical issues.

- Provide mentoring and guidance to junior engineers, ensuring knowledge sharing and skilled growth.

Qualifications

Core Technical Skills

- Proficiency in Linux administration (system tuning, SSH, log analysis with tools like grep and regex).

- Strong understanding of networking protocols and troubleshooting (Layer 1-3).

- Hands-on experience with Kubernetes and container orchestration.

- Automation experience with Ansible and scripting proficiency in Python.

- Knowledge of PKI certificate management and HashiCorp Vault or similar tools for secrets management.

- Expertise in monitoring and observability tools like Prometheus, Grafana, or Dynatrace.

Reliability Engineering Skills

- Experience defining and managing SLIs, SLOs, and Error Budgets.

- Proven ability to instrument and analyze system performance metrics.

- Deep familiarity with troubleshooting distributed systems and microservices architecture.

Soft Skills

- Strong initiative and curiosity for deep-diving into complex issues.

- Excellent communication skills to convey solutions and collaborate across teams.

- Ability to prioritize and resolve competing priorities in high-pressure situations.

Preferred





- Experience with observability frameworks like OpenTelemetry.

- Familiarity with ITIL frameworks and best practices in incident and problem management.

- Background in enterprise environments, navigating corporate processes and bureaucracy.

- Security experience in areas like RBAC, least privilege principles, and secure software development lifecycle processes.

- Certifications in Kubernetes or relevant technologies.

How to Apply: Interested candidates are encouraged to respond/submit their updated resumes, and for additional job opportunities, please visit Jobs In India – VARITE.

Unlock Rewards: Refer Candidates and Earn.

If you're not available or interested in this opportunity, please pass this along to anyone in your network who might be a good fit and interested in our open positions. VARITE offers a Candidate Referral program, where you'll receive a one-time referral bonus based on the following scale if the referred candidate completes a three-month assignment with VARITE.

Experience Level Bonus Referral:

0-2 years

INR 5,000

2-6 years

INR 7,500

6+ years

INR 10,000

About VARITE: VARITE is a global staffing and IT consulting company providing technical consulting and team augmentation services to Fortune 500 Companies in USA, UK, CANADA and INDIA. VARITE is currently a primary and direct vendor to the leading corporations in the verticals of Networking, Cloud Infrastructure, Hardware and Software, Digital Marketing and Media Solutions, Clinical Diagnostics, Utilities, Gaming and Entertainment, and Financial Services.

Equal Opportunity Employer:

VARITE is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, marital status, veteran status, or disability status.

📌 Site Reliability Engineer (SRE) (Bengaluru)
🏢 VARITE
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (sre) (bengaluru) / bengaluru