19 Sep
|
VARITE
|
Hyderabad
Company Name: VARITE India Private Limited
About The Client
A cloud computing company offers a platform for digital workflows, enabling organizations to automate and streamline business processes. The solutions include IT service management, human resources, customer service, and security operations. Designed to enhance efficiency and collaboration, the platform digitizes and automates workflows for diverse organizational needs.
Headquartered in the United States, the company is recognized for its innovative approach to workflow automation, playing a significant role in the IT service management and business process automation space. In 2018, Forbes magazine named it number one on its list of the world's most innovative companies. About The Job:
- We are seeking a Senior Systems Engineer (IC3) to join our Hyperscaler Performance & Reliability Engineering (HPRE) team.
- This role is focused on improving the reliability, performance, and operational excellence of infrastructure services running across AWS, Azure, and GCP.
- Operating within a Site Reliability Engineering (SRE)-inspired model, you will own infrastructure-level reliability challenges spanning compute, storage, and networking services.
- You will investigate production issues, develop performance insights from telemetry and observability platforms, validate infrastructure designs through testing, and implement engineering improvements that reduce customer-impacting incidents.
- This is a hands-on engineering role for someone who enjoys troubleshooting complex distributed systems, analyzing system behavior at scale, and turning operational learnings into lasting reliability improvements.
- The ideal candidate combines strong cloud infrastructure experience with a data-driven approach to performance tuning and operational excellence.
Essential Job Functions:
- Own reliability and performance investigations for compute, storage, and network-related issues across AWS, Azure, and GCP, driving incidents from detection through root-cause analysis and remediation.
- Design, implement,
and maintain observability solutions using hyperscaler-native monitoring platforms and internal telemetry systems to identify bottlenecks, capacity constraints, and performance regressions.
- Analyze customer workload behavior using metrics, logs, traces, and infrastructure telemetry to optimize resource utilization, latency, throughput, and availability.
- Develop and maintain Infrastructure-as-Code, automation, and operational tooling that improve reliability, reduce toil, and enable secure infrastructure changes at scale.
- Execute performance, load, stress, and resilience testing of cloud infrastructure platforms and validate infrastructure changes before production adoption.
- Identify and drive reliability improvements, operational processes, and technical projects within the team while serving as a trusted technical resource for peers and partners.
Qualifications:
Required Qualifications
- Strong understanding of cloud compute, storage, and networking fundamentals, including virtual machines, block storage, VPC/VNet design, load balancing, routing, and access management.
- Experience implementing and operating observability platforms, including metrics, logging, tracing, alerting, and performance analysis workflows.
- Proficiency with Infrastructure-as-Code and automation technologies such as Terraform, Git-based workflows, CI/CD pipelines, and scripting with Python or Bash.
- Demonstrated ability to independently troubleshoot complex production issues, perform structured root-cause analysis, and implement preventative solutions.
- Experience working within SRE, platform engineering, infrastructure engineering, or cloud operations environments where reliability, scalability, and operational accountability are core responsibilities.
Preferred Qualifications
- Experience with Kubernetes, GitOps, and declarative infrastructure management models.
- Familiarity with CloudWatch, Azure Monitor, Google Cloud Monitoring, Splunk, OpenTelemetry, or similar observability ecosystems.
- Cloud platform or Kubernetes certifications (AWS, Azure, GCP, CKA, CKAD, or equivalent).
Keywords:
- Education: 3+ years of hands-on experience operating and troubleshooting production infrastructure in at least one major public cloud, with working knowledge of AWS, Azure, and/or GCP services.
How to Apply: Interested candidates are encouraged to respond/submit their updated resumes, and for additional job opportunities, please visit Jobs In India – VARITE.
Unlock Rewards: Refer Candidates and Earn.
If you're not available or interested in this opportunity, please pass this along to anyone in your network who might be a good fit and interested in our open positions. VARITE offers a Candidate Referral program, where you'll receive a one-time referral bonus based on the following scale if the preferred candidate completes a three-month assignment with VARITE.
Experience Level Bonus Referral: 0-2 years INR 5,000 2-6 years INR 7,500 6+ years INR 10,000
About VARITE: VARITE is a global staffing and IT consulting company providing technical consulting and team augmentation services to Fortune 500 Companies in USA, UK, CANADA and INDIA. VARITE is currently a primary and direct vendor to the leading corporations in the verticals of Networking, Cloud Infrastructure, Hardware and Software, Digital Marketing and Media Solutions, Clinical Diagnostics, Utilities, Gaming and Entertainment, and Financial Services.
Equal Opportunity Employer:
VARITE is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate based on race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, marital status, veteran status, or disability status.
📌 Systems Engineer - III (Hyderabad)
🏢 VARITE
📍 Hyderabad