17 Sep
|
TRIGENT SOFTWARE PRIVATE
|
Hyderabad
17 Sep
TRIGENT SOFTWARE PRIVATE
Hyderabad
We are seeking a Senior Systems Engineer (IC3) to join our Hyperscaler Performance & Reliability Engineering (HPRE) team. This role is focused on improving the reliability, performance, and operational excellence of infrastructure services running across AWS, Azure, and GCP.
Operating within a Site Reliability Engineering (SRE)-inspired model, you will own infrastructure-level reliability challenges spanning compute, storage, and networking services. You will investigate production issues, develop performance insights from telemetry and observability platforms, validate infrastructure designs through testing, and implement engineering improvements that reduce customer-impacting incidents.
This is a hands-on engineering role for someone who enjoys troubleshooting complex distributed systems, analyzing system behavior at scale, and turning operational learnings into lasting reliability improvements. The ideal candidate combines robust cloud infrastructure experience with a data-driven approach to performance tuning and operational excellence.
Key Responsibilities
- Own reliability and performance investigations for compute, storage, and network-related issues across AWS, Azure, and GCP, driving incidents from detection through root-cause analysis and remediation.
- Design, implement, and maintain observability solutions using hyperscaler-native monitoring platforms and internal telemetry systems to identify bottlenecks, capacity constraints, and performance regressions.
- Analyze customer workload behavior using metrics, logs, traces, and infrastructure telemetry to optimize resource utilization, latency, throughput, and availability.
- Develop and maintain Infrastructure-as-Code, automation, and operational tooling that improve reliability, reduce toil,
and enable safe infrastructure changes at scale.
- Execute performance, load, stress, and resilience testing of cloud infrastructure platforms and validate infrastructure changes before production adoption.
- Identify and drive reliability improvements, operational processes, and technical projects within the team while serving as a trusted technical resource for peers and partners.
Skills:
Required Qualifications
- Strong understanding of cloud compute, storage, and networking fundamentals, including virtual machines, block storage, VPC/VNet design, load balancing, routing, and access management.
- Experience implementing and operating observability platforms, including metrics, logging, tracing, alerting, and performance analysis workflows.
- Proficiency with Infrastructure-as-Code and automation technologies such as Terraform, Git-based workflows, CI/CD pipelines, and scripting with Python or Bash.
- Demonstrated ability to independently troubleshoot complex production issues, perform structured root-cause analysis, and implement preventative solutions.
- Experience working within SRE, platform engineering, infrastructure engineering, or cloud operations environments where reliability, scalability, and operational accountability are core responsibilities.
Preferred Qualifications
- Experience with Kubernetes, GitOps, and declarative infrastructure management models.
- Familiarity with CloudWatch, Azure Monitor, Google Cloud Monitoring, Splunk, OpenTelemetry, or similar observability ecosystems.
- Cloud platform or Kubernetes certifications (AWS, Azure, GCP, CKA, CKAD, or equivalent).
Education:
3+ years of hands-on experience operating and troubleshooting production infrastructure in at least one major public cloud, with working knowledge of AWS, Azure, and/or GCP services.
📌 Senior Systems Engineer || Hyderabad || Hybrid
🏢 TRIGENT SOFTWARE PRIVATE
📍 Hyderabad