10 Aug
|
Cybage Software
|
Pune
10 Aug
Cybage Software
Pune
Experience: 8+ years .NET Application Reliability · Read, debug, and contribute to production C#/.NET code to diagnose and fix app-level reliability issues. · Identify and resolve memory leaks, thread pool exhaustion, and GC pressure before they manifest as incidents. · Partner with application engineers to embed reliability into new feature design and deployment practices. · Instrument .NET services with distributed tracing and structured logging to surface runtime anomalies early. AWS Compute &
• Infrastructure · Operate and optimize EC2 Auto Scaling, ECS Fargate, and Lambda workloads — with clear judgment on when each is the right fit. · Build and maintain infrastructure-as-code using CloudFormation or CDK for consistent, reproducible environments. · Automate operational tasks, deployment pipelines, and disaster recovery procedures. · Continuously reduce toil through tooling and automation, freeing the team for higher-impact engineering work. RDS SQL Server Operations · Manage RDS SQL Server deployments including Multi-AZ failover configuration and read replica setup. · Operate backup and point-in-time recovery (PITR) processes and validate restore procedures regularly. · Diagnose and resolve performance issues: slow queries, missing indexes, and blocking chains. · Capacity plan and scale database infrastructure to support transaction volume growth. Observability &
• Monitoring · Build and maintain observability stacks using CloudWatch metrics, log insights, and alarms
• AWS X-Ray for distributed tracing. · Own service health dashboards, SLOs/SLIs, and drive data-driven reliability improvements. · Design alerts that surface signal — not noise — and ensure on-call responders have the context to act quickly. · Conduct root cause analysis (RCA) on incidents and lead blameless post-mortems to capture lessons and prevent recurrence. Networking &
• Security · Design and maintain secure AWS network topologies: VPCs, subnets, security groups, and NACLs. · Configure and manage ALB/NLB routing, Route 53 DNS,
and TLS certificate lifecycle via ACM. · Author and review least-privilege IAM policies; audit roles and resource-based policies for over-permissioning. · Support compliance and security controls relevant to a PCI-regulated payments environment.
Incident
Response & On-Call · Participate in on-call rotation to respond to production incidents and drive swift resolution. · Define and track error budgets; use them to balance velocity and reliability investment. · Communicate status updates clearly during incidents and coordinate cross-functional response. · Maintain and improve runbooks, escalation paths, and on-call health over time. Cross-Functional Partnership · Collaborate with platform engineering teams on architecture decisions and scalability requirements. · Share observability and reliability best practices with application teams. · Mentor engineers on SRE principles and operational excellence. · 8+ years in SRE, DevOps, platform engineering, or a systems-focused software engineering role. · C#/.NET engineering ability — can read, debug, and contribute to production code;
experience diagnosing memory leaks, thread exhaustion, and GC pressure. · AWS compute fluency: hands-on depth across EC2 Auto Scaling, ECS Fargate, and Lambda, with informed opinions on when to use each. · RDS SQL Server operational experience: Multi-AZ failover, read replicas, backup/PITR, slow query analysis, and blocking chain resolution. · Native AWS observability proficiency: CloudWatch (metrics, logs, alarms), X-Ray, and infrastructure-as-code via CloudFormation or CDK. · AWS networking and security competence: VPCs, security groups, ALB/NLB, Route 53, TLS/ACM, and least-privilege IAM. · SLO discipline: experience defining SLIs/SLOs against real metrics, running blameless postmortems, and carrying an on-call pager. · Solid scripting ability (PowerShell, Python, or Bash) for automation and operational tooling. · Excellent communication skills and a collaborative, blameless engineering mindset. · Genuine openness to adopting AI tools and a willingness to experiment with new technology to work smarter and faster.
📌 Site Reliability Engineer (Pune)
🏢 Cybage Software
📍 Pune