06 Aug
|
Cybage Software
|
Maharashtra
06 Aug
Cybage Software
Maharashtra
Experience: 8+ years
.NET Application Reliability
· Read, debug, and contribute to production C#/.NET code to diagnose and fix app-level reliability issues.
· Identify and resolve memory leaks, thread pool exhaustion, and GC pressure before they manifest as incidents.
· Partner with application engineers to embed reliability into new feature design and deployment practices.
· Instrument .NET services with distributed tracing and structured logging to surface runtime anomalies early.
AWS Compute & Infrastructure
· Operate and optimize EC2 Auto Scaling, ECS Fargate, and Lambda workloads — with clear judgment on when each is the right fit.
· Build and maintain infrastructure-as-code using CloudFormation or CDK for consistent, reproducible environments.
· Automate operational tasks, deployment pipelines, and disaster recovery procedures.
· Continuously reduce toil through tooling and automation, freeing the team for higher-impact engineering work.
RDS SQL Server Operations
· Manage RDS SQL Server deployments including Multi-AZ failover configuration and read replica setup.
· Operate backup and point-in-time recovery (PITR) processes and validate restore procedures regularly.
· Diagnose and resolve performance issues: slow queries, missing indexes, and blocking chains.
· Capacity plan and scale database infrastructure to support transaction volume growth.
Observability & Monitoring
· Build and maintain observability stacks using CloudWatch metrics, log insights, and alarms; AWS X-Ray for distributed tracing.
· Own service health dashboards, SLOs/SLIs, and drive data-driven reliability improvements.
· Design alerts that surface signal — not noise — and ensure on-call responders have the context to act quickly.
· Conduct root cause analysis (RCA) on incidents and lead blameless post-mortems to capture lessons and prevent recurrence.
Networking & Security
· Design and maintain secure AWS network topologies: VPCs, subnets, security groups, and NACLs.
· Configure and manage ALB/NLB routing, Route 53 DNS,
and TLS certificate lifecycle via ACM.
· Author and review least-privilege IAM policies; audit roles and resource-based policies for over-permissioning.
· Support compliance and security controls relevant to a PCI-regulated payments environment.
Incident Response & On-Call
· Participate in on-call rotation to respond to production incidents and drive swift resolution.
· Define and track error budgets; use them to balance velocity and reliability investment.
· Communicate status updates clearly during incidents and coordinate cross-functional response.
· Maintain and improve runbooks, escalation paths, and on-call health over time.
Cross-Functional Partnership
· Collaborate with platform engineering teams on architecture decisions and scalability requirements.
· Share observability and reliability best practices with application teams.
· Mentor engineers on SRE principles and operational excellence. · 8+ years in SRE, DevOps, platform engineering, or a systems-focused software engineering role.
· C#/.NET engineering ability — can read, debug, and contribute to production code;
experience diagnosing memory leaks, thread exhaustion, and GC pressure.
· AWS compute fluency: hands-on depth across EC2 Auto Scaling, ECS Fargate, and Lambda, with informed opinions on when to use each.
· RDS SQL Server operational experience: Multi-AZ failover, read replicas, backup/PITR, slow query analysis, and blocking chain resolution.
· Native AWS observability proficiency: CloudWatch (metrics, logs, alarms), X-Ray, and infrastructure-as-code via CloudFormation or CDK.
· AWS networking and security competence: VPCs, security groups, ALB/NLB, Route 53, TLS/ACM, and least-privilege IAM.
· SLO discipline: experience defining SLIs/SLOs against real metrics, running blameless postmortems, and carrying an on-call pager.
· Strong scripting ability (PowerShell, Python, or Bash) for automation and operational tooling.
· Excellent communication skills and a team-oriented, blameless engineering mindset.
· Genuine openness to adopting AI tools and a willingness to experiment with new technology to work smarter and faster.
📌 Site Reliability Engineer (Maharashtra)
🏢 Cybage Software
📍 Maharashtra