12 Sep
|
The TJX Companies
|
Hyderabad
12 Sep
The TJX Companies
Hyderabad
What Youll Do The Senior Staff Engineer – Cloud Operations will lead complex cloud operations initiatives, improve reliability, reduce operational risk, and help mature how cloud services are supported at scale.
Key responsibilities include:
What You’ll Need
- Lead medium-to-very-high complexity cloud operations initiatives across Azure-first platforms, shared services, multi-cloud integrations, and production support capabilities.
- Own operational readiness for cloud services, including monitoring, alerting, incident response, runbooks, recovery patterns, release readiness, and support models.
- Improve reliability by analyzing incidents, telemetry, service health, capacity signals, and recurring operational issues across Azure, GCP, OCI, and related services.
- Use AI-enabled operations tooling, including SRE agents where appropriate, to improve incident triage, correlation, response, remediation guidance, and resolution time.
- Lead complex incident response and troubleshooting across cloud infrastructure, AKS, networking, identity, security tooling, application platforms, vendor-managed services, and multi-cloud dependencies.
- Improve incident, problem, and change management practices, including post-incident reviews, root cause analysis, corrective actions, and knowledge management.
- Design and implement automation to reduce toil, validate platform health, accelerate recovery, and improve repeatable operational execution across cloud environments.
- Define and influence cloud operations standards, reference patterns, and roadmap priorities that improve reliability, reduce incident recurrence, improve mean time to resolution, and mature operational practices across Azure-first multi-cloud platforms.
- Partner with engineering teams to define production readiness, operational acceptance criteria, support handoffs, and environment validation practices.
- Champion DevSecOps, site reliability engineering, infrastructure as code, policy-as-code, AI-assisted operations, and operational excellence practices.
- Maintain operational standards, runbooks, playbooks, diagrams, reusable patterns, and service support documentation.
- Mentor engineers, lead technical discussions, and contribute to onboarding, interviews, and knowledge-sharing activities.
- We are looking for a hands-on technical leader with robust cloud operations experience, sound engineering practices, and a customer-focused mindset.
- The ideal candidate brings deep Azure operations experience, understands multi-cloud operating models across GCP and OCI, and is comfortable leading complex troubleshooting,
creating durable fixes, mentoring engineers, and raising operational maturity across the cloud platform.
Required qualifications and experience include:
- 8+ years of engineering experience in cloud, infrastructure, platform engineering, site reliability engineering, or a related technical domain.
- Hands-on experience operating enterprise-scale Microsoft Azure environments, including compute, networking, identity, storage, monitoring, security, and platform services.
- Baseline familiarity with Google Cloud Platform and Oracle Cloud Infrastructure concepts, with enough operational exposure to collaborate in a multi-cloud environment and apply consistent reliability, support, and operational practices.
- Experience with incident, problem, and change management, including production readiness, operational support models, post-incident reviews, and corrective actions.
- Experience operating and troubleshooting Azure Kubernetes Service, including cluster health, node pools, ingress, networking, identity integration, observability, and platform upgrades.
- Strong knowledge of Azure networking, including virtual networks, network security groups, route tables, Azure Firewall, private endpoints, DNS, load balancing, and hybrid connectivity troubleshooting.
- Experience with monitoring, alerting, logging, dashboards, service health checks, operational reporting, and AI-enabled operations tools used to support incident triage and resolution.
- Hands-on automation experience using PowerShell, Python, Bash, Terraform, Bicep, ARM, GitHub Actions, Azure DevOps pipelines, or similar tools.
- Strong understanding of DevSecOps, CI/CD, infrastructure as code, policy-as-code, release controls, security, compliance, RBAC, least privilege, secrets management, and privileged access practices.
- Ability to lead complex troubleshooting, translate operational data into engineering fixes, and drive measurable reliability improvements.
- Experience creating technical documentation, runbooks, support models, reusable patterns, and operational standards.
- Demonstrated ability to mentor engineers, influence technical direction, communicate clearly with stakeholders,
and lead complex work independently in an Agile environment.
Preferred Qualifications
Great to have:
- Experience applying site reliability engineering practices, including service level indicators, service level objectives, error budgets, toil reduction, reliability reviews, and operational maturity assessments.
- Experience supporting Azure Landing Zones, enterprise-scale cloud operating models, subscription management, policy enforcement, and shared platform services.
- Hands-on experience supporting or improving GCP and/or OCI operations, including platform monitoring, identity and access management, networking, security controls, landing zone patterns, automation, and operational support models.
- Experience with Kubernetes operations and security, including network policies, workload identity, pod security, image scanning, certificate rotation, and cluster lifecycle management.
- Experience using or integrating AI-enabled SRE agents, AIOps platforms, or intelligent automation to improve incident detection, correlation, remediation recommendations, and mean time to resolution.
- Experience improving production readiness through release gates, operational acceptance criteria, environment validation, support handoffs, and service transition practices.
- Understanding of cloud resiliency practices, including backup and restore, disaster recovery, high availability patterns, capacity management, and dependency mapping.
- Exposure to cost management, tagging, capacity planning, utilization optimization, quota management, and operational reporting.
- Experience supporting global platforms, regulated environments, or business-critical technology services.
Key Competencies
- Technical leadership and engineering judgment
- Cloud operations and reliability mindset
- Ownership, urgency, and operational discipline
- Incident leadership and problem-solving
- Automation and continuous improvement
- DevSecOps and secure engineering practices
- Stakeholder communication and influence
- Mentoring and knowledge sharing
- Collaboration across global and cross-functional teams
- Agile ways of working
Minimum Education Bachelor's degree in computer science, Information Technology, Engineering, or a related discipline, or equivalent practical experience.
Minimum Experience
- 8+ years of relevant engineering experience in cloud, infrastructure, platform engineering, site reliability engineering, technology operations, or related technical areas.
- Experience working in large-scale enterprise environments with production-critical technology services.
📌 Senior Staff Engineer Cloud (Hyderabad)
🏢 The TJX Companies
📍 Hyderabad