02 Aug
|
Charles Schwab
|
Hyderabad
02 Aug
Charles Schwab
Hyderabad
What you have
Role Description:
Advisor Wealth & Asset Technology (AWAT) is an organization within Schwab Technology Services aligned to support the technology needs of Registered Independent Advisor (RIA), Schwab Asset Management (SAM), Wealth & Advice Solutions (WAS), the Schwab Center for Financial Research (SCFR), and Third-Party Platform teams.
Join a team that is redefining operational excellence at Schwab. As part of AWAT Technology Operations, you'll play a key role in supporting critical Advisory, Wealth and Asset Management platforms while driving the next generation of Site Reliability Engineering practices. As a Support Engineer, you will be part of a collaborative production support team focused on keeping critical applications and platforms stable and highly available, resilient, and ready to support business growth.
You will apply your technical expertise to solve complex operational challenges, strengthen system reliability, and improve the experience of both internal and external business partners. With a strong focus on automation, observability, and continuous improvement, you will help shape a more proactive, scalable SRE support model. If you are energized by ownership, up-to-date support practices, incident management, and building smarter tools that make a real difference, this is a great opportunity to grow your career with Schwab.
Required Qualifications:
- 3 to 5 years of experience in Site Reliability Engineering, Production Support, DevOps, or Systems Engineering supporting highly available, client-facing enterprise applications.
- Understanding of SRE principles, including reliability, scalability, resiliency, observability, automation, toil reduction, and operational excellence.
- Hands-on experience with incident management, problem management, change management, root cause analysis, post-incident reviews, and continuous service improvement practices.
- Experience defining, measuring, and improving SLIs, SLOs, SLAs, KPIs, availability, performance, and reliability metrics for production systems.
- Working knowledge of observability and monitoring platforms such as Splunk, Datadog, Grafana Cloud, OpenTelemetry, ThousandEyes, or similar tools, including dashboards, alerting, logs, metrics, and traces.
- Familiar with one or more programming or scripting languages such as Python, PowerShell, C#, Java, JavaScript, Bash,
or Shell to automate operational tasks and support troubleshooting.
- Familiar with CI/CD and engineering collaboration tools such as GitHub, Bitbucket, Bamboo, Atlassian products
- Experience supporting applications hosted on cloud or platform environments such as PCF, AWS, GCP, Azure, Kubernetes, or container-based platforms.
- Familiar with relational and non-relational databases such as SQL Server, PostgreSQL, MongoDB, or similar technologies, including query execution, scripting, and production support.
- Robust troubleshooting skills across applications, infrastructure, integrations, end-user workflows, and distributed systems, with the ability to diagnose complex production issues.
- Understanding of incident life cycle and strong background in incident and problem manage
- Strong written and verbal communication skills, with the ability to clearly explain technical issues, risks, workarounds, and resolution plans to technical and non-technical stakeholders.
- Ability to participate in rotational on-call support, respond to production incidents, and drive timely restoration of service while maintaining clear stakeholder communication.
- Experience working in a regulated enterprise environment, preferably financial services, with a focus on secure, compliant, and reliable technology operations.
- Bachelors degree in Computer Science, Information Technology, Engineering, or equivalent practical experience.
- Experience supporting Asset Management, Wealth Management, or Financial Services platforms, with knowledge of the trade lifecycle, investment operations, and front-to-back office workflows.
- Strong ownership mindset with the ability to drive incident response, service restoration, root cause analysis, post-incident reviews, and long-term corrective actions.
- Practical experience with observability, monitoring, alerting, dashboards, logs, metrics, and traces using tools such as Splunk, Datadog, Grafana, OpenTelemetry, ThousandEyes, or similar platforms.
- Ability to automate repetitive operational tasks, reduce toil, improve runbooks, and build scripts or tools using Python, PowerShell, Bash, Java, or similar technologies.
- Understanding of SRE practices such as SLIs, SLOs, SLAs, error budgets, reliability reporting, capacity awareness, performance monitoring, and operational risk reduction.
- Experience supporting applications across cloud, container, or platform environments such as AWS, Azure, GCP, Kubernetes, PCF, or similar enterprise platforms.
- Working knowledge of databases, APIs, integrations, job scheduling, file transfers, CI/CD pipelines, and release/change management processes in a production environment.
- Collaborative, curious, and adaptable problem solver who partners effectively with engineering, product, business, and operations teams while respectfully challenging the status quo.
- Strong verbal and written communication skills, with the ability to explain technical issues, business impact, workarounds, timelines, and resolution plans to technical and non-technical stakeholders.
- Exposure to AI-driven, data-assisted, or analytics-based operational practices that improve triage, anomaly detection, incident response, knowledge management, or continuous service improvement.
- Ability to participate in rotational on-call support and work effectively under pressure while maintaining clear priorities, sound judgment, and customer-focused communication.
Our base benefits, wellbeing, and total rewards include:
- Competitive compensation and retirement programs including Employee Provident Fund (EPF), Gratuity, and optional National Pension System (NPS) contributions
- Robust Paid Time Off, including annual/privilege leave, sick and casual leave, public holidays, maternity/paternity leave, and more
- Education assistance for continued learning to help you grow
- Comprehensive medical insurance with Outpatient Department (OPD) services, including vaccination, pharmacy, dental, and vision coverage
- Annual reimbursement for health check-ups and mental health support through our Employee Assistance Program (EAP)
- Childcare (creche) reimbursement for eligible employees
- Transportation and meal benefits that support your day-to-day work
- Group life, personal accident, and critical illness insurance
📌 Specialist, Software Development & Engineering - SRE (Hyderabad)
🏢 Charles Schwab
📍 Hyderabad