14 Aug
|
Tata Consultancy Services
|
Bengaluru
14 Aug
Tata Consultancy Services
Bengaluru
Job Location : Bangalore,Hyderabad
Experience - 6+ Only
Job Requirements
- Drive operational stability across ERP and other finance applications
- Implement and enhance automation and tooling to reduce manual effort and improve efficiency
- Own and execute Disaster Recovery (DR) planning and testing to ensure business continuity
- Lead service design and service transition activities for current and existing systems
- Manage incident, problem, and change processes aligned with SRE and ITIL practices
- Establish effective service communication frameworks for incidents and outages
- Drive continual service improvement (CSI) initiatives across finance systems
- Define and manage service metrics, SLAs, SLIs, and reporting dashboards
- Collaborate with engineering teams to improve system reliability, observability, and performance
- Ensure smooth onboarding and ownership of satellite applications within the finance ecosystem
- Proactively identify risks and implement preventive measures to minimize production issues
Preferred Skills and Experience
- Solid experience in Site Reliability Engineering / Production Support / DevOps roles
- Experience supporting Oracle Cloud ERP or similar enterprise SaaS platforms
- Knowledge of monitoring, alerting, and observability tools (e.g., Splunk, Grafana, OCI monitoring)
- Experience in incident management, RCA, and problem management
- Exposure to automation frameworks and scripting (e.g., Python, Shell, Terraform)
- Understanding of cloud platforms (OCI/AWS/Azure) and distributed systems
- Familiarity with ITIL processes and service management frameworks
- Experience working in global, distributed teams
- Exposure to financial systems and processes is an advantage
Key Responsibilities:
- Operational Stability & SRE Practices
- Maintain high system availability through proactive monitoring and incident prevention
- Define and track SLIs/SLOs to measure service health and reliability
- Lead root cause analysis (RCA) and implement preventive fixes
- Automation & Tooling
- Build and maintain automation for repetitive operational tasks
- Improve deployment pipelines and operational workflows
- Develop scripts/tools to enhance productivity and reduce human intervention
- Service Design & Transition
- Ensure new services are designed with reliability, scalability, and supportability in mind
- Lead service transition activities, including documentation, readiness, and handover
- Disaster Recovery & Resilience
- Develop and execute DR strategies and testing plans
- Ensure systems meet recovery objectives (RTO/RPO)
- Service Management & Communication
- Manage incident, change, and problem processes
- Ensure clear and timely communication during production incidents
- Collaborate with stakeholders during outages and recovery
- Continual Service Improvement (CSI)
- Identify and implement improvements to systems and processes
- Drive reliability engineering best practices across teams
- Service Metrics & Reporting
- Define and maintain dashboards for system health and performance
- Provide insights and reporting to leadership for decision-making
- Satellite Application Ownership
- Own end-to-end support for finance-related satellite applications
- Ensure integration stability between ERP and downstream/upstream systems
📌 Site Relibility Enginneer (SRE) (Bengaluru)
🏢 Tata Consultancy Services
📍 Bengaluru