05 Aug
|
Deutsche Bank
|
India
05 Aug
Deutsche Bank
India
Job Description:
Job Title: Technology Operations Specialist
Corporate Title: VP
Location: Pune, India
Role Description
You will be operating within the Production Support Services team of the TST domain, spanning SES, TAS and Trade Finance Lending within Corporate Bank Production Services. The role is focused on providing hands-on production support, troubleshooting and operational ownership for applications hosted on Google Cloud Platform (GCP), with strong emphasis on Google Kubernetes Engine (GKE), containerised workloads, PostgreSQL/Postgres-backed services and related cloud-native technologies. You will be expected to lead incident recovery, perform deep technical analysis across application, infrastructure, database and platform layers, improve observability, and drive automation to enhance service stability, resilience and operational efficiency.
- Lead technical and functional troubleshooting for production incidents, user requests and platform issues across GCP-hosted applications, including GKE, Cloud Run, Compute Engine, Cloud Storage, IAM, networking, PostgreSQL/Postgres databases and database connectivity.
- Drive end-to-end incident recovery for P3H and above incidents by analysing logs, metrics, traces, deployment history, configuration changes, infrastructure events and application behaviour.
- Use GCP observability tools such as Cloud Logging, Cloud Monitoring, Error Reporting, dashboards, uptime checks and alert policies to identify root cause, reduce mean time to detect and improve service reliability.
- Partner with application, infrastructure, SRE, security and engineering teams to troubleshoot cloud networking, container runtime, IAM, quota, performance, scaling and availability issues.
- Drive automation, toil reduction, platform hygiene, monitoring improvements and operational controls for cloud-native applications and supporting infrastructure.
- Prepare and distribute clear incident communications, technical updates, recovery timelines and service restoration summaries for senior stakeholders.
- Participate in CAB, change validation and post-change support activities to assess operational risk, ensure rollback readiness and support safer change delivery.
- Collaborate with Global Incident Management, L2/L3 support, SRE and platform teams to orchestrate recovery of major incidents and implement preventive actions.
What we’ll offer you
As part of our versatile scheme, here are just some of the benefits that you’ll enjoy,
- Best in class leave policy.
- Gender neutral parental leaves
- 100% reimbursement under childcare assistance benefit (gender neutral)
- Sponsorship for Industry relevant certifications and education
- Employee Assistance Program for you and your family members
- Comprehensive Hospitalization Insurance for you and your dependents
- Accident and Term life Insurance
- Complementary Health screening for 35 yrs. and above
Your key responsibilities
- Provide hands-on technical and functional production support for applications deployed on GCP and integrated enterprise platforms within the TST domain.
- Troubleshoot complex production issues across cloud runtime, containers, application services, middleware, databases, network connectivity, IAM permissions, certificates, secrets and deployment pipelines.
- Analyse GCP logs,
metrics and alerts using Cloud Logging, Cloud Monitoring, dashboards and log-based metrics to identify root cause and restore service quickly.
- Support containerised workloads running on GKE and Cloud Run, including pod/container restarts, scaling behaviour, node pressure, cluster events, health checks, readiness/liveness failures, resource saturation, ingress/service issues and deployment rollbacks.
- Troubleshoot PostgreSQL/Postgres production issues, including connection failures, query performance, locks, replication or failover symptoms, storage growth, backup/restore readiness and application-to-database connectivity.
- Build technical and functional subject matter expertise across supported applications, business flows, cloud architecture, service dependencies and infrastructure configuration.
- Identify proactive opportunities for automation, self-healing, alert rationalisation, toil reduction and improved operational resilience.
- Drive service requests and incidents to resolution within L2 scope, ensuring timely escalation to L3, SRE, infrastructure or vendor teams where deeper engineering intervention is required.
- Review monitoring coverage for critical services, SLIs, SLOs, error budgets, availability, latency, throughput, saturation and business-critical transaction flows.
- Maintain and continuously improve runbooks, KEDB articles, support procedures, troubleshooting guides and operational readiness documentation.
- Participate in BCP, DR, EDR, SSRTO and component failure tests for cloud and application services, validating recovery procedures and operational readiness.
- Understand data flow, request flow and service dependencies across cloud infrastructure to provide effective operational support during incidents and changes.
- Own managed risk, operational controls, production hygiene and continuous improvement initiatives for supported services.
- Mentor and coach global support team members on GCP troubleshooting, incident handling, observability and production support best practices.
- Apply an SRE mindset with practical understanding of SLIs, SLOs, error budgets, incident learning and reliability improvement.
Your skills and experience
Must Have : -
- 12+ years of IT experience in large corporate environments, with strong exposure to controlled production support environments, preferably within Financial Services Technology.
- Strong hands-on experience supporting and troubleshooting production workloads on Google Cloud Platform, with deeper focus on GKE, Cloud Run, Compute Engine, Cloud Storage, Cloud SQL/PostgreSQL, IAM, VPC networking, load balancers and service accounts.
- Proven ability to investigate incidents using Cloud Logging, Cloud Monitoring, Error Reporting, dashboards, alerts, log-based metrics and application traces. Strong troubleshooting skills across application, infrastructure and platform layers, including container failures, scaling issues, latency,
memory/CPU saturation, network connectivity, DNS, certificates, secrets, IAM permissions and deployment failures.
- Strong hands-on experience with Kubernetes concepts and GKE operations, including pods, deployments, services, ingress, node pools, namespaces, autoscaling, config maps, secrets, health checks, cluster events, workloads, logs and kubectl-based diagnostics.
- Working knowledge of cloud-native application architecture, microservices, REST APIs, service-to-service communication, event-driven flows and enterprise integration patterns.
- Strong understanding of UNIX/Linux operating systems, shell commands, process analysis, file systems, logs, network utilities and infrastructure troubleshooting.
- Working knowledge of scripting and automation using UNIX shell, Python, PowerShell, Perl or similar tools for diagnostics, reporting, remediation and toil reduction.
- Understanding of middleware and messaging platforms such as MQ, Kafka or similar, including operational troubleshooting of connectivity, message flow and performance issues.
- Understanding of web and application server environments such as Apache, Tomcat, WebLogic or equivalent cloud-hosted runtimes.
- Strong understanding of relational databases, especially PostgreSQL/Postgres and Cloud SQL for PostgreSQL, including connectivity, query performance, locks, indexing concepts, backup/restore, storage growth, availability and operational troubleshooting.
- Experience with enterprise monitoring tools such as Geneos, AppDynamics, New Relic, Dynatrace, Grafana or equivalent, alongside GCP observability tools. Strong understanding of ITIL Service Management practices, including Incident, Problem, Change, Request and Major Incident Management.
- Experience in securities services, asset management or trade finance domains will be an advantage.
- Flexibility to support incident callouts after office hours and during weekends, including participation in rotational weekend cover.
Nice to Have:
- Good analytical and problem-solving skills.
- Excellent communication skills, both written and verbal, with attention to detail.
- Ability to work in virtual teams and in matrix structures.
- Team management and program management skills
- ITIL / best practice service context. ITIL foundation is plus.
- Ticketing Tool experience – Service Desk, Service Now.
- Understanding of SRE concepts (SLA, SLO’s, SLI’s)
- Knowledge and development experience in Ansible automation.
- Working knowledge of one cloud platform (AWS or GCP).
How we’ll support you
- Training and development to help you excel in your career
- Coaching and support from experts in your team
- A culture of continuous learning to aid progression
- A range of flexible benefits that you can tailor to suit your needs
About us and our teams
Please visit our company website for further information:
https://www.db.com/company/company.html
We strive for a culture in which we are empowered to excel together every day. This includes acting responsibly, thinking commercially, taking initiative and working collaboratively.
Together we share and celebrate the successes of our people. Together we are Deutsche Bank Group.
We welcome applications from all people and promote a positive, fair and inclusive work environment.
📌 Technology Operations Specialist, VP (India)
🏢 Deutsche Bank
📍 India