Production Support Engineer (Bengaluru)

Production Support Engineer (Bengaluru)

30 Sep
|
Honeybee Tech Solutions
|
Bengaluru

30 Sep

Honeybee Tech Solutions

Bengaluru

Role & responsibilities

Required Technical Skills Java &

- Application - Java 21 - Spring Boot 3 - REST APIs and HTTP - Microservices architecture - Ability to read and analyze Java stack traces - Application log analysis DynamoDB - Amazon DynamoDB - Queries and scans - Partition keys and sort keys - Indexes - Capacity and throttling - Data retrieval and validation AWS &
- Kubernetes - AWS fundamentals - Kubernetes - Pods, deployments, services, namespaces, ConfigMaps, and Secrets - Helm - Basic troubleshooting of AWS and Kubernetes environments Monitoring &
- Logging - Application and infrastructure monitoring - Logs, metrics, and alerts - Centralized logging - APM and observability tools - Distributed tracing is an advantage Azure DevOps - Azure DevOps (ADO) - Bug and work-item creation and management - Incident tracking - Investigation and resolution documentation Data &
- Reporting - Strong data-analysis skills - Ability to write queries and scripts for data extraction - Experience creating CSV and Excel-based reports - Ability to validate extracted data - Understanding of data security and access controls - Scripting experience with Python, Java, shell, or similar technologies is desirable DevOps &
- Tools - Azure devops - CI/CD fundamentals - Linux/Unix command line - API testing tools such as Postman Key Responsibilities Production Monitoring &
- Support We need Developer with Support role Experience in Development along with L1 Production Support.
- Monitor production applications, microservices, Kubernetes workloads, and AWS infrastructure.
- Monitor application health, availability, performance, error rates, response times, CPU and memory utilization, and other operational metrics.
- Monitor alerts and take appropriate first-level corrective actions.
- Perform daily production health checks and confirm that critical services are operational.
- Monitor scheduled jobs, background processes, integrations, and dependent services.
- Proactively identify potential production issues before they affect users.
- Maintain production-support checklists and operational runbooks.

Production

Exception &

• Incident Investigation - Investigate production exceptions, application errors, and failed transactions.
- Analyze application logs, stack traces, and error messages to determine the likely root cause.
- Perform first-level troubleshooting of Java 21 and Spring Boot 3 microservices.
- Correlate application errors with Kubernetes, AWS, DynamoDB, network, and downstream-service issues.
- Investigate recurring exceptions and identify opportunities for permanent remediation.
- Determine whether an issue is related to application code, configuration, database, infrastructure, or an external dependency.
- Escalate complex issues to L2/L3 teams with complete and accurate diagnostic information.
- Track incidents through resolution and verify that implemented fixes have resolved the issue. Bug Identification, Tracking &
- Management - Identify application defects and bugs during production monitoring and incident investigations.
- Create and maintain bugs and work items in Azure DevOps (ADO).
- Ensure that every identified production defect is appropriately documented and tracked.
- Capture relevant information, including issue description, business impact, setting, reproduction steps, logs, stack traces, investigation findings, and resolution details.
- Continuously update ADO work items with investigation progress and resolution status.
- Monitor open production bugs and follow up with L2/L3 teams through closure.
- Identify recurring defects and recommend permanent corrective actions. Report &




- Data Extraction - Respond to authorized business-user requests for operational and business reports and data extracts.
- Extract data from DynamoDB for reporting, investigation, and reconciliation purposes.
- Develop and execute appropriate queries or scripts to retrieve required data.
- Generate reports in agreed formats, such as CSV, Excel, or other approved formats.
- Perform data validation and sanity checks before distributing reports to users.
- Investigate discrepancies between application data and requested reports.
- Support ad hoc data extraction required for production incidents, audits, reconciliations, and business operations.
- Maintain reusable queries and scripts for frequently requested reports.
- Document the purpose, query logic, and expected output for recurring reports.
- Ensure that data extraction complies with access-control, data-security, privacy, and approval requirements.
- Ensure that sensitive or confidential information is not unnecessarily exposed in reports or support tickets.
- Escalate complex reporting requirements or high-volume data extraction requests to the appropriate development or data team.

Important: Data extraction must be performed using approved access and processes. Production data must not be modified as part of an L1 reporting activity unless explicitly authorized through the organizations change-control process.» User &

- Business Support - Serve as the first point of contact for application-related user queries.
- Maintain and update all runbooks for user and system queries. All common issues should have an entry in the runbook.
- Understand and troubleshoot user-reported application issues.
- Determine whether an issue is functional, technical, configuration-related, or data-related.
- Provide timely and accurate responses to business users.
- Troubleshoot failed transactions and application errors.
- Provide authorized reports and data extracts.
- Explain technical issues in clear, non-technical language when required.
- Log relevant user-reported issues in ADO.
- Convert recurring user queries or issues into knowledge articles, known issues, or bugs where appropriate.
- Track user issues through resolution and communicate status appropriately.

Application

Logging &

• Log Analysis - Monitor and analyze application logs to identify errors, exceptions, and abnormal behaviour.
- Investigate Java and Spring Boot stack traces and correlate errors across microservices.
- Use centralized logging and observability platforms to investigate production issues.
- Ensure that sufficient diagnostic information is captured for incidents and bugs.
- Record relevant log information and error details in ADO.
- Identify recurring error patterns and proactively raise defects where required.
- Work with development teams to improve application logging and troubleshooting capabilities.
- Ensure that sensitive information is not unnecessarily included in logs, ADO tickets, or support communications. Kubernetes &
- AWS Support - Monitor Kubernetes pods, deployments, services, and application workloads.
- Troubleshoot pod failures, restarts, readiness and liveness probe failures, and resource-related issues.
- Analyze Kubernetes events and application logs.
- Use “kubectl” for basic production troubleshooting.
- Understand and troubleshoot Helm-based deployments at an operational level.




- Perform basic AWS troubleshooting related to application availability and connectivity.
- Monitor application infrastructure and identify resource and capacity-related issues.
- Collaborate with cloud and infrastructure teams on complex AWS or Kubernetes issues. DynamoDB Support - Perform basic Amazon DynamoDB troubleshooting and investigation.
- Investigate application errors related to DynamoDB connectivity and data access.
- Understand tables, partition keys, sort keys, queries, scans, and indexes.
- Understand fundamental concepts such as read/write capacity, throttling, and consistency.
- Investigate DynamoDB-related errors and performance issues.
- Perform approved data-validation and data-retrieval activities.
- Support reporting and data-extraction requirements.
- Coordinate with development and cloud teams on complex DynamoDB issues. Incident, Problem &
- Change Management - Own L1 incidents from detection through resolution or escalation.
- Follow established incident-management and escalation procedures.
- Create and maintain incidents and bugs in Azure DevOps (ADO).
- Maintain accurate incident timelines and investigation notes.
- Participate in critical incident calls when required.
- Provide regular status updates during critical incidents.
- Support root-cause analysis and problem-management activities.
- Identify recurring incidents and recommend permanent corrective actions.
- Follow change-management, release-management, and production-deployment procedures.
- Maintain and continuously improve operational runbooks and knowledge articles. Automation &
- Continuous Improvement - Identify repetitive manual support activities and recommend opportunities for automation.
- Develop simple scripts and utilities to improve operational efficiency.
- Automate recurring reporting and data-extraction requirements where appropriate.
- Improve monitoring and alerting capabilities.
- Identify gaps in application logging and observability.
- Create dashboards and operational reports where required.
- Develop troubleshooting guides and knowledge articles.
- Identify recurring production issues and work toward their permanent resolution.

Key Performance Indicators

The candidate will be measured on:
- Production availability and operational stability - Maintain Runbooks for user queries and system issues.
- Incident response and escalation effectiveness - Incident resolution and escalation times - Quality and completeness of exception investigations - Percentage of incidents resolved at L1 - Quality and timeliness of ADO bug and incident logging - Reduction in recurring production incidents - Quality and timeliness of user responses - Timeliness and accuracy of requested reports and data extracts - Accuracy of data validation and reconciliation - Effectiveness of production monitoring and alert handling - Quality of incident documentation and runbooks - Automation of repetitive support and reporting activities - Successful implementation of approved minor fixes - Adherence to security, data-access, and change-management processes Role Expectations This is an L1 technical production-support role, not a traditional helpdesk position. The engineer is expected to manage the complete L1 support lifecycle: Monitor Detect Log Investigate Diagnose Fix and Resolve small bugs Extract and validate data Respond to users Escalate when required Track in ADO Validate Document Prevent recurrence Update Runbook The successful candidate should be capable of independently performing first-level technical investigations across the application, microservices, Kubernetes, AWS, DynamoDB, and logging and monitoring layers, while also handling user queries, production bug tracking, and authorized operational data and report extraction.

📌 Production Support Engineer (Bengaluru)
🏢 Honeybee Tech Solutions
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: production support engineer (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: production support engineer (bengaluru) / bengaluru