skillset : Java , Observability (ELF , Grafan , Splunk) , Github , Service now(ticketing) , AWS/GCP
Replacement of Harikaran.
Roles & Responsibilities:
• Provide hands-on support for the runtime operation of our applications, ensuring high availability and performance.
• Collaborate with software engineering and infrastructure teams to troubleshoot and resolve runtime issues, including performance bottlenecks, scalability challenges, and system failures.
• Contribute to the design and implementation of monitoring, alerting, and logging solutions to proactively identify and address potential runtime issues.
• Participate in incident response and root cause analysis efforts to ensure the stability and resilience of the applications.
• Work closely with cross-functional teams to understand application requirements and provide input on runtime and operational considerations during the software development lifecycle.
• Contribute to the development and maintenance of runtime automation and tooling to streamline operational processes and improve efficiency.
• Develop common framework components (to be leveraged by enterprise applications), define standards for configuration, monitoring, reliability, and performance engineering
• Create automation and ensure automated tests are completed for recent features.
• Valuable attitude, communication, willingness to learn and collaborate.
• Continuously improve automated remediation tasks to ensure the highest levels of availability.
• Cloud: Manage secure, scalable, and highly available cloud infrastructure.
• Kubernetes & Containers: Deploy, operate, and troubleshoot containerized workloads.
• Observability: Implement monitoring, logging, tracing, dashboards, and actionable alerts.
• Reliability: Define SLOs/SLIs, manage error budgets, and improve service availability.
• Networking: Troubleshoot DNS, TCP/IP, HTTP/S, TLS, routing, and load balancing.
• Linux & Systems: Administer and troubleshoot Linux systems and performance issues.
•
📌 Consultant Chennai
🏢 Atos
📍 Chennai