27 Aug
|
Tavant
|
Bengaluru
A Tavant Reliability Engineering team has full vertical ownership of a system, from the server configuration up to the application interfaces. This enables the team to have full control of a service, avoiding situations where different teams own different areas of a system, causing some parts to fall between the cracks.
As systems and services grow in size and complexity, so too does the operational overhead. It is a fundamental principle of SRE to break this relationship between operational toil, system size, and complexity. Ultimately, fundamental software engineering skills coupled with strong systems and networking knowledge will guide the SRE to create more reliable
Key Responsibilities
Building software applications
- Responsible for building software applications using relevant development languages and applying knowledge of systems, services, and tools appropriate for the business area.
- Responsible for writing readable and reusable code by applying standard patterns and using standard libraries.
- Responsible for refactoring and simplifying code by introducing design patterns when necessary.
- Responsible for ensuring the quality of the application by following standard testing techniques and methods that adhere to the test strategy.
- Responsible for maintaining data security, integrity, and quality by effectively following company standards and best practices.
End-to-End System Ownership:
- Responsible for owning a service end-to-end by actively monitoring application health and performance, setting and monitoring relevant metrics, and acting accordingly when violated.
- Responsible for reducing business continuity risks by applying state-of-the-art practices and tools and writing appropriate documentation such as runbooks and OpDocs.
- Responsible for reducing risk and obtaining customer feedback by using continuous delivery and experimentation frameworks.
- Responsible for independently managing an application or service by working through deployment and operations in production.
Software Systems Design:
- Possesses sufficient knowledge to evaluate possible architectural solutions by taking into account cost, business requirements, technology requirements, and emerging technologies.
- Possesses sufficient knowledge to describe the implications of changing an existing system or adding a recent system to a specific area, by having a broad, high-level understanding of the infrastructure and architecture of our systems.
- Possesses sufficient knowledge to help grow the business and/or accelerate software development by applying engineering techniques (e.g., prototyping, spiking, and vendor evaluation) and standards.
- Possesses sufficient knowledge to meet business needs by designing solutions that meet current requirements and are adaptable for future enhancements.
Technical Incident Management:
- Responsible for addressing and resolving live production issues by mitigating customer impact within SLA.
- Responsible for improving the overall reliability of systems by producing long-term solutions through root cause analysis.
- Responsible for keeping track of incidents by contributing to postmortem processes and logging live issues.
Automation and toil reduction:
- Responsible for ensuring that infrastructure stays current by reducing technical debt, searching for bottlenecks, and preparing for scaling.
- Responsible for reducing the cost of operations and maintenance by leveraging new technologies, automation, and partnering with vendors to ensure we stay current.
- Responsible for reducing human labor by writing small software features that address availability, scalability, latency, and efficiency.
Monitoring and Alerting improvements:
- Responsible for reviewing and verifying the performance of production systems and network infrastructure by continuously monitoring appropriate observability metrics, business KPIs, and capacity planning.
- Responsible for improving application reliability by partnering with development teams to advise on setting appropriate observability metrics.
Architectural Guidance:
- Possesses basic knowledge to advise product teams towards a technical solution that meets the functional,
non-functional, and architectural requirements by challenging the rationale for an application design and providing context in the wider architectural landscape.
- Possesses basic knowledge to set a clear direction for a technical capability by evaluating and aligning target architecture improvements, reframing architectural designs, and decisions for wide-ranging stakeholders.
Critical Thinking:
- Responsible for systematically identifying patterns and underlying issues in complex situations, and for finding solutions by applying logical and analytical thinking.
- Responsible for constructively evaluating and developing ideas, plans, and solutions by reviewing them, objectively taking into account external knowledge, initiating SMART improvements, and articulating their rationale.
Continuous Quality and Process Improvement:
- Responsible for identifying opportunities for process, system, and structural improvements (i.e., performance gains) by examining and evaluating current process flows, methods, and standards.
- Responsible for designing and implementing relevant improvements by defining adapted/new process flows, standards, and practices that enable business performance.
Effective Communication:
- Responsible for delivering clear, well-structured, and meaningful information to a target audience by using suitable communication mediums and language tailored to the audience.
- Responsible for achieving mutually agreeable solutions by staying adaptable, communicating ideas in clear, coherent language, and practicing active listening.
- Responsible for asking relevant (follow-up) questions to properly engage with the speaker and truly understand what they are saying, by applying listening and reflection techniques.
- Responsible for technical implementation and maintenance of data processing services and storage systems in line with the Data Governance Framework.
- This role may require participation in an on-call rotation
Disclaimer: This has been sourced from a public domain and may have been modified by Naukri.com to improve clarity for our users. We encourage job seekers to verify all details directly with the employer via their official channels before applying.
📌 Site Reliability Engineer (Bengaluru)
🏢 Tavant
📍 Bengaluru