31 Jul
|
Pivotree
|
Bengaluru
31 Jul
Pivotree
Bengaluru
Position Summary:We are currently seeking a Senior Site Reliability Engineer (SRE) to join our team. Inthis role you will contribute to the reliability and enhancement of the technologyengine that powers multiple Pivotree solutions. The primary function of this role is thedirect responsibility for the availability of platform solutions, focusing on several keyareas, including availability, performance, change management, monitoring andemergency response.
You will work with other members of the platform, solutions,
operations, and application teams to understand and ultimately address changingand evolving requirements through extending and exposing capabilities in a simpleand consistent fashion. You will be a member of a team who maintains expertisewith Utility Computing services and will advise management and the organization asa whole on this mode of computing.
You will● Be responsible for the availability, performance, and reliability of platformservices.● Design and manage infrastructure in AWS, especially across multi-accountAWS Organization setups.● Lead efforts to automate cloud resource provisioning using IaC tools such asTerraform and AWS CloudFormation.● Contribute to ensuring pooled and independent utility services are highlyavailable● Actively take part and initiate continuous improvement: measure and reducemanual tasks and overhead● Be a subject matter expert for Utility Computing providers and respectiveservices both existing and emerging - with particular focus on AWS● Complete systems development, administration, and engineering tasksincluding integration, documentation and testing● Develop and maintain tools, processes, and workflows for automatedinfrastructure resource(s) and application deployment, configurationmanagement & maintenance● Own the responsibility for platform management, supporting services, and allrelated tooling and automation● Design, implement, and maintain monitoring, alerting, and observabilitysolutions to ensure system reliability, performance, and timely incidentdetection.● Investigate and troubleshoot relevant platform-based issues and incidents,(high availability, performance, security,
etc.)● Collaborate across distributed teams and act as a technical liaison for variousstakeholders.● Participate in change management, release automation, compliance, andaudit-readiness practices.
● Participate in recurring stand-ups with other team members located indifferent locations and time zones● Participate in on-call rotation, escalations, and shift work● Work with other team members to improve processes and advance relevantand related competencies● Provide technical mentorship, helping teammates grow through knowledgesharing and peer coaching.You are
● Experienced in production-grade AWS Cloud environments with a focus onscalability, security, and governance.● Well-versed in Linux (RHEL/Debian) environments and proficient in systemadministration.● Adept at supporting modern development workflows, Agile teams, and CI/CDtooling● A strong communicator, with the ability to interface with technical and non-technical stakeholders across geographies.● Motivated to mentor others and foster a cooperative engineering culture.● Capable of operating independently and making strategic infrastructuredecisions.● Passionate about reliability engineering and operational excellence.● Experienced at working on large projects with deadlines● Committed to high quality and attention to detail● Focused and committed to delivering high quality services● A strategic thinker who is able to link business and technical objectives● Someone that can go wide and deep, who works with several disparatesystems and services and ultimately acquires expert knowledge and who cannavigate accordingly
You have (MUST HAVE)
● 5+ years of experience in Site Reliability Engineering, Cloud Engineering, orDevOps roles.
● Minimum one Associate-level Amazon AWS certification.● 3+ years mature,
production level experience with infrastructure-as-codeconcepts and practices using Terraform, AWS CloudFormation or similar.● 3+ years of hands-on experience managing Kubernetes clusters (EKSpreferred), including container orchestration and troubleshooting.● Strong knowledge and practical experience in Linux systems administration(RHEL and Debian-based distros), networking, storage, and virtualization inproduction-grade environments.● Experience working with API-driven and/or Event-driven architectures at scalein AWS environments.● Demonstrated ability to manage and troubleshoot web applications,middleware, and databases in real-world deployments.● Expertise with observability stacks and performance monitoring using toolslike Grafana, Prometheus, CloudWatch, and Loki or similar.● Advanced scripting proficiency in Python, Bash, and basic PowerShell.● Solid experience with CI/CD pipelines, version control, and automated testingusing Git, Bitbucket, GitHub, Jenkins, or similar.● Proven track record in implementing security and compliance controls,particularly in regulated environments (SOC 2, PCI-DSS, or ISO 27001).● Strong understanding of systems security, including identity, permissions,network policies, and audit tooling.● Exceptional troubleshooting skills, attention to detail, and a strong drive forcontinuous improvement.● Excellent communication skills with the ability to clearly articulate complexconcepts to both technical and non-technical audiences, and collaborateeffectively across distributed teams.● Demonstrated ability to mentor junior engineers, share knowledge, and act asa regional technical leader.● Ability to work both independently and collaboratively, learn new technologiesquickly, and help set standards and best practices.
Nice to Have
● Experience and/or exposure to the Serverless Framework
● Experience with APM tools such as AppDynamics, NewRelic, Grafana orDynatrace, Amazon X-Ray● Experience with the following Amazon AWS services in a productionenvironment (API Gateway, Cognito, RDS, DynamoDB, ECS, EMR, Lambda)● AWS Certified Developer● AWS Certified SysOps Administrator● AWS Certified Solution Architect
📌 Senior Site Reliability Engineer (Bengaluru)
🏢 Pivotree
📍 Bengaluru