24 Sep
|
Qualcomm
|
Chennai
Company: Qualcomm India Private Limited
Job Area: Engineering Group, Engineering Group > Software Engineering
Job Summary
We are looking for a Senior Site Reliability Engineer to join a 24/7, follow-the-sun SRE team responsible for keeping our critical production systems reliable, scalable, and secure around the clock. As part of a globally distributed team spanning multiple time zones, you will share on-call coverage that follows daylight hours rather than night shifts, and partner closely with development teams to automate operations, strengthen observability, and continuously improve system resilience.
Responsibilities
- System Reliability: Ensure the reliability, availability, and performance of critical systems.
- Service Level Objectives: Define, measure, and report on SLIs, SLOs, and error budgets, and use them to prioritize reliability work.
- Automation: Develop and maintain automation scripts and tools to streamline operations.
- Monitoring: Develop and maintain monitoring dashboards alerts
- Incident Management: Lead incident response efforts and post-mortem analysis to prevent future occurrences.
- Performance Tuning: Optimize system performance and scalability.
- Cost Optimization: Monitor and optimize cloud spend (e.g., right-sizing and autoscaling with Karpenter) to balance reliability with cost efficiency.
- Security: Implement and maintain security best practices.
- Documentation: Create and maintain comprehensive documentation for systems and processes.
- Mentorship: Mentor junior engineers and champion SRE best practices across cross-functional teams.
- On-Call Shifts: Own front-line 24/7 on-call rotations and incident response for critical production systems, acting as the reliability shield for the platform and observability engineering teams so they are not paged for production incidents.
Requirements
- Education: Bachelor's degree in Computer Science, Engineering, or a related field. Advanced degrees are a plus.
- Experience: 5+ years of experience in a similar role, with a strong background in software engineering and systems administration.
Technical Skills
- Programming Languages: Proficiency in one or more programming languages such as Python, Go.
- Cloud Platforms: Extensive experience with AWS cloud platform.
- Infrastructure as Code: Hands-on experience with tools like Terraform, Ansible, or CloudFormation.
- Containerization and Orchestration: Expertise in Docker and Kubernetes. Kubernetes - Hands-on experience is a must.
- Kubernetes technologies (with preferences): ArgoCD, Linkerd, Prometheus, Karpenter, etc
- Monitoring and Logging: Proficiency with monitoring tools like Prometheus, Grafana, and logging tools like ELK stack or Loki stack.
- CI/CD Pipelines: Experience with continuous integration and continuous deployment tools such as Jenkins, GitLab CI, or CircleCI.
- Networking: Strong understanding of networking concepts, protocols, and security.
Minimum Qualifications
- Bachelor's degree in Engineering, Information Systems, Computer Science, or related field and 2+ years of Software Engineering or related work experience.
- OR Master's degree in Engineering, Information Systems, Computer Science, or related field and 1+ year of Software Engineering or related work experience.
- OR PhD in Engineering, Information Systems, Computer Science, or related field.
- 2+ years of academic or work experience with Programming Language such as C, C++, Java, Python, etc.
Soft Skills
- Problem-Solving: Excellent analytical and troubleshooting skills.
- Communication: Strong verbal and written communication skills.
- Collaboration: Ability to work effectively in a team setting and collaborate with cross-functional teams.
- Leadership: Proven leadership skills and the ability to mentor junior engineers.
- Remote Work: Comfortable working in a fully distributed, offshore setup and collaborating effectively with development teams across multiple locations.
📌 Senior Site Reliability Engineer (Chennai)
🏢 Qualcomm
📍 Chennai