29 Aug
|
Alion
|
Anupgarh
You will be responsible for the operations and reliability of our SaaS application and the underlying infrastructure on GCP, with a focus on automation and continuous improvement. You will tackle the most complex technical challenges, mentor other engineers, and set the standard for operational excellence. This is not a traditional system administration role; you are expected to write code and build systems to eliminate manual work.
Key Responsibilities
- Design and implement scalable, self-healing infrastructure on GCP using Terraform and Kubernetes (GKE).
- Develop automation tooling to reduce operational toil and improve system efficiency.
- Provide technical guidance and support to the application team on application reliability, performance, observability, and operational best practices.
- Lead the technical response during major incidents, performing deep-dive analysis to identify root causes.
- Architect and implement comprehensive monitoring, logging, and alerting solutions to ensure proactive issue detection.
- Participate in the 24/7 on-call rotation.
- Drive post-mortem analysis and ensure follow-up actions are implemented.
Required Qualifications
- 5+ years of experience in a Site Reliability, DevOps, or Software Engineering role with a focus on infrastructure.
- Strong, hands-on experience with GCP, specifically with GKE, VPC, Cloud Load Balancing, IAM,
and Cloud Monitoring.
- Proficient in writing production-quality code for automation (Python or Go preferred).
- Understanding of Java application operations.
- Expertise in Terraform for managing complex infrastructure.
- In-depth knowledge of Kubernetes architecture and operations in a production environment.
- Experience with CI/CD systems (e.g., GitLab CI) and integrating operational controls into pipelines.
- A systematic problem-solving approach, coupled with strong communication skills and a sense of ownership.
What We Offer
- Participation in a strategic Cloud & SaaS transformation in a global company.
- Long-term growth opportunities opportunities within an international organization.
- The chance to design and develop PSI's proprietary products.
- A team of highly experienced specialists eager to share knowledge.
- Comfortable office environment (small rooms, chillout space, parking).
- Stability and security of employment in a company with a 50-year tradition.
- Flexible working hours and a friendly atmosphere with no artificial hierarchy.
- Well-defined career development paths.
- Access to technology conferences, training, and language courses (English and German).
- Benefits package: private medical care, group insurance, benefits platform, Multisport card
📌 Senior Site Reliability Engineer (SRE) (Anupgarh)
🏢 Alion
📍 Anupgarh