Key Accountabilities:
- Code Contribution : Review and contribute to code in both application and infrastructure stacks.
- Operational Efficiency : Promote automation and develop recent tools to improve operational processes.
- Incident Management : Participate in on-call incident management, from event handling to root cause analysis and postmortems.
- Continuous Improvement : Provide operations-focused feedback to product teams for process improvement.
Skills & Experience Required:
- Experience : 2+ years in software/systems engineering, preferably in AWS.
- Scripting & Programming : Proficient in shell scripting and at least two high-level languages (e.g., Python, Golang, Bash).
- Agile : Experience working in an Agile Scrum environment.
- DevOps : Deep knowledge of DevOps practices (e.g., 12-factor apps, Infrastructure-as-Code, 'shift left').
- Observability :
Experience designing and implementing infrastructure and application observability (e.g., Datadog, Terraform).
Preferred Experience:
- Complex Environments : Engineering experience in a globally available, complex environment.
- Greenfield Processes : Comfortable with contributing to and developing new processes.
- On-call & SRE : Experience with on-call rotations and site reliability engineering.
- Cloud Expertise : Experience with multi-cloud providers (AWS, Azure, GCP) and HashiCorp Terraform.
- Kubernetes : Familiarity with Kubernetes & Helm deployments, GHA.
- Customer Relationships : Build strong relationships with customers and consistently exceed expectations.