06 Aug
|
lambdatest is now testmu ai
|
Noida
06 Aug
lambdatest is now testmu ai
Noida
The Candidate Will Have Responsibilities Across The Following Functions
Platform Reliability (50%):
- Ensure robust cloud infrastructure and deployment pipelines.
- Improve observability tools and practices.
- Automate and streamline service workflows.
- Enhance platform systems for developer efficiency.
Infrastructure Engineering (30%):
- Build and maintain cloud-based systems.
- Optimize performance and scalability.
- Implement infrastructure improvements based on feedback.
Incident Management (20%):
- Lead incident response with strong SRE principles.
- Analyze and resolve production issues swiftly.
- Develop strategies to prevent future incidents.
Requirements
- Must have strong cloud, scripting, and core DevOps fundamentals.
- Must have hands-on experience with observability tools like Current Relic and Sumo Logic.
- Must demonstrate real coding ability.
- Must have a strong grasp of incident management and reliability/SRE thinking.
- Must explain service-layer architecture and reasoning effectively.
- Cloud Expertise: Designed and optimized cloud infrastructure for scalability and reliability.
- Coding Proficiency: Developed scripts and tools to automate deployment processes.
- Incident Management: Led a team to resolve critical production incidents effectively.
- Architecture Understanding: Explained complex service architectures to non-technical stakeholders.
This job was posted by Himanshii Tomer from TestMu AI.
📌 Platform Engineer (Noida)
🏢 lambdatest is now testmu ai
📍 Noida