The Candidate Will Have Responsibilities Across The Following Functions Platform Reliability (50%): Ensure robust cloud infrastructure and deployment pipelines.
Improve observability tools and practices.
Automate and streamline service workflows.
Enhance platform systems for developer efficiency.
Infrastructure
Engineering (30%): Build and maintain cloud-based systems.
Optimize performance and scalability.
Implement infrastructure improvements based on feedback.
Incident
Management (20%): Lead incident response with strong SRE principles.
Analyze and resolve production issues swiftly.
Develop strategies to prevent future incidents.
Requirements Must have strong cloud, scripting, and core DevOps fundamentals.
Must have hands-on experience with observability tools like New Relic and Sumo Logic.
Must demonstrate real coding ability.
Must have a solid grasp of incident management and reliability/SRE thinking.
Must explain service-layer architecture and reasoning effectively.
Cloud Expertise: Designed and optimized cloud infrastructure for scalability and reliability.
Coding Proficiency: Developed scripts and tools to automate deployment processes.
Incident Management: Led a team to resolve critical production incidents effectively.
Architecture Understanding: Explained complex service architectures to non-technical stakeholders. This job was posted by Himanshii Tomer from TestMu AI.
📌 Platform Engineer (Noida)
🏢 lambdatest is now testmu ai
📍 Noida
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.