03 Sep
|
NetSkope Software
|
Kolkata
03 Sep
NetSkope Software
Kolkata
Job Summary
Were looking for an engineer to ensure the reliable operation of production environments for our Data Infrastructure and products, running at scale on large-volume distributed cloud systems. Youll focus on maximizing system reliability, automating routine tasks, and improving production efficiency. This role offers hands-on exposure to modern cloud technologies - Docker, Kubernetes, networking, and platforms like AWS and GCP - while you contribute directly to system uptime and user experience for a large-scale distributed application.
Youll lead production monitoring and incident response, drive automation to reduce manual work, and collaborate with Engineering on root cause analysis, CI/CD support, and capacity planning.
Whats in it for you:
In this role, you will be responsible for seamless operation of production environments for our Data Infrastructure and products within large-scale, high-volume distributed cloud systems. You will concentrate on maximizing system reliability, automating routine tasks, and maintaining the efficiency of our production systems.
This role provides practical cloud exposure to help you deepen your expertise in modern technology stacks, such as Docker, Kubernetes, Networking, and major public cloud providers like AWS and GCP. You will have the chance to contribute to enhancing user experience and system uptime for a large-scale distributed application.
What you will be doing:
- Monitoring: Use observability dashboards to monitor system performance, error rates, and resource utilization. Define current dashboards and alerts as and when required and write technical runbooks for Incident response.
- Incident Response: Act as the initial point of contact for production alerts,
mitigate/solve the ongoing issues, escalate highly complex problems to the Engineering team, and conduct thorough root cause analysis for incidents.
- Automation: Identify repetitive manual tasks and streamline them through automation.
- Support CI/CD: Help create, test, and maintain automated application deployment pipelines.
- Capacity Management: Assist in tracking system resource usage (CPU, memory, storage) and traffic volume to help forecast and scale infrastructure needs
Required skills and experience:
- 5+ years of overall industry experience in a relevant technical role
- Proficiency in at least one scripting language, preferably Python
- Hands-on experience with containerization and orchestration technologies such as Docker, Kubernetes, and related cluster concepts
- Strong understanding of public cloud infrastructure, with a preference for AWS (GCP experience also considered)
- Working knowledge of Linux/Unix command-line environments and core web protocols including HTTP, gRPC, DNS, and TCP/IP
- Experience with observability and monitoring tools such as Grafana and Prometheus is a plus
- Familiarity with Infrastructure as Code (IaC) using Terraform, along with workflow automation via GitHub Actions, is a plus
- Willingness to participate in an on-call rotation and respond to production incidents with flexibility to support critical systems
Education
- BSCS or equivalent required, MSCS or equivalent strongly preferred
Disclaimer: This job posting and Location has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Sr. DevOps Engineer, Data Platform (Kolkata)
🏢 NetSkope Software
📍 Kolkata