06 Sep
|
NPCI
|
Hyderabad
The opportunity
NPCI is looking for a highly skilled DevOps Engineer specializing in automation, Kubernetes, and system reliability engineering, with a solid focus on AI-driven testing and resilience engineering. In this role, you will architect and execute automated reliability testing frameworks, simulate real-world failure scenarios, and ensure high availability, fault tolerance, and scalability of mission-critical payment systems operating at national scale.
This is a high-impact role where you'll work at the intersection of:
- DevOps + SRE + Chaos Engineering + AI-driven testing
- Ensuring platforms are failure-ready, not just failure-resistant
Job details
- Job Title: DevOps Engineer – Reliability & Automation
- Division: Quality Control & Monitoring Systems
- Experience: 3–7 Years
- Employment Type: Full-time
- Location: Hyderabad
- Education: Bachelor’s degree in Computer Science / Engineering or related field
- Preferred Certifications: CKA / CKAD / Cloud DevOps (AWS/Azure/GCP)
- Work Mode: 5 Days Work From Office (WFO)
Key Responsibilities
DevOps & Platform Engineering
• Design, implement, and manage CI/CD pipelines for application delivery and infrastructure provisioning
- Containerize applications using Docker and orchestrate via Kubernetes
- Deploy, manage, and optimize production-grade Kubernetes clusters
Reliability & Resilience Engineering
• Design and implement automated resilience and reliability testing frameworks
- Develop and execute failure simulation strategies (node failures, latency, network partitioning, etc.)
- Build and run fault injection mechanisms for controlled chaos testing
- Identify system weaknesses and improve fault tolerance & recovery strategies
Monitoring, Observability & Incident Readiness
• Implement monitoring, logging, and alerting solutions (Prometheus, Grafana, ELK/EFK, Open Telemetry)
- Define and track SLIs, SLOs, and error budgets
- Enhance observability for distributed systems
Infrastructure, Networking & Security
• Manage Linux systems and troubleshoot performance and reliability issues
- Configure networking components including:
➤ Load balancing
➤ DNS
➤ Firewalls & VPN
- Ensure security, compliance, and best practices across infrastructure
Collaboration & Productivity
• Work closely with SRE, DevOps, QA, and development teams
- Improve developer productivity through automation and self-service platforms
- Contribute to platform reliability culture and engineering excellence
Requirements
Required Technical Skills
Core DevOps Stack
Strong hands-on experience with:
- Docker & Kubernetes (Administration + Automation)
- CI/CD tools: Jenkins, GitLab CI, GitHub Actions, ArgoCD, Flux
Experience in Infrastructure as Code (IaC):
- Terraform, Ansible, Helm, Pulumi
Systems & Networking Expertise
Strong Linux administration skills (Ubuntu, CentOS, RHEL)
Deep understanding of networking
- TCP/IP, DNS, routing
- Load balancing, firewalls, VPN
Observability & Monitoring
Experience with
- Prometheus
- Grafana
- ELK/EFK stack
- Open Telemetry
Programming & Automation
Proficiency in
- Bash / Shell
- Python / Go
Strong scripting for automation and testing frameworks
Systems Thinking
Understanding of
- Distributed systems behavior
- Failure modes & recovery patterns
- Storage systems (Ceph, physical storage)
- Network stack (Overlay networking)
Good to Have Skills and Experience Required
Advanced Reliability Engineering
• Exposure to Chaos Engineering frameworks (Chaos Mesh, Litmus, Gremlin)
- Understanding of SRE practices and production reliability
AI-Driven Testing
• Experience with AI-powered testing or automation tools
- Familiarity with intelligent fault detection & predictive failure analysis
Distributed Systems Expertise
Understanding of
- System behavior under large-scale failures
- High-throughput, low-latency architectures
Security & Advanced Platform
Knowledge of
- Kubernetes security (RBAC, Network Policies)
- GitOps methodologies
- Advanced networking (CNI plugins, eBPF)
📌 Senior Associate ART Engineer (Hyderabad)
🏢 NPCI
📍 Hyderabad