07 Aug
|
National e Governance Division
|
New Delhi
07 Aug
National e Governance Division
New Delhi
Key Responsibilities
1. Own the platform infrastructure strategy for the AI capability programme across all six pods, covering CI/CD, environment provisioning, secrets management, observability, and reliability engineering
2. Design and maintain multi-workplace infrastructure (development, staging, production) with appropriate isolation, promotion gates, and security controls for classified, internal, and public workloads
3. Define infrastructure-as-code standards (Terraform, Ansible, or equivalent) and drive their adoption across pods
4. Establish the observability stack metrics, logs, tracing, alerting covering both AI-model and application layers, in coordination with the MLOps Lead
5. Own incident response protocols and on-call rotation across the programme; conduct post-incident reviews and drive corrective actions
6. Define availability and reliability SLOs for each pod's production services; track error budgets and coordinate with Product Managers on trade-offs
7. Oversee infrastructure security hardening network policies, WAF operations, DDoS protection, VAPT remediation at infrastructure layer in coordination with the Security & Compliance Engineer
8. Manage relationships with NIC, IndiaAI Compute, MeghRaj, and other government infrastructure providers; coordinate capacity planning and reserved-versus-on-demand mixing
9. Ensure Data Residency, Data Storage, and Data Lifecycle Management compliance per the LoE (data within India, AES-256 at rest, TLS 1.3 in transit, sanitisation at project closure)
10. Mentor DevOps/SRE Engineers and coordinate with MLOps Lead on shared platform responsibilities
Technical Competencies
1. Container & Orchestration: Docker, Kubernetes (production-grade), Helm, service mesh (Istio, Linkerd),
operators; certified Kubernetes administrator credentials preferred
2. Infrastructure as Code: Terraform, Ansible, Pulumi, or equivalent; GitOps workflows (Argo CD, Flux)
3. Cloud Platforms: AWS, Azure, GCP; government infrastructure (NIC, MeghRaj); familiarity with IndiaAI Compute service model
4. CI/CD: Jenkins, GitLab CI, GitHub Actions, Argo Workflows; artefact repositories (Nexus, Harbor)
5. Observability: Prometheus, Grafana, OpenTelemetry, Loki, Tempo, Jaeger; SLO/SLI methodology; incident management platforms
6. Infrastructure Security: Network security, secrets management (HashiCorp Vault, Infisical), certificate management, zero-trust patterns
7. Reliability Engineering: SRE principles, error budgets, capacity planning, chaos engineering basics, disaster recovery and business continuity
8. Scripting & Programming: Bash, Python, Go for tooling and automation
9. Compliance & Governance: MeitY Security Policy, ISO 27001, incident reporting timelines (24-hour breach reporting per LoE), audit trail management.
Experience
- 8+ years in infrastructure engineering, DevOps, or Site Reliability Engineering, with minimum 3 years in a lead or architect role
- Demonstrated experience running production infrastructure for high-availability platforms at scale
- Prior experience with government-hosted infrastructure (NIC, MeghRaj) or MeitY-empanelled CSPs preferred
- Experience with on-premise, hybrid, and sovereign-cloud deployment models
Educational Qualification
- B.Tech./B.E. in Computer Science, Information Technology, or related engineering discipline (Must have)
- M.Tech./M.S. in Computer Science, Information Technology, or related field desirable
- Certifications in DevOps, SRE, or Cloud Infrastructure (AWS, Azure, GCP, or Kubernetes CKA/CKS) preferred
- ITIL or equivalent service management certification desirable
📌 DevSecOps / SRE Lead (New Delhi)
🏢 National e Governance Division
📍 New Delhi