06 Sep
|
R-LOGICS SOLUTIONS
|
Chennai
06 Sep
R-LOGICS SOLUTIONS
Chennai
Platform Engineering | DevSecOps | Kubernetes Security | SRE | AI Infrastructure
ROLE TYPE - Permanent, full-time | Hybrid / location aligned to project needs
CAREER PATH - Platform Automation Engineer → AI Platform SRE → Principal Platform Engineer
CERTIFICATION - Kubernetes security certification (CKS) mandatory - held already or completed during probation
WHO WILL THRIVE HERE - DESIGN ENGINEER
Turns an outcome into architecture, controls, automation, tests and operating evidence - rather than waiting for step-by-step tasks.
SECURITY-FIRST BUILDER
Treats identity, secrets, network boundaries, software supply chain and runtime security as part of the platform design.
SYSTEMS THINKER
Can move between Kubernetes, Linux, networking, storage, GPU and application behaviour to find the real cause of a problem.
ROLE PURPOSE The RIGORIX Platform AI Automation Engineer will help engineer, secure, automate and operate the infrastructure underpinning the RIGORIX™ Enterprise AI Factory. The role deliberately sits across platform engineering and DevSecOps, with a planned progression toward Site Reliability Engineering and AI-assisted operations.
THE ENGINEERING MINDSET A requirement such as “deploy RIGORIX securely onto Kubernetes” should naturally lead you to think about topology, failure domains, ingress and egress, tenancy, identity, certificate lifecycle, secrets, network policy, persistent storage, GPU scheduling, supply-chain controls, observability, upgrade strategy, rollback, disaster recovery and reproducibility.
Requirement → Architecture → Threat Model → Design → Automation → Test → Evidence → Documentation → Runbook
CORE RESPONSIBILITIES
ADVANCED KUBERNETES ENGINEERING
Design production-grade clusters and deployments; Helm packaging; RBAC; Pod Security; network policies; quotas; affinities; disruption budgets; autoscaling; storage; upgrades; multi-tenancy and air-gap patterns.
PLATFORM AUTOMATION
Build reusable, version-controlled deployment profiles using Helm, Terraform/OpenTofu, GitOps, CI/CD and scripting. Manual configuration should steadily become technical debt.
DEVSECOPS INTEGRATION
Integrate security into build and release workflows: SAST, SCA, dependency and image scanning, IaC/Kubernetes scanning, secret detection, SBOMs, policy gates, signed artifacts and remediation automation.
AI INFRASTRUCTURE
Engineer GPU node pools and inference environments; understand GPU scheduling, drivers, CUDA compatibility, VRAM, MIG where applicable, model placement, throughput and latency.
SERVICE-TO-SERVICE SECURITY
Engineer service mesh, mTLS, workload identity, certificate rotation, ingress/egress control, authorization policies and distributed telemetry.
OPERATIONAL ENGINEERING
Build observability, resilience, upgrade/rollback, backup/recovery, runbooks and automated verification so deployments are supportable at enterprise scale.
CLOUD-AGNOSTIC DEPLOYMENT EXPECTATION
RIGORIX must not become architecturally dependent on a single hyperscaler. You should be able to pick up an unfamiliar environment, understand its compute, network, storage, identity and security constraints, and adapt the platform safely.
- Azure AKS, AWS EKS and Google GKE
- VMware / private cloud / sovereign cloud Kubernetes
- Bare-metal and GPU appliance deployments
- Disconnected and air-gapped environments
- Hardware-agnostic Kubernetes patterns with environment-specific profiles
02 | SECURITY, DEVSECOPS & PRODUCTION HARDENING
SECURITY IS MANDATORY,
NOT DELEGATED A platform engineer changing a service, Helm chart or cluster control should automatically consider exposure, TLS, authentication, authorization, secrets, attack surface, container privileges, dependencies, network policy and auditability.
RIGORIX is designed around explicit trust boundaries and controlled service access; the Kubernetes implementation must preserve and strengthen those properties.
MANDATORY KUBERNETES SECURITY CERTIFICATION
Certified Kubernetes Security Specialist (CKS) is mandatory for this role. Candidates who do not already hold it must complete the required certification path during probation. RLogics will provide access to training, but successful completion remains an individual role requirement.
DEVSECOPS SECURITY AUTOMATION
Using integrated platforms such as Aikido Security as an example of the desired experience, the goal is to continuously assess application, dependency, container, infrastructure-as-code, cloud/Kubernetes and runtime risk - and increasingly turn findings into controlled automated engineering workflows.
- SAST, SCA and dependency vulnerability management
- Container/image scanning, CVE prioritisation and hardened base images
- IaC, Helm and Kubernetes manifest security scanning
- Secret detection, SBOM generation, provenance and license risk
- API and application security checks
- Policy-driven security gates before promotion
- Automated evidence collection and remediation workflows
ARTIFACT TRUST
SBOM generation, vulnerability scanning, image signing, signature verification and provenance before deployment.
PROMOTION CONTROL
Build → Evaluation → Run separation with controlled approvals, quarantine, release gates and rollback.
DRIFT & EVIDENCE
Git as source of truth, configuration drift detection, auditable deployment history and machine-readable evidence.
IDENTITY, SECRETS AND ZERO TRUST
- Keycloak / OIDC / OAuth2 and enterprise identity federation
- OpenBao / Vault concepts, agile credentials and short-lived secrets
- PKI, internal CA, certificate lifecycle and workload identity
- OPA/policy enforcement and admission controls
- mTLS and least-privilege service communication
- RBAC / project or tenant isolation and secure machine identities
Service mesh: experience with Istio, Istio Ambient, Envoy or equivalent is highly desirable, including mTLS, workload identity, policy, certificates, traffic controls, observability and troubleshooting.
03 | PLATFORM DEPTH, SRE & AGENTIC OPERATIONS A role that reaches below Kubernetes and beyond traditional DevOps.
LINUX AND NETWORK ENGINEERING DEPTH
LINUX ENGINEERING
Processes, systemd, permissions, namespaces, cgroups, capabilities, filesystems, sockets, DNS, routing, nftables/iptables, kernel tuning, I/O, memory, CPU and performance troubleshooting.
NETWORK ENGINEERING
IP/subnetting, routing, NAT, DNS, TLS, load balancing, firewalls, L4/L7, CNI, NetworkPolicy, Gateway API, MTU, east-west and north-south traffic.
TROUBLESHOOTING
Comfortable following a fault through kubelet, runtime, DNS, CNI, storage, service mesh, Linux and application layers rather than stopping at “pod unhealthy”.
AUTOMATION LANGUAGES
Strong Bash and/or Python; Go is desirable. Comfortable reading APIs, creating tooling and using automation rather than relying only on portals.
MOVING FROM DEVOPS TOWARD SRE
During the next 12 months, the role will deliberately move toward Site Reliability Engineering. We want engineers who design for measurable reliability and recoverability, not only deployment success.
- Define SLIs and SLOs for critical platform services
- Establish availability, latency and recovery objectives
- Create runbooks and incident response automation
- Perform capacity and performance baselining
- Engineer high availability, backup and recovery
- Test failure scenarios and validate self-healing behaviour
OBSERVABILITY ENGINEERING A production AI platform requires layered telemetry. You will work with OpenTelemetry, Prometheus/Grafana and related tooling across infrastructure, Kubernetes, GPU, application and AI execution signals.
- Infrastructure and node health
- Kubernetes workload health and saturation
- GPU utilisation, memory, temperature and error signals
- API and dependency latency
- Agent execution, tool calls and workflow duration
- Model/inference latency, throughput and failure behaviour
AGENTIC DEVOPS - AN R&D; PRIORITY
Alert → AI-assisted investigation → Kubernetes/log/metric checks → probable cause → policy check → human approval where required → controlled remediation → verification → incident evidence The goal is not unrestricted LLM access to production. It is controlled AI-assisted operations built around least-privilege tools, policy gates, deterministic automation and human approval for sensitive actions.
R&D; MINDSET
You may be asked to evaluate technologies such as eBPF, Cilium, NATS, policy engines, GPU schedulers, Kubernetes sandboxing, confidential computing or distributed inference.
A useful investigation should end with:
Problem → Requirements → Alternatives → PoC → Security/Performance Assessment → Recommendation
04 | WHAT GOOD LOOKS LIKE - FIRST 12 MONTHS
Clear expectations for growth, ownership and impact.
0-3 MONTHS | FOUNDATION
Understand RIGORIX architecture and security model; deploy the platform to Kubernetes; work safely with Helm/Git; troubleshoot Linux/Kubernetes; begin or complete mandatory security certification.
3-6 MONTHS | AUTOMATION
Own deployment components; build secure reusable Helm and GitOps automation; integrate security gates; complete at least two repeatable platform deployments in different environments.
6-9 MONTHS | PLATFORM
Contribute to service mesh, identity/secrets, network policy, high availability, observability, GPU profiles, backup/recovery and cloud portability.
9-12 MONTHS | SRE + AGENTIC
Define SLOs and runbooks; improve automated recovery and diagnostics; contribute to Agentic DevOps and security remediation; lead significant production-readiness workstreams.
BY 12 MONTHS, A STRONG ENGINEER CAN...
Discovery → Design → Infrastructure Automation → Security → Deployment → Validation → Documentation → Operational Handover.
CANDIDATE & TECHNICAL PROFILE
- Strong: Linux, Kubernetes, containers, Git/CI-CD, scripting and networking
- Good: Helm, Terraform/OpenTofu, security, PKI/secrets, GitOps and observability
- Desirable: Istio/Envoy, Cilium, Keycloak, OpenBao/Vault, OPA, GPU/NVIDIA Kubernetes and OpenTelemetry
- Essential: strong communication, documentation, collaboration and willingness to work across technology boundaries
📌 AI Automation Engineer (Chennai)
🏢 R-LOGICS SOLUTIONS
📍 Chennai