Platform Engineer, AI Platform Operations (Maharashtra)

Platform Engineer, AI Platform Operations (Maharashtra)

16 Sep
|
Cephas Consultancy Services Private
|
Maharashtra

16 Sep

Cephas Consultancy Services Private

Maharashtra

AI Platform Operations Engineer (L1)

Position Summary

Join our 24x7 operations team supporting the Central AI Kitchen platform and its AI services. As an AI Platform Operations Engineer, you will be responsible for continuous monitoring, first-level incident detection and troubleshooting, executing approved recovery procedures, and escalating issues to the appropriate support teams. The platform runs AI services on Red Hat OpenShift, hosted on on-premise infrastructure. You will primarily use established dashboards, alerts, runbooks, and GitOps-based recovery procedures.

This entry-level role is ideal for engineers with foundational Linux, networking, and Kubernetes/OpenShift knowledge who are eager to develop advanced platform engineering and SRE skills.

Key Responsibilities

- 24x7 Platform Monitoring
- Monitor platform and application dashboards, alerts, and operational mailboxes.
- Identify availability, infrastructure, application, and service degradation events.
- Acknowledge alerts promptly and determine initial severity based on established procedures.
- Maintain accurate shift handover and operational records.

- First-Level Incident Troubleshooting

- Perform basic network connectivity checks (ping, DNS resolution, endpoint connectivity).
- Check OpenShift/Kubernetes resource status (nodes, pods, deployments, services).
- Review basic application and platform logs for common failure conditions.
- Check infrastructure and service health using approved dashboards and operational tools.
- Collect relevant diagnostic information before escalation.

- Incident Recovery

- Execute documented recovery procedures and operational runbooks.
- Restart or redeploy affected workloads using approved GitOps processes.




- Verify service recovery through dashboards, health checks, and application endpoints.
- Escalate when recovery procedures are unsuccessful or when an incident falls outside the approved operating scope.

- Incident Coordination & Escalation

- Create and maintain incident tickets with accurate timestamps, symptoms, actions, and observations.
- Engage appropriate application, platform, infrastructure, or network teams based on established escalation procedures.
- Provide transparent status updates during active incidents.
- Support incident bridges by providing operational information and executing actions requested by L2/L3 engineers.
- Ensure effective handover of unresolved incidents between shifts.

- Operational Procedures

- Follow established Standard Operating Procedures (SOPs), runbooks, and change-management processes.
- Document newly encountered symptoms and successful troubleshooting steps.
- Highlight recurring alerts or operational problems to senior platform engineers.
- Participate in operational drills and recovery exercises.

Scope of Authority
- Engineer may independently:
- Acknowledge and investigate alerts.
- Perform approved diagnostic commands and health checks.
- Execute documented L1 recovery procedures.
- Trigger approved GitOps redeployments.
- Open incidents and engage predefined support teams.




- Escalate incidents based on severity and runbook criteria.

- Engineer must escalate:

- Changes requiring manual modification of production infrastructure.
- Unauthorized configuration or source-code changes.
- Security incidents or suspected compromises.
- Infrastructure failures requiring RE intervention.
- Incidents where documented recovery procedures fail.
- Major incidents requiring business or management decisions.

Qualifications & Experience
- Essential
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline.
- 2 years of IT operations, infrastructure, cloud, or application support experience.
- Basic Linux command-line knowledge.
- Basic networking knowledge (IP addressing, DNS, ports, connectivity troubleshooting).
- Basic understanding of containers and Kubernetes concepts.
- Ability to follow technical procedures accurately.
- Good written and verbal communication skills.
- Willingness and ability to work in a 24x7 shift environment.

- Good to Have

- Exposure to Kubernetes or Red Hat OpenShift.
- Exposure to Git and GitOps concepts.
- Familiarity with monitoring tools (e.g., Grafana, Prometheus).
- Basic scripting experience (Bash or Python).
- Familiarity with incident-management or ITIL processes.
- Exposure to cloud or data-centre infrastructure.
- Interest in AI/ML infrastructure and GPU-based platforms.

Key Competencies
- Systematic troubleshooting
- Attention to detail
- Ability to remain structured during incidents
- Clear communication and escalation
- Discipline in following operational procedures
- Willingness to learn
- Teamwork across geographically distributed support teams

📌 Platform Engineer, AI Platform Operations (Maharashtra)
🏢 Cephas Consultancy Services Private
📍 Maharashtra

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: platform engineer, ai platform operations (maharashtra) / maharashtra

Subscribe to this job alert:

Get the latest job offers by email for: platform engineer, ai platform operations (maharashtra) / maharashtra