11 Sep
|
Infinityquest It Services
|
Bengaluru
11 Sep
Infinityquest It Services
Bengaluru
: Platform / API Engineer (GCB5, India)
Role purpose
We are seeking a Platform / API Engineer to design, build, and maintain RESTful APIs that expose enterprise infrastructure capabilities, with a strong emphasis on Day-2 operational automation. The role sits within our Platform & API Engineering function and focuses on delivering secure, scalable, and automated tooling that accelerates developer and infrastructure workflows.
Role & responsibilities
- Platform/API engineering
- Design, develop, and maintain RESTful APIs using Python (FastAPI) to expose infrastructure capabilities (e.g., provisioning, configuration, compliance checks, lifecycle operations) to internal consumers.
- Define API contracts, versioning, authentication/authorisation, rate limiting, auditing, and developer-facing documentation (OpenAPI/Swagger).
- Package and deploy services using Docker and Kubernetes/OpenShift, following secure engineering practices.
- OpenShift + Kubernetes engineering
- Engineer and support solutions on Red Hat OpenShift, including cluster and workload lifecycle operations.
- Plan and execute OpenShift Day-2 activities, including:
- Cluster upgrades (control plane/worker upgrades, operator lifecycle, compatibility checks, upgrade windows, rollback/restore approach)
- Security patching and vulnerability remediation aligned to enterprise controls
- Capacity management (quota/limits, node sizing, autoscaling approaches, utilization monitoring)
- Backup/restore and DR readiness (where applicable to platform scope)
- Operational readiness: runbooks, monitoring/alerting, SLOs, incident response support
- Enterprise infrastructure lifecycle automation
- Build automation for common infrastructure operations, such as:
- OS upgrades / patch orchestration (upgrade planning, pre-checks, dependency validation, scheduling, post-validation, audit evidence)
- File system expansion automations (volume extension workflows, LVM/filesystem resize steps, validation, and safe guards)
- Standardized health checks, compliance validations, and automated remediation where appropriate
- Integrate API/automation services with enterprise components (identity, secrets, certificates, networking, observability/logging, CMDB/ITSM workflows).
- Reliability & operability
- Ensure services meet non-functional requirements: resilience, scalability, performance, availability, auditability, and security-by-design.
- Troubleshoot production issues across API, container, OS, and platform layers; drive RCA and permanent fixes.
- Collaborate with SRE/operations, security, and engineering teams to deliver self-service platform capabilities that reduce manual operational toil.
- Developer productivity / AI (good to have)
- Prototype and integrate AI/LLM-enabled features into platform/developer tooling (positive to have),
including RAG patterns for knowledge retrieval (runbooks, standards, operational procedures) while meeting governance requirements.
Must-have skills & experience
- Strong platform/API engineering background in an enterprise environment.
- OpenShift experience (engineering and/or operations) including real-world Day-2 cluster lifecycle responsibilities.
- Proven experience with enterprise infrastructure operations and governance.
- Hands-on development of RESTful APIs using Python (FastAPI).
- Strong hands-on experience with Docker + Kubernetes/OpenShift.
- Demonstrable experience in automation/operations, including:
- OS upgrade/patching workflows (planning, validation, automation, evidence)
- File system expansion / storage lifecycle operations automation (safe execution and verification)
Good-to-have
- Ansible (playbooks/roles) for automating OS and platform tasks; integration into CI/CD pipelines.
- Practical experience with AI/LLMs including RAG and integration into platform/tooling.
- Familiarity with observability stacks (metrics, logging, tracing) and alerting best practices.
Success measures
- Repeatable, low-risk OpenShift upgrade and patching processes with strong automation and clear audit trails.
- Reduced manual effort through reliable OS upgrade and filesystem expansion automation.
- APIs that are secure, well-documented, and adopted by engineering teams to self-serve Day-2 operations safely.
- Improved stability and reduced incident recurrence through robust monitoring, runbooks, and RCA-driven improvements.
📌 Platform / API Engineer (Bengaluru)
🏢 Infinityquest It Services
📍 Bengaluru