10 Sep
|
Infinityquest It Services
|
Hyderabad
10 Sep
Infinityquest It Services
Hyderabad
Job Description: Infrastructure Automation Engineer
Role purpose
Deliver enterprise-grade infrastructure automation that reduces operational toil and improves consistency, security, and speed of change. Youll engineer reusable automation (primarily Ansible) to support platform and infrastructure Day-2 operations, with a strong focus on OpenShift-aligned environments and enterprise controls.
Key responsibilities
- Design, build, test, and maintain infrastructure automation using Ansible (playbooks, roles, inventories, reusable modules) to enable repeatable, auditable operations at scale.
- Own automation for Day-2 operational use cases, such as patching/maintenance workflows, configuration drift management, standard health checks, compliance validation, and remediation orchestration (aligned to the scope of the domain).
- Support automation for OpenShift platform operations where required (e.g., integration with cluster processes, configuration updates, operator lifecycle support, platform hygiene tasks).
- Build and maintain automation guardrails: input validation, idempotency, error handling, rollback steps, logging, and evidence capture to meet audit and operational requirements.
- Integrate automation with enterprise infrastructure ecosystems (identity/access, secrets/certificates, networking, monitoring/observability, CMDB/ITSM workflows), ensuring alignment to governance and change controls.
- Partner with SRE/operations, security, and engineering teams to translate domain requirements into automation features that are easy to consume and safely run.
- Troubleshoot automation failures end-to-end (Ansible, OS, network, platform dependencies) and drive root-cause analysis and continuous improvement.
- Contribute to CI/CD integration for automation (testing, linting, packaging, release/versioning) and maintain clear documentation/runbooks for consumers.
- Where applicable, contribute to developer enablement via self-service patterns (e.g., pipelines, catalog items, standard templates).
- Good to have develop lightweight RESTful APIs in Python/FastAPI to expose automation/infrastructure capabilities as services, and/or package supporting components in containers for scalable execution.
- Good to have explore AI/LLM-enabled enhancements to automation/developer tooling (including practical RAG patterns) to improve discoverability of runbooks, standards, and automation usage.
Must-have skills & experience
- Strong hands-on Ansible experience in enterprise environments (roles/playbooks, inventories, best practices for idempotency and reusability).
- OpenShift knowledge/experience (working knowledge of OpenShift concepts and how infrastructure automation supports platform operations).
- Demonstrable enterprise infrastructure experience (controls, change management,
dependency management across shared services).
- Domain & feature knowledge relevant to the infrastructure area in scope (as defined by the hiring team).
- Solid operational mindset: automation built for safety, auditability, and production support.
Good-to-have
- Building RESTful APIs using Python/FastAPI to expose infrastructure capabilities.
- Hands-on Docker + Kubernetes/OpenShift experience (packaging/running automation services or supporting tooling).
- Practical knowledge of AI/LLMs, including RAG and integration into developer tooling/platforms.
Success measures
- Measurable reduction in manual operational effort (e.g., repeatable automation adopted for high-volume Day-2 tasks) and improved time-to-deliver standard infrastructure changes.
- Automation is reliable and safe: idempotent runs, clear logging, predictable outcomes, and defined rollback/failure handling; reduced repeat incidents linked to manual change.
- Strong auditability and compliance: evidence capture is automated, approvals/controls are integrated where needed, and automation aligns to enterprise change processes.
- Improved platform/infrastructure stability through consistent configuration management and reduced drift.
- High usability and adoption: clear documentation/runbooks, self-service enablement, and positive feedback from operations and engineering teams.
- Reduced mean time to recover (MTTR) for automation-supported workflows through better diagnostics, monitoring, and rapid remediation.
📌 Infrastructure Automation Engineer (Hyderabad)
🏢 Infinityquest It Services
📍 Hyderabad