24 Sep
|
Infosys
|
Bengaluru
Relevant Yrs. of experience
5+ years (overall experience 10 years)
Detailed JD (Roles and Responsibilities)
Qualifications:
- Experience: 5+ years as an SRE, DevOps Engineer, or Software Engineer with a focus on infrastructure and operations.
- Programming Skills: Proficient in at least one programming language, such as Python, Go, Java, or similar.
- IaC Expertise: Proven experience with infrastructure automation tools like Terraform, CloudFormation, or similar.
- Containerization: Experience with containerization and orchestration technologies such as Docker and Kubernetes.
- CI/CD Tools: Solid experience with CI/CD pipelines and tools (Jenkins, GitLab CI, CircleCI, etc.).
- Cloud Platforms: Deep understanding of cloud architecture like AWS, GCP, or Azure and their respective services.
- Monitoring & Logging: Solid experience with monitoring and logging tools such as Prometheus, Grafana, ELK stack, or Datadog.
- Networking & Security: Strong grasp of networking concepts and security practices.
- Problem-Solving: Excellent analytical and troubleshooting skills to diagnose and resolve complex system issues.
- Collaboration: Strong communication skills and the ability to work effectively within cross-functional teams
Responsibilities:
- Infrastructure as Code (IaC): Design, develop, and manage infrastructure automation using tools such as Terraform, Ansible, or similar to provision and maintain cloud resources efficiently.
- CI/CD Pipeline Development: Build and enhance continuous integration and deployment pipelines, automating testing, build, and release processes to ensure faster, more reliable deployments.
- Monitoring & Alerting Systems: Develop and improve monitoring,
logging, and alerting systems to track the health, performance, and security of applications and infrastructure, while implementing automated responses to resolve issues.
- Tooling & Automation Development: Use programming languages like Python, Go, or Java to create custom tools that automate operational tasks, streamline workflows, and improve system efficiency.
- Incident Management & Troubleshooting: Participate in an on-call rotation to investigate and resolve complex production issues, conduct root cause analysis, and implement long-term preventative solutions.
- Capacity & Performance Optimization: Analyze system performance, identify bottlenecks, implement solutions to optimize scalability and efficiency and forecast capacity needs and ensure systems can handle growth without degradation.
- Collaboration & Mentorship: Work closely with software engineering teams to drive system reliability improvements, influence design decisions, and provide mentorship to junior SREs.
- Security Best Practices: Implement and maintain security standards across infrastructure and applications, addressing vulnerability management and supporting incident response.
- Documentation: Develop and maintain explicit and comprehensive documentation for systems, procedures, and operational practices.
- Research & Technology Evaluation: Stay current with industry trends and evaluate new technologies to improve infrastructure, enhance operational efficiency, and ensure cutting-edge solutions.
Mandatory skills
AWS SRE, DevOps, Python/Java/Go, Terraform/CloudFormation, Docker/Kubernetes, Jenkins/GitLab CI/CircleCI, Prometheus/Grafana/ELK stack/ Datadog
📌 AWS SRE (Bengaluru)
🏢 Infosys
📍 Bengaluru