02 Sep
|
evoluteIQ
|
India
Job Type:
Permanent
Experience:
9+ Years
Location:
Bangalore
Life at EvoluteIQ:
We at EvoluteIQ believe in the power of transformation. We are committed to building a industry-leading technology that will revolutionize the way enterprises conduct business.
To make that happen, we need people who are generous, genuine, self-driven and collaborative. People who not only want to be a part of a fast-growing and radical thinking company, but who are kind and care—about each other. We at EvoluteIQ thrive in the company of each other and make each other a better version of ourselves every day.
Could that be you?
The EIQ Mission:
Our culture is who we are, and it defines us. We are nimble, fast, dedicated, humble and bold. We see opportunities everywhere around us to make a “dent in the universe”. Our core fabric is Innovation! And this is not just a word, but the way of life within EvoluteIQ.
Changing the way business is done, reimagining the future of work, and enabling people to improve their quality of work every day, is a daunting task. That is what we all strive every day at EvoluteIQ. To enable everyone to automate complicated processes with their instinct and knowledge.
We make building of automation easy for everyone, so that it is not mystery for a few.
Would you like to be part of this journey?
What you'll do at EvoluteIQ:
As an SRE you will be responsible for ensuring the reliability, availability, scalability, performance, and operational excellence of EvoluteIQ’s platform across AWS, Azure, and Google Cloud.
You will work closely with Engineering, Product, and Customer-facing teams to troubleshoot complex production issues, improve platform resilience, automate operational processes, and drive continuous improvements in the reliability of the platform.
The role requires strong hands-on experience with Kubernetes, Docker, Linux, cloud platforms, Java/Spring Boot applications, databases, networking, APIs, and enterprise integrations.
What will you bring to the team?
Platform Reliability & Production Operations
- Own and drive in incident management, troubleshooting, root-cause analysis, and post-incident reviews.
- Monitor production environments and proactively identify potential reliability and performance issues.
- Establish and maintain operational runbooks and troubleshooting procedures.
- Define and implement reliability improvements, operational best practices, and self-healing mechanisms.
Application & Database Troubleshooting
- Debug and troubleshoot Java/Spring Boot applications running on Docker and Kubernetes.
- Analyze application logs, stack traces, resource utilization, and performance issues.
- Troubleshoot database connectivity, queries, connection pools, and performance issues.
- Troubleshoot JDBC/ODBC connectivity, REST APIs, enterprise connectors, and integration workflows.
- Work with Engineering teams to identify application-level reliability and performance improvements.
Infra & Automation
- Contribute to self-healing and automated remediation capabilities by using AI tools, Python, Bash and cloud native tools for quick proof of concepts.
- Manage and troubleshoot production workloads across AWS, Microsoft Azure, Google Cloud Platform (GCP) and on-prem.
- Build and enhance monitoring, alerting, logging, tracing, and observability capabilities.
Required Skills & Experience
- 9+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, Cloud Engineering, or a similar role.
- Experience troubleshooting Java/Spring Boot applications and good understanding of SQL and databases.
- Experience with REST APIs, enterprise integrations, JDBC/ODBC, and application connectivity.
- Strong hands-on experience with Kubernetes and Docker.
- Strong experience with at least one major cloud platform and preferably exposure to AWS, Azure, and GCP.
- Strong Linux administration and troubleshooting skills.
- Strong scripting/programming skills using AI tools, Python, Bash and cloud native tools.
- Strong knowledge of networking concepts, DNS, HTTP/HTTPS, TCP/IP, load balancing, and connectivity troubleshooting.
- Strong incident management, problem management, root-cause analysis, and production troubleshooting skills.
- Ability to work effectively with Engineering, Product, and Customer-facing teams.
- Strong analytical and problem-solving skills with a hands-on approach.
Good To Have:
- Experience with AI/ML platforms and production AI workloads.
- Experience working with enterprise automation, RPA, or workflow platforms.
What We Offer:
- Opportunity to shape the strategy of a next-gen hyper-automation platform.
- Work with a cross-disciplinary team in a fast-growing, innovation-driven environment.
- Competitive compensation and growth opportunities.
- A culture of innovation, ownership, and continuous learning.
Ready to take the next step in your career journey with us? Apply now and become part of a dynamic team that values expertise, collaboration, and a positive work culture!
Don’t Meet Every Requirement? We Still Want to Hear From You!
Please go ahead and apply anyway.
We know that experience comes in all shapes and sizes and passion can’t be learned. If you believe, then we would like to meet you and explore that belief!
We value a range of diverse backgrounds, experiences, and ideas. We pride ourselves on our diversity and inclusive workplace that provides equal opportunities to all persons regardless of age, race, colour, religion, sex, sexual orientation, gender identity and expression, national origin, disability, neurodiversity, military and/or veteran status, or any other protected classes.
📌 Site Reliability Engineer (India)
🏢 evoluteIQ
📍 India