11 Sep
|
Netrolynx AI
|
India
11 Sep
Netrolynx AI
India
About The Company Cogrion is at the forefront of building an Autonomous Data & AI Infrastructure Platform tailored for modern enterprises. Our mission is to empower organizations to harness the full potential of their data and artificial intelligence capabilities through innovative, scalable, and secure cloud-native solutions.
We specialize in integrating cutting-edge technologies such as Kubernetes, data platforms, and AI workloads to deliver robust infrastructure that drives digital transformation. Our team is composed of forward-thinking professionals dedicated to creating platforms that are reliable, efficient, and adaptable to the evolving needs of our clients. At Cogrion, we foster a culture of innovation, collaboration, and continuous learning, ensuring that our solutions stay ahead of industry trends and meet the highest standards of security and performance.
About The Role We are seeking a highly skilled Senior DevOps Engineer to join our dynamic team. In this role, you will be instrumental in designing, automating, securing, and operating large-scale cloud-native platforms across multiple cloud providers including AWS, Azure, GCP, and Alibaba Cloud.
Your expertise will be vital in managing Kubernetes environments, data infrastructure, and AI workloads to ensure high availability, security, and optimal performance. You will work closely with cross-functional teams to develop scalable solutions, implement automation strategies, and enhance platform reliability. The ideal candidate will possess a deep understanding of cloud infrastructure, container orchestration, and data platform technologies, with a passion for building resilient and efficient systems that support enterprise-grade applications and AI initiatives.
Qualifications
- 7–10 years of experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering roles.
- Strong expertise in Kubernetes and Docker containerization technologies.
- Hands-on experience with multiple cloud platforms, particularly AWS,
and at least one additional cloud provider such as Azure, GCP, or Alibaba Cloud.
- Proficiency in infrastructure as code tools like Terraform and OpenTofu, along with Helm for package management.
- Extensive experience designing and implementing CI/CD pipelines using tools like ArgoCD, GitHub Actions, GitLab CI, or similar.
- Solid understanding of Linux operating systems, networking concepts, DNS, ingress controllers, load balancing, and cloud networking.
- Experience with identity and access management (IAM), secrets management, workload identity, and network security protocols.
- Strong scripting skills in Python, Bash, or equivalent languages for automation and troubleshooting.
- Proven ability to troubleshoot complex issues related to Kubernetes, storage, networking, and distributed systems in production environments.
- Knowledge of scalability, reliability, disaster recovery, and infrastructure cost optimization strategies.
Responsibilities
- Design, deploy, and maintain production-grade Kubernetes platforms to support diverse workloads.
- Build and manage infrastructure components using Terraform/OpenTofu, Helm, and other automation tools.
- Implement GitOps practices and automate CI/CD pipelines to streamline application deployment and updates.
- Manage multi-cloud environments, ensuring seamless operation across AWS, Azure, GCP, and Alibaba Cloud.
- Develop autoscaling, capacity planning, and cost optimization strategies to maximize resource efficiency.
- Operate and optimize platforms running data and AI workloads such as Spark, Trino, Airflow, Jupyter, MLflow, and Kafka.
- Establish observability frameworks using Prometheus, Grafana,
Loki, OpenTelemetry, and alerting systems for proactive monitoring.
- Enhance platform security through IAM, workload identity, secrets management, network security, and RBAC implementations.
- Troubleshoot and resolve complex issues related to Kubernetes, networking, storage, and distributed systems in production.
- Automate deployment, upgrades, patching, backup, and disaster recovery processes to ensure platform stability and security.
Benefits At Cogrion, we offer a comprehensive benefits package designed to support our employees' professional growth and personal well-being. Our offerings include competitive salary packages, health insurance, and retirement plans to ensure financial security. We promote a flexible work environment with options for work from home and flexible hours to accommodate diverse lifestyles. Employees have access to continuous learning opportunities, including training programs, certifications, and industry conferences, to stay ahead in the rapidly evolving cloud and data platform landscape. We also foster a collaborative and inclusive culture that values innovation, diversity, and work-life balance. Additionally, Cogrion provides wellness programs, paid time off, and other perks to support our team members’ overall health and happiness.
Equal Opportunity
Cogrion is an equal opportunity employer committed to fostering an inclusive environment for all employees. We do not discriminate based on race, ethnicity, gender, sexual orientation, age, disability, or any other protected characteristic. We believe that diversity drives innovation and excellence, and we are dedicated to providing equal employment opportunities to all qualified candidates.
Our hiring practices are designed to ensure fairness and transparency, and we welcome applicants from all backgrounds to join our team and contribute to our mission of building cutting-edge cloud-native infrastructure solutions.
📌 DevOps Engineer (India)
🏢 Netrolynx AI
📍 India