06 Aug
|
DeepIQ
|
Hyderabad
About the Role
DeepIQ is looking for a hands-on Site Reliability Engineer (SRE) to help improve the reliability, security, scalability, and operational efficiency of our cloud platform.
Our applications are primarily hosted on AWS and deployed on Amazon EKS. We use Argo CD for GitOps deployments, Terraform for infrastructure as code, GitHub for source control and CI/CD workflows, and Linear for project and incident tracking.
In this role, you will work closely with engineering teams to operate our production environments, improve observability, automate infrastructure and deployments, respond to incidents, and strengthen the security of our cloud platform.
ResponsibilitiesCloud Infrastructure and Kubernetes
· Operate, maintain, and improve production and non-production environments running in AWS.
· Manage applications and supporting services deployed on Amazon EKS.
· Troubleshoot Kubernetes workloads, networking, storage, resource utilization, and application deployment issues.
· Manage Kubernetes resources such as deployments, services, ingress controllers, ConfigMaps, secrets, autoscaling policies, and Helm charts.
· Improve the scalability, availability, and cost efficiency of the platform.
· Support AWS services such as EC2, IAM, VPC, S3, RDS, Route 53, CloudFront, load balancers, and CloudWatch.
Infrastructure as Code and GitOps
· Build and maintain reusable Terraform modules for AWS and Kubernetes infrastructure.
· Manage application deployments using Argo CD and GitOps practices.
· Maintain clear separation between development, staging, and production environments.
· Review infrastructure changes through pull requests and automated validation.
· Identify manual operational processes and replace them with reliable automation.
· Help maintain GitHub-based CI/CD workflows for application builds, testing, security checks, and deployments.
Monitoring, Logging, and Observability
· Build and maintain monitoring, alerting, logging, and dashboarding across applications and infrastructure.
· Monitor platform health, application performance, Kubernetes workloads, and AWS services.
· Define actionable alerts that minimize noise and identify real production issues.
· Centralize and improve application, infrastructure, audit, and security logs.
· Help implement and maintain tools such as Amazon CloudWatch, Prometheus, Grafana, OpenTelemetry, OpenSearch, or similar observability platforms.
· Work with developers to improve application metrics, structured logging, distributed tracing, and health checks.
· Define and track service-level indicators, service-level objectives, availability, latency, and error rates.
Reliability and Incident Management
· Participate in production support and an on-call rotation.
· Investigate production incidents and restore services quickly and safely.
· Perform root-cause analysis and document incident findings, corrective actions, and preventive measures.
· Create and maintain operational runbooks and troubleshooting documentation.
· Improve backup, disaster recovery, high availability, and business continuity procedures.
· Conduct capacity planning, resilience testing, and failure scenario reviews.
· Track operational improvements, incidents, and follow-up work in Linear.
Security and Compliance
· Apply security best practices across AWS, Kubernetes, CI/CD pipelines, and application deployments.
· Review and improve AWS IAM roles, policies, service accounts, and least-privilege access.
· Help manage secrets securely using AWS Secrets Manager, Kubernetes secrets, or similar solutions.
· Implement container image scanning, dependency scanning, infrastructure scanning, and vulnerability management.
· Support Kubernetes security practices such as RBAC, network policies, workload identity, pod security controls, and secure container configurations.
· Monitor security events and assist with incident investigation and remediation.
· Support patching, certificate management, access reviews, audit logging, and security compliance activities.
Collaboration and Continuous Improvement
· Work closely with software engineers to improve application reliability and production readiness.
· Participate in architecture, infrastructure, and deployment design reviews.
· Help engineering teams understand operational risks and reliability requirements.
· Promote automation, observability, security, and infrastructure-as-code best practices.
· Document platform architecture, operational procedures, and technical decisions.
· Identify opportunities to reduce cloud costs without negatively affecting reliability or performance.
Required Qualifications
· 2–3 years of experience in site reliability engineering, DevOps, cloud infrastructure, platform engineering, or a related role.
· Hands-on experience working with AWS.
· Practical experience operating Kubernetes environments, preferably Amazon EKS.
· Experience troubleshooting Kubernetes applications, networking, resource constraints, and deployment failures.
· Experience with Terraform or another infrastructure-as-code tool.
· Experience with Git-based development and deployment workflows.
· Familiarity with GitOps and continuous delivery tools such as Argo CD.
· Experience implementing or supporting monitoring, logging, dashboards, and alerts.
· Working knowledge of Linux administration, networking, DNS, TLS, HTTP, and load balancing.
· Ability to write automation scripts using Python, Bash, or a similar language.
· Understanding of cloud security, IAM, secrets management, and least-privilege access.
· Robust troubleshooting, documentation, and communication skills.
Preferred Qualifications
· Experience with Prometheus, Grafana, Amazon CloudWatch, OpenTelemetry, OpenSearch, or the Elastic Stack.
· Experience managing Helm charts and Kubernetes ingress controllers.
· Experience with GitHub Actions or another CI/CD platform.
· Familiarity with service meshes, Kubernetes autoscaling, and cluster autoscaling.
· Experience with container and infrastructure security tools such as Trivy, Checkov, tfsec, Snyk, or similar platforms.
· Familiarity with AWS security services such as GuardDuty, Security Hub, Inspector, CloudTrail, AWS Config, and IAM Access Analyzer.
· Experience supporting multi-tenant SaaS applications.
· Familiarity with SOC 2, ISO 27001, or other security and compliance frameworks.
· Exposure to FinOps, AWS cost monitoring, and cloud resource optimization.
· AWS or Kubernetes certifications are helpful but not required.
What Success Looks Like
Within your first several months, you will:
· Develop a strong understanding of DeepIQ’s AWS, EKS, and application architecture.
· Improve visibility into application and infrastructure health.
· Reduce noisy alerts and improve incident detection.
· Automate recurring infrastructure and operational tasks.
· Strengthen cloud and Kubernetes security controls.
· Improve deployment reliability through Terraform, GitHub, and Argo CD.
· Create useful operational documentation and incident-response runbooks.
· Help reduce production incidents, recovery time, and unnecessary cloud costs.
What We Are Looking For
We are looking for someone who is curious, dependable, and comfortable taking ownership of operational problems. You should enjoy troubleshooting complex systems, automating repetitive work, and collaborating with developers to make applications easier and safer to operate.
You do not need to know every tool listed in this description. However, you should have a strong foundation in AWS, Kubernetes, infrastructure automation, and production troubleshooting, along with the willingness to learn and take ownership.
📌 Site Reliability Engineer (Hyderabad)
🏢 DeepIQ
📍 Hyderabad