29 Aug
|
Pitchers Internet
|
Bengaluru
29 Aug
Pitchers Internet
Bengaluru
Responsibilities:
- Own the reliability, scalability, and performance of production systems and services.
- Design and implement highly available, fault-tolerant, and distributed infrastructure.
- Define and drive observability strategy, including monitoring, logging, and alerting.
- Build and maintain scalable CI/CD pipelines to enable fast and reliable deployments.
- Automate infrastructure provisioning and operational workflows using IaC tools.
- Lead incident management, root cause analysis (RCA), and implement preventive measures.
- Define and track SLIs, SLOs, and SLAs aligned with business and product requirements.
- Collaborate closely with engineering teams to improve system design, deployment processes, and operational excellence.
- Optimize cloud infrastructure for cost, performance, and efficiency.
- Own and improve on-call processes; mentor engineers in handling production incidents.
- Drive best practices for security, networking, and infrastructure reliability.
- Participate in architecture and design reviews to ensure system resilience and scalability.
- Document system architecture, runbooks, and operational processes.
Requirements:
- 8+ years of experience in DevOps, SRE, or Cloud Engineering roles.
- Proven experience building software using Python, Go, Rust, or JavaScript, with strong scripting capabilities.
- Hands-on experience with cloud platforms such as AWS, GCP, or Azure.
- Expertise in Infrastructure as Code tools like Terraform, Ansible, or CloudFormation.
- Strong experience with containerization (Docker) and orchestration (Kubernetes).
- Solid understanding of Linux systems, networking concepts, and security best practices (IAM, VPNs,
firewalls).
- Experience with monitoring and observability tools like Prometheus, Grafana, ELK Stack, Recent Relic, or Datadog.
- Proven experience in building and maintaining CI/CD pipelines (Circle CI, ArgoCD, Jenkins, GitHub Actions, GitLab CI, etc.).
- Deep hands-on experience owning cloud security end-to-end.
- Strong debugging and troubleshooting skills for complex production systems.
- Comfortable operating in fast-moving engineering environments, balancing long-term infrastructure investment with immediate operational needs.
- Experience owning and continuously improving on-call processes - including rotation design, escalation policies, runbook culture, and post-incident review cadence.
Nice to Have:
- Experience with large-scale distributed systems and microservices architecture.
- Experience building and running operators on Kubernetes / knowledge of internals.
- Hands-on experience with cost optimization and capacity planning in cloud environments.
- Experience mentoring engineers and driving engineering best practices.
- Prior involvement in architectural reviews and cross-team technical decision-making.
- Strong understanding of compliance, security standards, and DevSecOps practices.
- Experience building internal developer platforms or self-service infrastructure tools.
- Development exposure experience contributing to backend services or application code (e.g., APIs, microservices), with strong software engineering fundamentals.
- MLOps experience – familiarity with deploying, monitoring, and managing ML models in production, including tools like ML pipelines, model versioning, and data workflows.
📌 Devops SDE 3 (Bengaluru)
🏢 Pitchers Internet
📍 Bengaluru