This is a high-growth role designed for an early-career engineer who is passionate about infrastructure, automation, and problem-solving. You won’t just be fixing things when they break; you’ll be helping us build resilient systems so they don’t break in the first place.
What You'll Do (Responsibilities)
Monitor & Alert: Keep a watchful eye on system health using tools like [Prometheus / Datadog / Grafana]. You will help tune alerts to reduce fatigue and ensure we only wake up for real emergencies.
Incident Response: Serve as an early responder for system issues. You'll help troubleshoot, mitigate downtime, and participate in blameless post-mortems to ensure we learn from every outage.
Automate Everything: Write scripts in [Python / Bash / Go] to automate manual, repetitive toil so the team can focus on higher-value engineering work.
Maintain Infrastructure: Assist in provisioning and managing cloud infrastructure on [AWS / GCP / Azure] using Infrastructure as Code (IaC) tools like [Terraform / Ansible].
Support CI/CD: Help maintain deployment pipelines (e.g., [GitHub Actions / GitLab CI / Jenkins]) to ensure developers can ship code quickly and safely.
Documentation:
Write and update runbooks, system documentation, and architecture diagrams so the whole team can share knowledge effectively.
What We’re Looking For (Requirements)
Education/Experience: B.S. in Computer Science, IT, or a related field (or equivalent hands-on experience/bootcamp completion).
OS Fundamentals: Solid understanding of Linux/Unix environments and command-line utilities.
Scripting Skills: Proficiency in at least one programming or scripting language (Python, Bash, Go, or Ruby).
Networking Basics: A foundational understanding of networking concepts (TCP/IP, DNS, HTTP/S, Load Balancing).
Curiosity & Communication: A solid desire to learn complex systems, ask positive questions, and collaborate across engineering teams.
Nice to Haves (Not Required, But a Plus)
Cloud Exposure: Familiarity with deploying applications to AWS, GCP, or Azure.
Containers: Basic understanding of Docker and container orchestration (Kubernetes, ECS).
Version Control: Experience using Git and team-oriented workflows (Pull Requests, code reviews).
📌 Junior 2 5 Yrs Site Reliability Engineer Chennai Onsite
🏢 OSpectra AI
📍 Chennai
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.