03 Aug
|
Match Point Solutions
|
India
03 Aug
Match Point Solutions
India
Match Point Solutions is a fast-growing, young, energetic global IT-Engineering services company with clients across the US. We provide technology solutions to various clients like Uber, Robinhood, Netflix, Airbnb, Google, Sephora, and more! More recently, we have expanded to working internationally in Canada, China, Ireland, UK, Brazil, and India. Through our culture of innovation, we inspire, build, and deliver business results, from idea to outcome. We keep our clients on the cutting edge of the latest technologies and provide solutions by using industry-specific best practices and expertise.
We are excited to be continuously expanding our team. If you are interested in this position, please send over your updated resume. We look forward to hearing from you!
Job Title: Production Operations Site Reliability Engineer (Linux Systems)
Location: Hyderabad/Bengaluru/Pune, India
Employment Type: Full-time
This is a Linux-first SRE role. The systems you protect run on Linux, the automation you write runs on Linux, and the incidents you resolve are debugged at the operating-system level. We need an engineer who is genuinely fluent at the command line and reasons confidently about the system beneath the tooling - processes, storage, networking, permissions, the kernel - not only through managed abstractions or orchestration platforms.
When a system is slow, a service won't start, a mount fills up, or a port is unreachable, we expect you to drop to the shell and work the problem directly. You'll partner across Engineering, Operations, Security, Platform, and Release teams to keep production observable, resilient, and secure to deploy - owning reliability, incident response, and the automation that reduces both toil and risk.
Much of our data-center footprint runs on VMware vSphere, managed with Ruby and Puppet, so that's the environment you'll work in day to day. Strong vSphere experience is highly important here. We don't require prior Ruby or Puppet experience - strong Linux fundamentals and transferable scripting and configuration-management skills matter more - but if you've lived in a vSphere/Ruby/Puppet shop, you'll feel at home fast.
We operate a hybrid environment today and are working toward a full migration to AWS over time. This person will help run production reliably across both worlds now, and may play a meaningful part in shaping and executing that migration - so strong AWS experience is close to essential, and a track record of moving workloads to the cloud is a real plus.
What You'll Do
Keep production reliable. Diagnose and resolve complex problems directly at the OS level - performance, resource contention, service failures, storage,
and networking - including cases where surface metrics look fine but the system is unhealthy. Own SLIs, SLOs, and the reliability improvements that reduce incidents and shorten recovery.
Make production observable. Build dashboards, monitors, logs, metrics, traces, and synthetic checks. Improve early detection and cut alert noise so alerts are actionable, timely, and tied to customer impact.
Automate the toil away. Write robust production tooling in Bash, plus Ruby and/or Python - pipelines, stream handling, text processing, safe error handling. Build automated remediation with proper safeguards, approvals, and rollback.
Deploy safely. Assess production readiness before changes ship - observability coverage, rollback plans, dependency and capacity impact. Advocate for progressive rollout, validation gates, and fast rollback.
Respond and prevent. Drive hands-on Linux diagnosis during incidents; coordinate, communicate, and mitigate under pressure. Run blameless post-incident reviews and close the loop on fixes that prevent recurrence. Support an on-call rotation.
What You Bring
Core Linux - the primary requirement:
Strong, demonstrable Linux systems administration across enterprise distributions (RHEL/CentOS, Ubuntu, Amazon Linux), earned through hands-on production operation - not managed-platform use alone.
Command-line fluency you can demonstrate: navigate, diagnose, and operate a Linux system confidently from the shell without reference material.
Working knowledge of Linux internals: process and service management, init/systemd, file systems and storage (including LVM), permissions, users and authentication, resource limits, logging, and performance troubleshooting.
Practical host-level networking: connectivity, name resolution, routing, ports and sockets, and firewalling from the operating system itself.
Strong Bash scripting - I/O redirection and streams, pipelines, text processing, control flow, and error handling - enough to write and debug production automation independently.
Production operations - also required:
5+ years in SRE, Production Operations, Systems/Platform Engineering, Dev Ops, or Cloud Operations, with Linux administration central and continuous throughout.
Strong grasp of incident management, monitoring, alerting, troubleshooting, and root cause analysis,
plus at least one scripting language beyond Bash - Ruby or Python - for real automation work.
Hands-on with CI/CD (Jenkins or similar), Git, Infrastructure-as-Code (Terraform or similar), configuration management (we run Puppet; Ansible or Chef experience transfers well), and containers (Docker).
Experience supporting production in a hybrid environment, with hands-on AWS experience strongly expected - given our migration path, AWS depth is close to essential for this role. Azure experience is a welcome bonus but not required. Includes a working understanding of deployment readiness, change management, and rollback.
Strong hands-on VMware vSphere experience - highly important. Much of our production runs on vSphere, so you should be comfortable operating, troubleshooting, and managing virtualized infrastructure at scale (ESXi hosts, vCenter, VM lifecycle, resource and capacity management).
Clear documentation and strong communication, especially under high-pressure incidents.
Nice to Have
SRE practice at depth: error budgets, toil reduction, reliability reviews, and production readiness reviews.
Deeper Linux: performance tuning, kernel-level troubleshooting, or systems-programming familiarity.
Centralized identity for Linux fleets: FreeIPA, IDM, LDAP, or Active Directory integration.
Data-center-to-AWS migration: hands-on experience moving production workloads from data center to AWS, and hybrid-cloud operations at scale.
Broader virtualization and storage: hypervisor technologies beyond vSphere, and enterprise SAN / shared-storage experience (fabric, LUNs, multipathing, capacity).
Network security: Palo Alto firewalls or comparable enterprise firewall/security platforms.
Kubernetes / EKS, Git Ops practices, and common observability and platform tooling (e.g., Datadog, Splunk, Grafana/Prometheus, Cloud Watch, Pager Duty, Sonar Qube, Snyk).
Match Point Solutions is a fast-growing, young, energetic global IT-Engineering services company with clients across the US. We provide technology solutions to various clients like Uber, Robinhood, Netflix, Airbnb, Google, Sephora, and more! More recently, we have expanded to working internationally in Canada, China, Ireland, UK, Brazil, and India. Through our culture of innovation, we inspire, build, and deliver business results, from idea to outcome. We keep our clients on the cutting edge of the latest technologies and provide solutions by using industry-specific best practices and expertise.
We are excited to be continuously expanding our team. If you are interested in this position, please send over your updated resume. We look forward to hearing from you!
📌 Production Operations Site Reliability Engineer (Linux Systems) (India)
🏢 Match Point Solutions
📍 India