System Health Monitoring and Telemetry (Bangalore Metropolitan Area)

System Health Monitoring and Telemetry (Bangalore Metropolitan Area)

04 Oct
|
Tsavorite Scalable Intelligence
|
Bangalore Metropolitan Area

04 Oct

Tsavorite Scalable Intelligence

Bangalore Metropolitan Area

About Company

Tsavorite Scalable Intelligence is building a new foundation for AI computing. We design silicon, systems, and software that work as one, bringing together general-purpose processing, AI acceleration, and scalable memory in a single architecture. From edge devices to data centers, our goal is to make advanced AI more powerful, efficient, and accessible.

Our founding team has a track record of shipping the highest performance silicon and we have engineering teams in the US and India. We are looking for curious, driven people across disciplines who want to turn ambitious engineering into real products and help shape how the world computes.

Job Summary

Architect and lead hardware health monitoring, telemetry, and RAS infrastructure for AI accelerator SoCs, boards, and datacenter platforms. The role requires robust C/C++,Python, Linux kernel interfaces, hardware telemetry, PCIe/CXL, and observability expertise, with ownership across firmware, drivers, and user-space monitoring services.

Technical Skills & Expertise

- Experience in platform/system software, embedded software, or firmware development, with demonstrated ownership of complex subsystems.
- Expert-level C/C++ skills; strong proficiency with Python for tooling/pipeline development.
- Deep experience with Linux kernel interfaces relevant to hardware monitoring (sysfs, procfs, netlink, hwmon subsystem, IOCTLs), including having designed or extended such interfaces.
- Strong understanding of hardware telemetry sources: sensors (temperature/voltage/current), PMIC/PMBus interfaces, I2C/SPI/PCIe register access, error/status registers.
- Proven experience architecting data collection/aggregation pipelines and time-series storage systems (e.g., Prometheus, InfluxDB) at scale.
- Solid grounding in RAS concepts: error correction/detection, fault isolation, predictive maintenance,



and experience defining RAS requirements/policy, not just implementing them.
- Experience with AI accelerator, GPU, or large-scale datacenter server platforms specifically.
- Deep familiarity with PCIe/CXL link status and error reporting (AER, CXL RAS features).
- Experience with BMC (Baseboard Management Controller) firmware, IPMI/Redfish interfaces, or equivalent out-of-band management systems.
- Exposure to ML-based anomaly detection on telemetry data, and judgment on when it is (and isn't) the right tool.

What You'll Drive & Bring to the Role

- Architect end-to-end health monitoring and telemetry systems spanning firmware, drivers, kernel, and user-space services, defining the technical direction for the domain.
- Lead design and implementation of health monitoring agents/daemons polling or subscribing to hardware telemetry (temperature, voltage rails,power, clock/PLL lock status, error counters, link/PHYstatus, ECC/parity errors,thermal throttling events).
- Define statistics collection pipeline architecture: counters, histograms, event logs, periodic snapshots, and their integration into fleet-scale observability systems.
- Drive interface design between low-level firmware/driver telemetry sources and higher-level monitoring/orchestration software(sysfs, IOCTL, netlink,gRPC, vendor-specific APIs), setting standards other engineers follow.
- Own fault detection and classification strategy (thresholding, anomaly detection, cross-subsystem correlation),



distinguishing transient vs. persistent vs. fatal errors, and defining escalation/mitigation policies.
- Define health/statistics data schemas at a platform level and lead integration with datacenter monitoring stacks (e.g., Prometheus, Grafana, or internal telemetry pipelines) for fleet-wide visibility.
- Partner directly with firmware, driver, hardware validation, and reliability engineering leads to define monitoring requirements, signal granularity, and RAS (Reliability, Availability, Serviceability) strategy.
- Lead RAS feature development error logging architecture, predictive failure indicators, graceful degradation and failover triggers.
- Drive bring-up and production test tooling strategy, ensuring health/statistics visibility from earliest silicon bring-up through fleet deployment.
- Lead post-silicon debug efforts, correlating telemetry data with system failures or performance anomalies, and drive root-cause investigations across teams.
- Represent the health monitoring domain in cross-functional design reviews and influence hardware/firmware design decisions upstream to improve observability.

Required Qualification

- BS/MS in Computer Engineering, Electrical Engineering, Computer Science, or related field (MS preferred).
- 10 to 14 yrs of work experience.

Why Join Us?

- Own. Innovate. Make an Impact. Join a talented team building next-generation compute solutions and working on cutting-edge technology. Take ownership, collaborate with great engineers, grow your skills, and make an impact in a culture that values innovation and teamwork.

Benefits

- Competitive Compensation aligned with industry benchmarks.
- Comprehensive Health Insurance covering employees and eligible dependents.
- Stock Options & Ownership Opportunities.

📌 System Health Monitoring and Telemetry (Bangalore Metropolitan Area)
🏢 Tsavorite Scalable Intelligence
📍 Bangalore Metropolitan Area

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: system health monitoring and telemetry (bangalore metropolitan area) / bangalore metropolitan area

Subscribe to this job alert:

Get the latest job offers by email for: system health monitoring and telemetry (bangalore metropolitan area) / bangalore metropolitan area