02 Sep
|
Quantiphi Analytics Solutions
|
Thiruvananthapuram
02 Sep
Quantiphi Analytics Solutions
Thiruvananthapuram
Strategic Engagement & Leadership:
- Client & Stakeholder Management: Serve as the primary technical and functional point of contact for key clients and internal business units regarding platform reliability, performance, and service delivery. Translate complex technical issues and resolutions into clear, business-centric communications.
- Support Strategy & Roadmap: Define, champion, and execute the long-term support strategy, incorporating SRE principles (SLOs, SLIs, Error Budgets) to proactively enhance platform reliability and user experience. Align this strategy with overall product and business goals.
- Team Leadership & Development: Lead, mentor, and empower a team of Senior and Junior Support Engineers. Foster a culture of continuous improvement, technical excellence, and client-centricity. Conduct performance reviews, manage shift rosters, and facilitate knowledge sharing.
- Major Incident Management & Communication: Act as the Major Incident Manager and the primary escalation POC. Lead technical resolution war rooms, drive Root Cause Analysis (RCA), and own the Corrective Action processes. Crucially, manage all stakeholder communications during critical incidents, providing timely and transparent updates.
Technical Architecture & Automation:
- SRE Implementation & Governance: Drive the adoption and continuous improvement of SRE practices, defining and tracking key metrics (SLOs, SLIs, Error Budgets) to ensure platform reliability and performance meet business expectations.
- Infrastructure as Code (IaC) & Automation: Architect and enforce IaC best practices using Terraform/CloudFormation, ensuring environments are self-healing, reproducible, and compliant. Oversee the design and implementation of comprehensive monitoring, logging, and alerting stacks to enable proactive support.
- Release & Deployment Management: Oversee the maturity of the release pipeline, ensuring zero-downtime deployments, automated rollbacks, and robust change management processes.
- Cloud-Native Innovation: Introduce and integrate advanced cloud-native tools and services to improve operational efficiency, enhance automation, and drive cost optimization.
Collaboration & Influence:
- Design for Supportability: Collaborate closely with Solution Architects, Product Managers, and Engineering teams to influence the design of new features and platforms, ensuring they are inherently supportable, maintainable, and scalable from inception.
- Cross-Functional Alignment: Bridge the gap between Data Engineering, ML Engineering, and Operations,
ensuring seamless integration and productive resolution of cross-functional issues.
- Process Optimization: Continuously evaluate and optimize support processes, leveraging automation and best practices to improve efficiency and reduce Mean Time To Resolution (MTTR).
Required Technical Skills: Expert Proficiency in Cloud Infrastructure (AWS/GCP preferred):
- Compute & Serverless: Deep expertise in configuring and optimizing EC2/GCE instances, Lambda/Cloud Functions, and auto-scaling groups based on custom metrics.
- Networking: Advanced knowledge of VPC peering, Transit Gateways, Load Balancers (ALB/NLB), Route53/Cloud DNS, and VPN/Direct Connect setups.
- Security & IAM: Ability to design least-privilege IAM policies, manage KMS keys for encryption, and implement WAF rules to protect public-facing endpoints.
- Storage Lifecycle: Managing tiered storage strategies (S3/GCS lifecycle policies) and high-performance block storage optimization (EBS/Persistent Disks).
Software Engineering Proficiency:
- Core Development: Strong command of Python (using frameworks like Flask/Django/FastAPI) or Java (Spring Boot) to build production-grade applications and tooling.
- White-Box Debugging: Ability to clone application repositories, navigate complex codebases, attach remote debuggers, and identify the exact line of code causing logic errors or memory leaks.
- Code Contribution & Review: Comfortable submitting Pull Requests (PRs) for bug fixes, conducting thorough code reviews for peers, and writing robust unit/integration tests to prevent regressions.
- Performance Profiling: Experience using profilers (e.g., cProfile, JProfiler) to analyze thread dumps and heap dumps to resolve performance bottlenecks.
Microservices & API Architecture:
- Protocols & Patterns: Deep understanding of RESTful API design, gRPC protobufs, and asynchronous communication patterns (Pub/Sub, Kafka, SQS).
- Resiliency Patterns: Implementation of Circuit Breakers, Retry logic (exponential backoff), and Rate Limiting to prevent cascading failures.
- Service Mesh: Hands-on experience with traffic splitting, mutual TLS (mTLS), and observability sidecars.
- Distributed Tracing:
Analyzing request flows across microservices using tools like OpenTelemetry to pinpoint latency hotspots.
Container Orchestration (Kubernetes - GKE/EKS preferred):
- Cluster Administration: Managing GKE/EKS upgrades, node pools, and understanding the control plane components (API Server, Scheduler, Controller Manager, etcd).
- Networking & Security: Configuring Ingress Controllers (Nginx/ALB), Network Policies (Calico/Cilium), and Pod Security Policies/Contexts.
- Resource Management: Tuning resource requests/limits, setting up Horizontal/Vertical Pod Autoscalers (HPA/VPA), and debugging errors.
- Helm & Operators: Writing and managing complex Helm charts and understanding how Kubernetes Operators automate stateful applications.
Automation & Tooling:
- Internal Tools Development: Building custom CLI tools or web dashboards to empower L1/L2 support teams (e.g., a "one-click" user reset tool or a log analyzer dashboard).
- Event-Driven Automation: Creating self-healing workflows (e.g., a Lambda function that automatically restarts a hung process or cleans up temporary files when disk alerts fire).
Infrastructure as Code (IaC):
- Terraform Mastery: Structuring Terraform projects using reusable modules, managing remote state locking (S3 + DynamoDB), and using workspaces for multi-environment management.
- Configuration Management: Using Ansible playbooks for OS-level hardening, patch management, and configuration drift detection.
- Policy as Code: Implementing tools like OPA (Open Policy Agent) or Sentinel to enforce compliance checks within the IaC pipeline (e.g., preventing public S3 buckets).
Databases (SQL & NoSQL):
- Relational (PostgreSQL): Analyzing outputs to optimize slow queries, managing connection pooling (PgBouncer), and configuring replication/WAL archiving for disaster recovery.
- NoSQL (Cassandra/MongoDB): Understanding consistency levels, partition keys, sharding strategies, and diagnosing compaction or garbage collection issues.
- Caching: Implementing and troubleshooting caching layers (Redis/Memcached) to offload database read pressure.
DevOps Toolchain:
- CI/CD Pipelines: Designing complex Jenkins pipelines (Groovy shared libraries) that include parallel build stages, automated testing, and canary deployments.
- Version Control: Advanced Git strategies (Gitflow/Trunk-based), handling large merge conflicts, and using git hooks for pre-commit checks.
- Artifact Management: Managing Docker registries and dependency repositories (Artifactory/Nexus) with lifecycle policies to clean up old builds.
📌 Engagement Manager - Support Lead (Thiruvananthapuram)
🏢 Quantiphi Analytics Solutions
📍 Thiruvananthapuram