01 Aug
|
Praveza Consultancy Private
|
Bengaluru
01 Aug
Praveza Consultancy Private
Bengaluru
Company Name: Cognizant
Title : Site Reliability Engineer
Job Type: Full Time , Work from Office
Location : Bangalore
Years of Experience : 7+Years
Salary : Open Experience: 7 to 11 years (SA), 11+ years (M)
Location: Bangalore & Ready to relocate to Bangalore only
Skills Required:
• Docker/Container/Java/Splunk
• Appdynamics/Dynatrace/Load runner/Jmeter/Capacity planning/RCA/Production support
• AWS/GCP Mandatory skills
• Container platforms like Docker/Kubernetes/Docker EE/OpenShift/Meosphere • Splunk
• Java
• AWS / GCP
Roles & Responsibilities: * On Call responsibilities to help minimize MTTD and MTTR * Experience with containerization and container platforms. (e.g., Docker, Kubernetes, Docker EE, OpenShift, Mesosphere) * Should have skills to understand debugging info , “Drain” traffic away from a cluster, Rollback a bad software push , block or rate limiting unwanted traffic, bring up additional serving capacity thru autoscaling features and use the monitoring systems(for alerting and dashboards) * Engage with enterprise and business/infrastructure functions to establish, track, and optimize operational metrics and targets in line with SRE principles (SLO/SLI, Latency percentiles , error budgets, tech debt and setup alert guidelines ) * Work with Observability tools and enterprise monitoring solutions like Dynatrace, AppDynamics, New Relic, Prometheus, Graphite, Grafana, Nagios, Sensu and Splunk . Should be able to write promQLs and Splunk queries .
* Programming/Tooling and Automation experience in one or more of the following languages: Golang, Java, Python, Typescript, Node and Shell . * Positive understanding of Kafka internals , SQL/noSQL databases like Cassandra , Elasticsearch and Postgress and In-Memory Caching frameworks like Memcached . * Influence, design and create new architectures, standards, and methods for large-scale enterprise systems. * Design, write and build tools to improve the reliability, latency, availability and scalability of Walmart e-commerce/Retail and Enterprise products. * Engender reliability and availability starting with metrics and measurements. * Enable scaling by providing tools, developing training and/or augmenting processes. * Build tools/automate to prevent re-occurrence of problem to mission critical products/services. * Augment existing instrumentation to build a cohesive picture of the characteristics of our systems with special attention to points of failure. * Participate in capacity planning, demand forecasting, software performance analysis and system tuning. * Develop a deep understanding of the numerous services and applications that come together to deliver Walmart e-commerce/Retail and Enterprise products * Root-cause analysis complex problems involving multiple parties, networks, hardware, and software that relate to scaling and performance. * Secure the system from issues, be they real, perceived, or notional.
📌 Site Reliability Engineer (Bengaluru)
🏢 Praveza Consultancy Private
📍 Bengaluru