08 Oct
|
Ford Technology Services India
|
Chennai
08 Oct
Ford Technology Services India
Chennai
Overview:
TekWissen is a global workforce management provider throughout India and many other countries in the world. The below client is a global company with shared ideals and a deep sense of family. From our earliest days as a pioneer of modern transportation, we have sought to make the world a better place – one that benefits lives, communities and the planet
Job Title: Platform Engineering Engineer 3
Location: Chennai
Work Type: Hybrid (4 Days work From Office)
Position Description:
- A mid-to-senior level infrastructure and sytem engineering role tasked with owning platform availability, Infrastructure as Code (IaC), CI/CD pipelines, and deep runtime performance tuning (especially JVM-at-scale).
- The engineer bridges the gap between cloud architecture, SRE disciplines, runtime optimization, distributed backend systems, and modern AI-augmented platform operations.
Core Responsibilities:Platform Management & Reliability
- Define/own SLA/SLO/SLI and error budgets; build and maintain CI/CD pipelines (Jenkins, GitLab CI, GitHub Actions, ArgoCD)
- Manage Kubernetes at cluster level (RBAC, quotas, autoscaling — HPA/VPA/Cluster Autoscaler) and Infrastructure as Code (Terraform, Pulumi, Ansible)
- Own on-call rotation, incident response, and blameless postmortems
JVM Configuration & Tuning
- Heap/generational sizing (-Xms/-Xmx, Young/Old ratios) and GC selection/tuning (G1GC default, ZGC/Shenandoah for low-latency, Parallel for throughput)
- Thread pool sizing vs. core count/I/O profile; JIT/tiered compilation flags; Metaspace and off-heap (direct buffer) management
- Diagnostics via jstack, jmap, async-profiler, JFR; GC log monitoring integrated into observability stack
Application Architecture Optimization
- Capacity planning and load testing (JMeter, Gatling, k6); bottleneck profiling (CPU/I/O/GC-bound)
- Microservices vs.
monolith tradeoffs and service decomposition
- Caching strategy (Redis, Memcached, CDN) with invalidation design; connection pooling (HikariCP)
- Resilience patterns: circuit breakers, retries/backoff, bulkheads (Resilience4j)
Storage & Data Infrastructure
- Storage selection (block/object/file: EBS, S3, NFS, Ceph) and tiering for cost optimization
- DB performance tuning (indexing, query plans, replication lag) and consistency model tradeoffs (strong vs. eventual)
- DR strategy: RPO/RTO targets, snapshot policy, cross-region replication
Backend Composition & Architecture
- Service mesh (Istio, Linkerd): traffic management, mTLS, observability
- API gateway/load balancer configuration (rate limiting, routing algorithms); message queue/streaming design (Kafka, RabbitMQ, SQS) — partitioning, consumer scaling
- Deployment strategy: blue-green, canary, multi-tenancy design
Leveraging AI
- Infra automation: AI-assisted IaC generation/review (Terraform/K8s manifest drafting), config drift detection, auto-generated runbooks from incident data
- Incident response: AI-driven log/anomaly triage (correlating metrics, traces, logs pre-alert), root-cause suggestion during on-call, auto-summarized postmortems
- Capacity & performance: ML-based autoscaling/forecasting (beyond static HPA thresholds), predictive capacity planning from historical load patterns
- Code/config review: LLM-assisted PR review for IaC and pipeline configs,
JVM flag recommendation based on workload profiling data
- Observability: Natural-language query interfaces over metrics/logs (e.g., Grafana/Datadog AI assistants), automated dashboard/alert-rule generation
- Toolchain fluency expected: working knowledge of AI coding assistants (Claude Code, Copilot, Cursor) integrated into IaC/pipeline workflows; increasingly listed as a preferred/required skill in senior postings
Required Technical Skills:
- Java/JVM (required for tuning depth), Python/Go/Bash for automation
- Cloud: AWS/GCP/Azure (one at expert level); Docker/Kubernetes (CKA-level often expected)
- Observability: Prometheus/Grafana, ELK/OpenSearch, Datadog, OpenTelemetry/Jaeger/Zipkin
- Networking: DNS, LB, TCP/IP, VPC/subnet design, security groups
- Security: IAM, secrets management (Vault/KMS), patching, vuln scanning
Skills Required:
- Java/JVM (required for tuning depth), Python/Go/Bash for automation Cloud: AWS/GCP/Azure (one at expert level); Docker/Kubernetes (CKA-level often expected)
- Observability: Prometheus/Grafana, ELK/OpenSearch, Datadog, OpenTelemetry/Jaeger/Zipkin Networking: DNS, LB, TCP/IP, VPC/subnet design, security groups Security: IAM, secrets management (Vault/KMS), patching, vuln scanning
Experience Required:
- Engineer 3 Exp: Proficient In 2 coding lang. or adv. Prac. in 1 lang. 6+ years in IT; 4+ years in development
Experience Preferred:
- Runbook documentation, architecture decision records (ADRs)
- Cross-team collaboration (SRE, security, app dev)
- Capacity forecasting and cost optimization (FinOps)
Education Required:
- Bachelor's Degree
- BS CS/Engineering or equivalent;
TekWissen® Group is an equal prospect employer supporting workforce diversity.
Pay: ₹693,000.00 - ₹1,910,000.00 per year
Work Location: In person
📌 Platform Engineering Engineer 3 (Chennai)
🏢 Ford Technology Services India
📍 Chennai