Come work at a place where innovation and teamwork come together to support the most exciting missions in the world!
We are seeking a talented Lead Software Engineer – Performance to deliver roadmap features of the Unified Asset Inventory ETM Platform, which helps customers Measure, Communicate, and Eliminate Cyber Risks.
You will lead performance engineering efforts across Java microservices, Spark, Kafka, Elasticsearch, and middleware APIs, ensuring that our real-time data pipelines and services meet enterprise-grade SLAs.
As part of our high-performing engineering team, you will design and execute performance testing strategies, identify system bottlenecks, and work with development teams to implement performance improvements that support the processing of billions of cybersecurity events per day across our data platform.
Responsibilities
- Own the performance strategy across distributed systems, including Spring Boot microservices, Hadoop, Spark, Kafka, Elasticsearch/OpenSearch, Big Data components, and APIs for each release.
- Demonstrate strong expertise in performance engineering, including analysis of Heap Dumps, Thread Dumps, GC Logs, and CPU Profiling Reports, with hands-on experience using Apache JMeter, AppDynamics, Grafana, and Prometheus, and a deep understanding of JVM architecture, Garbage Collection, Threading, and Concurrency.
- Define, develop, and execute performance test plans, load tests, stress tests, and soak tests.
- Create realistic performance test scenarios for data pipelines and microservices based on production-like workloads.
- Proactively identify bottlenecks, resource contention, and latency issues using tools such as JMeter, Spark UI, Kafka Manager, Elastic Monitoring,
and AppDynamics.
- Provide deep-dive analysis and recommendations for tuning and scaling Spark jobs, Kafka topics/partitions, Elasticsearch queries, and API endpoints.
- Collaborate with developers, architects, and infrastructure teams to integrate performance feedback into system design and implementation.
- Simulate and benchmark real-time and batch data flows at scale using synthetic and production-like datasets, and own the performance framework end-to-end for the synthetic data generator.
- Lead the initiative to build a performance testing framework that integrates with CI/CD pipelines.
- Establish and track SLAs for throughput, latency, CPU/memory utilization, and Garbage Collection.
- Create performance dashboards and visualizations using Prometheus/Grafana, Kibana, or equivalent tools.
- Document performance test findings and prepare technical reports for leadership and engineering teams.
- Recommend performance optimizations to Development and Platform teams.
- Take responsibility for optimizing the overall cost.
- Contribute to feature development and fixes in addition to performance benchmarking.
Required Qualifications
- Bachelor's degree in Computer Science, Engineering, or a related field.
- 8+ years of overall experience in distributed systems and backend performance engineering.
- 4+ years of Java development experience with Microservices architecture.
- Proficiency in scripting using Python and Bash for automation and test data generation.
- 4+ years of hands-on experience with Apache Spark, including performance tuning, memory management, and DAG optimization.
- 3+ years of experience with Kafka, including topic optimization, producer/consumer tuning, and lag monitoring.
- 3+ years of experience with Elasticsearch/OpenSearch, including query profiling, indexing strategies, and cluster optimization.
- 3+ years of experience with performance testing tools such as JMeter or similar.
- Excellent programming and design skills, with hands-on experience with Spring and Hibernate.
- Deep understanding of middleware and microservices performance, including REST APIs.
- Strong knowledge of profiling, debugging, and observability tools, such as Spark UI, Athena, Grafana, and ELK.
- Experience designing and running benchmarks at scale for high-throughput environments in PBs.
- Experience with containerized workloads and performance testing in Kubernetes/Docker environments.
- Solid understanding of cloud-native architecture (OCI) and distributed systems design.
- Strong knowledge of Linux operating systems and performance-related improvements.
- Familiarity with CI/CD integration for performance testing, such as Jenkins and GitHub.
- Knowledge of data lake architecture, caching solutions, and message queues.
- Strong communication skills and experience influencing cross-functional engineering teams.
Preferred Qualifications
- Prior experience with analytics platforms on Big Data would be a solid plus.
📌 Lead Software Engineer - Performance (Pune)
🏢 Qualys
📍 Pune