13 Aug
|
Appiness Interactive
|
Bengaluru
13 Aug
Appiness Interactive
Bengaluru
Company Description
Appiness Interactive Pvt. Ltd. is a Bangalore-based product development and UX firm that
specializes in digital services for startups to fortune-500s. We work closely with our clients to
create a comprehensive soul for their brand in the online world, engaged through multiple
platforms of digital media. Our team is young, passionate, and aggressive, not afraid to think
out of the box or tread the un-trodden path in order to deliver the best results for our clients.
We pride ourselves on Practical Creativity where the idea is only as good as the returns it
fetches for our clients.
About the Role
You will own the reliability of the distributed data systems, the streaming runtime and
processing engines that move hundreds of billions of rows per day for top-tier enterprises. This
is an SRE role for our big data stack: Kafka, Spark, Flink, Ray, Redis, and data warehouses, all
running on Kubernetes.
This is not a cloud-provisioning role. We are looking for someone who has lived inside stateful,
high-throughput systems in production who has chased down a broker outage, a checkpoint
stall, a crashlooping cache, and a sink that silently stopped writing, and who fixes the
architecture rather than the symptom. If keeping a large, busy data platform alive and fast is the
kind of problem you find satisfying, you will have a lot of fun working with us. This is a unique
opportunity to shape the foundation of a product that is defining the next wave of intelligent,
context-aware data movement.
Responsibilities
● Streaming & Data Plane Reliability: Own the health of our Kafka-based runtime
(managed via Strimzi on Kubernetes) - broker health, topic lifecycle and count
management, partition and throughput tuning, certificate/secret rotation, and version
upgrades - at a scale of hundreds of thousands of topics and hundreds of billions of rows
per day.
● Distributed Processing Engines: Operate and tune distributed system workloads in
production in collaboration with backend teams, resource allocation, autoscaling,
checkpointing, backpressure, and failure recovery for both batch and streaming jobs.
● Stateful Services: Run Redis clusters and other stateful systems reliably - failover,
persistence, liveness/readiness tuning, and capacity planning under heavy and bursty
load.
● Kubernetes & Operators: Take end-to-end ownership of Amazon EKS, Google GKE and
the operators (Strimzi and others) running our stateful data workloads - cluster lifecycle,
scaling, version upgrades, and resource governance.
● Observability: Build deep, data-aware monitoring - consumer lag, throughput, partition
skew, job latency, error rates - not just host and CPU metrics. Make the data plane's
behavior legible before it breaks.
● Incident Management: Lead root-cause analysis for distributed-systems failures (broker
outages, crashloops, sink decommissions, control-plane race conditions) and drive
durable fixes. Mitigate fast, but design out the recurrence.
● Infrastructure as Code & Automation: Provision and manage cloud infrastructure with
Terraform; build operational runbooks and automation, including for air-gapped/private
enterprise installs (pre-staged images, operator-facing procedures).
● Collaboration: Partner with platform, runtime, and connector engineering - and with
SREs and support - to ship and scale new data-movement features reliably in a
large-scale Linux workplace.
Qualifications
● Experience: 6+ years in infrastructure, SRE, or DevOps, with significant time spent
operating production distributed data systems (not just application/cloud infra).
● Kafka: Deep, hands-on operational experience running Kafka at scale in production -
ideally on Kubernetes via Strimzi - including upgrades, topic/partition management,
performance tuning, and TLS/secret rotation.
● Distributed Processing (Strong Plus): Production experience operating one or more of
Spark, Flink, or Ray - resource tuning, checkpointing, failure recovery.
● Stateful Systems (Must Have):
Production experience with Redis (clustering, persistence,
failover) and a solid understanding of operating stateful workloads on Kubernetes
(StatefulSets, PVCs, probes, operators).
● Data Warehouses: Familiarity operating against Snowflake, BigQuery, or similar, and an
understanding of JDBC connectivity and sink reliability.
● Kubernetes & EKS: Strong hands-on EKS cluster creation, scaling, version upgrades, and
operator management.
● Infrastructure as Code: Advanced proficiency with Terraform.
● Programming: Proficiency in Python (or similar) for automation and tooling. Comfort
reading and debugging JVM-based systems is a strong plus.
● Reliability Mindset: Demonstrated ownership of incident management, RCA, capacity
planning, and performance tuning for high-throughput systems.
● CI/CD: Solid understanding of CI/CD methodology (Jenkins, GitHub Actions, or GitLab CI)
for containerized and non-containerized apps. Supporting, not the core of the role.
● Nice to Have: Configuration management (Ansible preferred); broader AWS services
(IAM, VPC, EC2, S3, Lambda); AWS CloudFormation.
● Soft Skills: Excellent communication and organizational skills; ability to coordinate
effectively within a team and with customers.
Why This Might Be Worth It
● You own the hard part. The stateful, distributed systems that move billions of rows are
the platform's most demanding reliability problems - and they'd be yours.
● Impact at scale from day one. Your work keeps mission-critical data flowing for
companies like DoorDash and LinkedIn.
● The AI wave is real for us. We're not bolting AI onto a legacy product. Intelligent
connectors, context-aware data movement, and agentic workflows are the core of what
we're building next - on top of the runtime you'd run. ○ Small team, big problems. Direct
access to the CTO, real influence over product direction, and the autonomy to make
significant technical bets. ○ Recognized platform, startup energy. Enterprise validation
with the speed and ownership of an early-stage company.
Skills:- DevOps, Python, Bash, Shell Scripting, Terraform and cicd
📌 DevOps Engineer (Bengaluru)
🏢 Appiness Interactive
📍 Bengaluru