18 Sep
|
Cdcx Technologies
|
Bangalore Rural
18 Sep
Cdcx Technologies
Bangalore Rural
The CoinDCX Journey: Building the Future of Finance
At CoinDCX, our mission is explicit - to make crypto and blockchain accessible to every Indian and enable them to participate in the future of finance. As Indias first crypto unicorn valued at $2.45B, we are reshaping the financial ecosystem by building safe, transparent, and scalable products that power adoption at scale. We believe that change starts together. It begins with bold ideas, relentless execution and people who want to build whats next. If youre driven by purpose and thrive in environments where your work defines the next chapter of an industry, youll feel right at home here.
About the Role
You'll establish CoinDCX's independent second-line function for operational resilience and infrastructure risk, across the exchange and the Web3 ecosystem we plug into, including Okto's protocol and wallet infrastructure. Today, capacity limits and release decisions for trading-critical systems are made entirely inside Engineering, with limited independent check. You'll build the gates that test whether our infrastructure claims are actually true: can we handle peak load, can we recover from a failure, is our chain/wallet infrastructure resilient. You won't build or run these systems, and you won't own custody or key-management security, that stays with InfoSec. Your job is to independently verify that what Engineering has built actually holds up, with the authority to say no to a release that isn't ready.
What Youll Do
Ongoing / BAU
- Standing independent risk review of every major infrastructure decision — new chain integrations, new order types, scaling events — before it ships.
- Operate the Release Governance Gate for every P0/P1 production release touching matching engine, ledger, or wallet services.
- Independently monitor multi-chain node redundancy, RPC infrastructure resilience, and key-signing quorum availability across exchange and Okto wallet architecture (resilience lens only).
- Quarterly resilience reporting into the CRO/Board risk pack.
- Recurring capacity and DR test cadence, owned jointly with the Assistant Manager.
- First point of escalation for any production incident with a systemic-resilience dimension.
Tentative Deliverables — 3 Months
- Build a complete map of single points of failure across the matching engine, trading API gateways, database clusters, and wallet/key infrastructure — including Okto's chain/wallet layer and node redundancy.
- Run the first independent capacity/load stress test against matching engine throughput at a defined multiple of peak historical volume (e.g. 5) to find our actual, tested breaking point — not our assumed one.
- Establish formal 2nd-line incident severity definitions (P0/P1/P2) and a mandatory post-mortem sign-off workflow for tier-1 systems.
- Meet with Engineering, SRE, and InfoSec/CISO leadership to agree on the resilience-testing vs. cybersecurity boundary — including where operational resilience testing of key-signing availability ends and InfoSec's custody/key-security ownership begins — and secure data access for independent testing.
- Draft an initial technology risk taxonomy and register aligned to the CoinDCX ERM Framework's Operational and Technology/security risk rows, referencing relevant external regimes (DORA, MAS TRMG, NYDFS Part 500) where they apply to in-scope entities.
Tentative Deliverables — 6 Months
- Stand up a formal Release Governance Gate: mandatory for all P0/P1 production releases touching the matching engine, ledger, or wallet services, requiring a documented rollback plan and automated test verification before deployment.
- Run the first unannounced, end-to-end disaster recovery failover test against Board-approved recovery time/point objectives, benchmarked against DORA/MAS TRMG-style supervisory expectations.
- Independently verify multi-chain node redundancy, RPC infrastructure resilience, and geographic/quorum availability for key-signing operations (failover only — not key-security configuration).
- Publish the first technology resilience dashboard feeding CRO/Board risk reporting.
- Hire, onboard, and stand up an operating rhythm with the Assistant Manager, Technology & Resilience Risk.
Tentative Deliverables — 1 Year
- Integrate automated capacity/load testing into the CI/CD pipeline so a deployment is gated automatically if engine performance regresses — not just tested after the fact.
- Capacity limits across every major system are tied to real, current test data — not assumptions — and refreshed at every material scaling event.
- The Release Governance Gate has a demonstrable track record of catching high-risk issues before they became incidents, with zero tier-1 releases shipped ungated since go-live.
- Every gap identified in the first two DR tests is closed and independently re-verified.
- Technology risk register and traffic-light escalation thresholds (aligned to DORA/MAS TRMG/NYDFS Part 500 where relevant) are fully operational and in active use.
You’ll Excel in This Role If You
- Have 8+ years leading infrastructure risk, SRE, or technology risk management at a high-throughput trading venue, fintech, or tier-1 cloud environment.
- Have deep technical grounding in matching engine architecture, low-latency APIs, distributed systems, container orchestration, and key-management availability/quorum design (the resilience layer — not the cryptography or security controls themselves, which sit with InfoSec).
- Are hands-on with performance/load testing tools (k6, JMeter or equivalent), observability platforms (Datadog, Grafana or equivalent), and cloud infrastructure (AWS/GCP) — you can run a test yourself, not just commission one.
- Hold a Bachelor's or Master's in Computer Science, Engineering, or a related field, or equivalent hands-on experience; certifications such as AWS/GCP solutions architect or a recognized DR/BCM credential are a plus, not a requirement.
- Working knowledge of the technology-risk expectations in DORA, MAS TRMG, or NYDFS Part 500 — enough to know what a regulator will ask for, even if you're not the one filing it.
- Can hold a release-gate "no" against Engineering leadership under commercial pressure, backed by data rather than opinion — direct, assertive communication is a requirement of this seat, not a nice-to-have.
- Think like an independent auditor of your own infrastructure, not like the person who'd defend it.
- Are comfortable owning genuinely ambiguous, cross-chain scope (including Okto) without an existing playbook to follow.
- Can translate deep technical risk into board-ready, non-technical language, and think in terms of building a repeatable function and cadence — not delivering a single audit report.
You’ll Know You’re Winning When
At 6 Months
- The release gate is live and enforced for all trading-critical changes; the first independently-verified DR test is complete.
- The Board has seen a tested — not merely asserted — DR result for the first time.
- The second Technology Risk hire is onboarded and executing testing under your direction.
At 12 Months
- Zero trading-critical releases have shipped without a documented rollback plan since the gate went live.
- Capacity ceilings across every major system are backed by real test data, not assumption.
- Both gaps identified in year-one DR tests are closed and independently re-verified.
- You are reviewing test results, not personally running each test — the register and cadence operate independently under the Assistant Manager.
Scope, Ownership & Boundaries (cross-referenced to the Risk Architecture)
- You own: independent testing and limit-setting for infrastructure resilience and capacity (CEX Risk Architecture, Pillar 3 — Operational, Technology & Custody Risk, resilience component); the release-governance gate; DR/BCP test verification; Okto chain, wallet, and node resilience monitoring.
- You partner on, but do not own: cybersecurity and exploit prevention, which the CoinDCX ERM Framework assigns to the CISO/InfoSec ("Technology/security" risk row) — your resilience testing and their security testing sit side by side and should be sequenced together, not merged.
Not this role: day-to-day system operation (stays with Engineering/SRE as 1st line); wallet key-custody security and penetration testing execution (stays with InfoSec — you consume and challenge their output, you don't produce it).
Why This Role Matters & What’s In It For You
You'll be one of the early hires standing up an independent risk function at an exchange running 24/7 trading infrastructure with no second-line check on it today. You'll have real veto authority over production releases, direct visibility to the CRO and Board, and the mandate to build the technology-resilience discipline from the ground up rather than inherit someone else's process.
Where We Work
We believe the best ideas emerge when people build together. Collaboration, speed and trust come alive when teams share the same space.
With this belief, we operate as a work-from-office organisation. This role is based out of our Bangalore office, where energy, alignment and innovation move in real time.
Perks That Empower You
We believe great people deserve great experiences.
- Design Your Own Benefits: Flexible perks to match your lifestyle
- Unlimited Wellness Leaves: Rest and recharge as you need
- Mental Wellness Support: Access to therapy and wellness resources
- Learning Sessions: Bi-weekly learning and growth opportunities
Ready to Build What’s Next?
If you’re looking for a role that gives you direct access to high-stakes decisions, deep impact and a chance to build the future of finance, this is it. Join CoinDCX and help us make crypto accessible to every Indian, together.
📌 Director/Associate Director - Technology Risk (Bangalore Rural)
🏢 Cdcx Technologies
📍 Bangalore Rural