10 Aug
|
Cognite - AI for Industry
|
India
10 Aug
Cognite - AI for Industry
India
About The Role
We're seeking a Software Engineer who excels at building high-performance distributed systems and thrives in a quick-paced startup environment. You'll be working on cutting-edge data infrastructure challenges that directly impact how Fortune 500 industrial companies manage their most critical operational data.
How you’ll demonstrate Ownership
- Platform Ownership: Design, build, and operate the core serverless execution engine and Workflows orchestration layer that serve as foundational primitives for CDF’s AI and automation capabilities.
- Reliability Engineering: Own uptime, latency SLOs, and incident response for platform services ensuring Functions execute deterministically and Workflows progress without data loss or silent failures.
- Scalability: Architect for multi-tenant & multi-cloud, high-throughput workloads. Design scheduling, queueing, and retry mechanisms that degrade gracefully under pressure.
- API Design: Define and evolve clean API-first architecture, versioned REST and event-driven APIs that downstream engineering teams and external customers depend on.
- Observability: Instrument services with distributed tracing, structured logging, and alerting (Open-telemetry / Prometheus / Grafana / Honeycomb stack) so failures surface before customers notice.
- CI/CD & Testing: Champion test automation - unit, integration, and smoke tests and maintain deployment pipelines that ship to production with confidence.
- Performance: Profile and resolve bottlenecks in execution throughput, cold-start latencies,
and cross-service call chains driving a “snappy” platform experience for industrial workloads.
- Cost Efficiency (Bonus): Model compute and storage costs for functions execution; identify and implement optimizations that reduce cloud spend without sacrificing reliability.
The Impact you bring to Cognite
- 6–8 Years of Engineering: Proven track record building and operating production backend services at scale.
- Expertise: Deep mastery of JVM languages (Kotlin preferred, Java acceptable), Python(FastAPI), distributed systems patterns, and cloud-native service design (Kubernetes, Azure, GCP, AWS, Private cloud).
- Workflow & Orchestration: Hands-on experience with workflow engines (Conductor, Apache Airflow, or equivalent) and event-driven architectures (Kafka, Pub/Sub).
- Data & Storage: Comfortable working with relational databases (PostgreSQL) & non-relational databases, object storage(Data-lakes), and caching layers (Redis) in multi-tenant environments.
- Observability Stack: Practical experience with Open-telemetry, Prometheus, and Grafana for instrumentation and operational insight.
- ML Platform Exposure: experience supporting ML workloads & notebooks in production, whether through job scheduling, resource management, experiment tracking integration,
or model serving infrastructure.
- Contextualisation Domain (Bonus): Familiarity with industrial knowledge graph construction, entity resolution, or NLP/CV pipelines as they relate to industrial asset data is a strong differentiator.
- Full-Stack Awareness (Bonus): Familiarity with React or TypeScript is a plus for consuming and dogfooding your own platform’s developer tooling.
- The Platform Thinking Spirit: A passion for building composable, well-documented, and automated platform systems that empower other engineers including ML engineers to build faster.
Good to have
ML Platform & Contextualisation:
- ML Workload Support: Build and extend platform primitives compute scheduling, environment management, and secrets handling, that enable ML engineers to run model training, fine-tuning, and batch inference jobs reliably.
- Contextualisation Pipelines: Support the engineering infrastructure behind Cognite’s Contextualisation capabilities (entity matching, asset hierarchy inference, P&ID; parsing) by ensuring the platform can orchestrate long-running, GPU-aware, and data-intensive ML workflows without manual intervention.
- Vector & Embedding Infrastructure (Bonus): Familiarity with serving or storing vector embeddings to support semantic search and RAG-based contextualisation use cases.
- Model Lifecycle Awareness: Understand model versioning, A/B experiment tracking, and the boundary between platform concerns and ML framework concerns, so the platform stays lean while ML teams stay unblocked.
📌 Senior Software Engineer - Platform (India)
🏢 Cognite - AI for Industry
📍 India