Principal AI Architect (Keesara)

Principal AI Architect (Keesara)

12 Sep
|
Cutshort
|
Keesara

12 Sep

Cutshort

Keesara

Role: Principal AI Architect — Multimodal Video Intelligence

Location: India Remote, with overlap with Singapore working hours

Employment Type: Full-time

Reporting to: Founder / CEO

Function: AI Architecture, Multimodal AI, Video Intelligence, Media Representation

About The Client The client is building an AI-native media intelligence platform that transforms long-form video into structured, searchable, reusable and monetisable media intelligence.

The Platform Is Not Simply a Video-clipping Tool.

We Are

Developing a Persistent Intelligence Layer For Media, Where Video, Audio, Speech, Text, Objects, Scenes, Events, Entities, Emotions, Narrative Arcs And Commercial Signals Are Processed Into a Reusable Representation That Can Support Multiple Downstream Use Cases, Including

- short-form clip generation;
- semantic search;
- scene and narrative understanding;
- contextual advertising;
- shoppable video;
- creator and content analytics;
- automated editing workflows;
- future media-intelligence APIs.

We are looking for a Principal AI Architect who can define and guide the AI architecture behind this platform. Role Summary The Principal AI Architect — Multimodal Video Intelligence will own the technical architecture for AI systems, including multimodal video understanding, persistent media representation, model orchestration, evaluation frameworks, and production AI design.

This is a hands-on architecture role. The ideal candidate can move between research papers, model selection, system design, data schemas, prototype review, engineering trade-offs, and implementation guidance.

You will work closely with the Founder / CEO, senior AI engineers, computer vision engineers, backend engineers and external vendors to convert the product and IP vision into a robust technical system.

Key Responsibilities

- AI System Architecture
- Define the end-to-end AI architecture for long-form video understanding.
- Design the processing pipeline from video ingest to structured media intelligence.
- Define how vision, audio, speech, text, metadata and user signals should be fused.
- Design the architecture for reusable media intelligence rather than one-time clip generation.
- Ensure the system can support multiple downstream applications from the same processed media layer.
- Persistent Media Representation
- Design persistent media representation layer across multiple levels, including frame, object, shot, scene, segment, entity, event and full-video levels.
- Define what intelligence must be stored permanently versus computed on demand.
- Design schemas for temporal, spatial, semantic, narrative and commercial metadata.
- Define provenance, confidence, model versioning and evidence-tracking requirements.
- Ensure the representation remains usable even when underlying AI models are replaced or upgraded.
- Multimodal Model Strategy
- Select and evaluate appropriate models for video, image, audio, speech, OCR, entity extraction, scene understanding, action recognition, embeddings, reranking and LLM/VLM reasoning.
- Decide where to use open-source models, commercial APIs, fine-tuning or custom models.
- Define model interfaces so models can be swapped without breaking downstream systems.
- Guide model benchmarking for accuracy, latency, cost and scalability.
- Prevent over-dependence on any single model vendor or API.
- Temporal and Narrative Intelligence
- Design approaches for understanding long-form video structure, including scenes, events, story arcs, character/entity continuity and engagement peaks.
- Define methods to identify clip-worthy moments across different content types.
- Support narrative scoring, highlight ranking,



scene segmentation and coherence validation.
- Ensure that clips are not only visually interesting but contextually and narratively coherent.
- Evaluation and Benchmarking
- Define objective evaluation frameworks for AI outputs.
- Build or guide creation of benchmark datasets and UAT criteria.
- Define metrics for clip quality, scene accuracy, entity continuity, timestamp alignment, hallucination control, ranking quality, retrieval precision and cost efficiency.
- Establish model and prompt evaluation processes.
- Create regression-testing methodology when models, prompts, schemas or scoring logic change.
- Search, Retrieval and Knowledge Layer
- Design hybrid search architecture across transcript, visual events, metadata, embeddings and structured knowledge.
- Define when to use relational storage, vector databases, graph databases and object storage.
- Design queryable media intelligence for downstream APIs and applications.
- Support knowledge-graph or ontology-based representation where useful.
- Ensure retrieved outputs are evidence-backed and timestamp-grounded.
- Production AI Architecture
- Work with AI engineers to convert architecture into deployable services.
- Guide decisions on batching, GPU inference, model serving, queues, retries, observability and cost controls.
- Review pipeline designs involving FFmpeg, GStreamer, DeepStream, TensorRT, Triton, ONNX, cloud services and model APIs.
- Define failure-handling, reprocessing, versioning and rollback mechanisms.
- Support scalable design without premature overengineering.
- IP and Technical Differentiation
- Help translate AI architecture into defensible technical differentiation.
- Support patent-related technical disclosures where required.
- Identify what is proprietary versus commodity.
- Avoid building a generic wrapper over existing models.
- Ensure the architecture reinforces the core thesis of persistent, reusable media intelligence.
- Team Guidance
- Provide technical direction to senior AI engineers and computer vision engineers.
- Review designs, experiments, evaluation results and architecture decisions.
- Mentor engineers without becoming a pure people manager.
- Help define technical milestones for the first 90, 180 and 365 days.
- Support hiring, technical interviews and vendor evaluation where needed.

Required Experience The ideal candidate should have:

- 8+ years of AI/ML experience, with significant exposure to computer vision, video AI, multimodal AI, retrieval systems or production ML architecture.
- Strong experience designing AI systems, not only implementing isolated models.
- Hands-on experience with video understanding, temporal modelling, multimodal pipelines, VLMs, LLMs, embeddings, ranking or retrieval.
- Experience taking AI systems from prototype to production.
- Strong knowledge of Python and modern AI/ML frameworks such as PyTorch, TensorFlow, Hugging Face or equivalent.
- Experience with model evaluation, benchmarking, error analysis and dataset design.
- Understanding of production architecture: APIs, queues, databases, cloud, model serving, observability and deployment trade-offs.
- Ability to work with founders and engineers in a high-ambiguity startup environment.

Strongly Preferred Experience

- Video understanding, action recognition, scene segmentation,



event detection or video retrieval.
- Multimodal AI involving video, audio, speech, text and metadata.
- LLM/VLM orchestration for structured outputs.
- Prompt/version management, schema validation and hallucination control.
- Embedding search, vector databases, reranking and retrieval evaluation.
- Knowledge graphs, ontologies, entity resolution or temporal knowledge representation.
- Model serving using TensorRT, Triton, ONNX, vLLM, DeepStream or similar.
- Experience with long-form video, OTT, sports media, entertainment, creator platforms, advertising technology or social commerce.
- Experience contributing to patents, technical disclosures or investor diligence.

Technical Areas The candidate should be comfortable discussing and making architecture decisions across:

- Computer vision;
- video AI;
- multimodal fusion;
- speech-to-text;
- OCR;
- image/video embeddings;
- VLMs and LLMs;
- semantic search;
- vector databases;
- graph databases;
- temporal reasoning;
- ranking and scoring systems;
- prompt orchestration;
- model evaluation;
- model versioning;
- data lineage;
- GPU inference;
- cloud AI deployment.

What This Role Is Not This is not a role for someone who has only built:

- chatbots;
- basic RAG demos;
- LangChain prototypes;
- prompt-engineering workflows;
- simple OpenAI/Gemini API wrappers;
- dashboards over model outputs;
- classical computer vision demos without production architecture;
- MLOps pipelines without AI system-design depth.

The role requires architectural depth in AI systems, not just familiarity with AI tools. First 90-Day Expectations First 30 Days

- Review product thesis, patent direction, prototype plans and existing technical assumptions.
- Assess current team capability and architecture gaps.
- Define the first version of AI architecture.
- Identify immediate technical risks and validation priorities.

First 60 Days

- Deliver a detailed architecture document covering media representation, model stack, pipeline design, storage strategy, evaluation framework and implementation roadmap.
- Define the canonical media-intelligence schema.
- Define model-selection and benchmarking criteria.
- Guide senior engineers on first implementation milestones.

First 90 Days

- Help the team implement and validate the first working version of the persistent media-intelligence layer.
- Establish evaluation datasets and UAT metrics.
- Review prototype outputs and improve architecture based on evidence.
- Produce a 6-month AI roadmap with technical risks, milestones and resourcing needs.

Success Metrics The Principal AI Architect Will Be Successful If

- They have a explicit AI architecture that the engineering team can execute.
- The platform does not collapse into a generic clip-generation pipeline.
- The media representation is reusable across multiple use cases.
- Models, prompts and schemas are versioned and testable.
- AI outputs are measurable through objective benchmarks.
- Snehashish, Abhishek and other engineers have clear technical direction.
- The architecture supports both product execution and investor/IP defensibility.

Candidate Personality Fit The Right Candidate Should Be

- intellectually strong but practical;
- hands-on enough to review code and experiments;
- comfortable with ambiguity;
- willing to challenge assumptions with evidence;
- able to simplify complex AI architecture for engineers and investors;
- disciplined about evaluation, cost and production constraints;
- not attached to one model, tool or vendor;
- able to work in a founder-led early-stage startup.

Skills:- Multi-modal AI, Artificial Intelligence (AI), Computer Vision, Machine Learning (ML) and Python

📌 Principal AI Architect (Keesara)
🏢 Cutshort
📍 Keesara

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: principal ai architect (keesara) / keesara