09 Oct
|
Staffing Services
|
Bengaluru
09 Oct
Staffing Services
Bengaluru
Role & responsibilities
- Agentic Architecture & Interoperability:
- Define endtoend Edge AI system architecture, covering data acquisition, preprocessing, model execution, orchestration, and edgecloud integration.
- Evaluate and select hardware accelerators (GPU, NPU, DSP, TPU, VPU) based on workload characteristics.
- Architect solutions using frameworks like NVIDIA Jetson, Intel OpenVINO, Qualcomm AI Engine, ARM Ethos, or TPU Edge.
- Define model pipelines for realtime analytics, including vision, audio, signal processing, or sensor fusion workloads.
- Design and implement decentralized multi-agent architectures using Google ADK (Agent Development Kit) and LangGraph.
- Implement Agent-to-Agent (A2A) communication protocols to enable seamless collaboration between heterogeneous agents across diverse ecosystems.
- Integrate Model Context Protocol (MCP) servers to standardize how agents access data, tools, and enterprise resources securely.
- Generative AI Customization (SLMs):
- Lead the selection and customization of Small Language Models (e.g., Phi, Gemma, Llama, Qwen variants) for domain-specific tasks.
- Apply parameter-effective fine-tuning (PEFT) techniques such as LoRA, QLoRA,
to adapt models without incurring massive computational overhead.
- Edge AI & Inference Optimization:
- Optimize models for deployment on resource-constrained edge devices (Smartphones, Smart Glasses, Wearable, IoT gateways, Embedded Linux).
- Utilize techniques like Post-Training Quantization (PTQ), Quantization-Aware Training (QAT), and weight pruning to minimize latency and memory footprint.
- Execute inference using TFLite, PyTorch Mobile or ONNX Runtime, leveraging hardware acceleration (Qualcomm NPU/DSP) where available.
- Quantization (INT8, INT4, mixed precision), Pruning and sparsity, KV-cache optimization, Speculative decoding (where applicable), Batch and streaming inference optimization, Profile and optimize inference pipelines across CPU, GPU, NPU, and DSP accelerators, reduce cold-start latency and improve real-time responsiveness.
- Embedded & System Integration:
- Develop high-performance inference engines and middleware in C/C++ to interface AI models with hardware sensors and actuators.
- Build Android-native AI services using Java/Kotlin and Android NDK, ensuring efficient background execution and battery management.
📌 Opening For Edge AI Architect Bangalore/ Trivandrum/ Chennai/ Pune (Bengaluru)
🏢 Staffing Services
📍 Bengaluru