12 Aug
|
Dusker AI
|
Jaipur
About Dusker AI
Dusker AI specializes in benchmarking and evaluating AI agents to help organizations understand real-world performance. Using expert-driven frameworks, we assess AI systems across reasoning, reliability, adaptability, and safety to ensure they are truly production-ready. From conversational AI to autonomous agentic systems, we build cutting-edge evaluation frameworks that enable organizations to develop trustworthy, high-performing AI solutions.
Role Overview
We are seeking an experienced AI Safety and Red-Teaming Engineer to lead adversarial testing of large language models and autonomous agents before they reach production. You will design structured attack campaigns spanning prompt injection, jailbreaks, tool-calling abuse, data exfiltration, and multi-turn manipulation, then translate findings into reproducible safety benchmarks and mitigation guidance. You will work alongside evaluation scientists, agent engineers, and domain experts to turn one-off exploits into automated regression suites that run against every model release.
This role is ideal for someone passionate about adversarial machine learning, agentic system security, safety evaluation methodology, and responsible AI deployment.
Key Responsibilities
- Design adversarial test suites targeting prompt injection, jailbreaks, and unsafe tool invocation in agentic systems.
- Build automated red-teaming harnesses that scale manual attack strategies across models and configurations.
- Define safety taxonomies and severity rubrics used to score refusals, over-refusals, and harmful completions.
- Conduct hands-on red-team campaigns against RAG pipelines, retrieval paths, and multi-agent workflows.
- Architect reproducible evaluation infrastructure using Python, Docker, Kubernetes, and CI/CD pipelines.
- Collaborate with evaluation scientists and annotation leads to calibrate human review of borderline safety cases.
- Guide engineering teams on mitigations such as input sanitization, sandboxed tool execution, and guardrail policies.
- Document attack methodologies, reproduction steps, and residual risk in clear technical reports.
Required Qualifications
- Bachelor's or Master's degree in Computer Science, Security Engineering, Machine Learning, or a related technical field.
- 4+ years of engineering experience, including 2+ years working directly with large language models or agentic systems.
- Solid Python skills and proven experience building evaluation or testing frameworks from scratch.
- Practical knowledge of LLM attack surfaces, including prompt injection, jailbreak taxonomies, and tool-calling exploitation.
- Hands-on experience with agent frameworks such as LangGraph, LangChain, CrewAI, or AutoGen.
- Familiarity with RAG architectures, embeddings, and vector databases, along with their common failure modes.
- Working knowledge of Docker, Kubernetes, CI/CD, and at least one major cloud platform such as AWS, Azure, or GCP.
- Clear written communication and the judgment to report risk precisely without overstating it.
Preferred Qualifications
- Background in application security, penetration testing, or vulnerability research.
- Experience with RLHF, preference data collection, or post-training safety alignment.
- Contributions to open-source safety, evaluation, or red-teaming tooling.
- Published research or public write-ups on model safety, robustness, or adversarial evaluation.
📌 AI Safety and Red-Teaming Engineer. (Jaipur)
🏢 Dusker AI
📍 Jaipur