About Xepmi AI
Xepmi AI is building Explainable Persistent Machine Intelligence: AI systems that remember context over time and can show the reasoning behind every decision they make. Our flagship product, VectorStream
, is an agentic go-to-market automation suite that turns scattered market data, buyer signals, and sales conversations into a pipeline that runs itself. It captures leads, detects intent, engages prospects across WhatsApp, email, and AI voice, and traces every outcome back to the signal that caused it. We're starting with real estate in India as our beachhead and building for global markets from day one.
Why You Should Apply Now
- High impact. The data you collect is the first step in VectorStream's pipeline. Every lead score, intent signal, and AI conversation depends on it.
- Interesting challenges. Energetic websites, JavaScript-heavy portals, messy unstructured listings, and turning all of it into clean, structured intelligence.
- Founding team member. You'll be one of the first people on the data team, with real input on our tools, processes, and culture.
About the Role We're looking for a Data Intelligence Analyst to own how VectorStream collects and structures data from the web. You'll design, build, and maintain the scrapers and extraction pipelines that feed our agentic system with market listings, project data, pricing trends, and buyer signals. This is a hands-on role that combines engineering with analytical thinking: you'll write the scrapers, and you'll also make sure the data they produce is accurate, useful, and ready for our AI agents.
As Our Data Intelligence Analyst, You Will:
Build and Maintain Data Pipelines
- Design, build, and maintain scrapers for real estate portals, public registries, project websites, and other market data sources.
- Write modular, well-documented code that's easy to maintain as source websites change.
- Build scheduled and event-driven collection workflows using n8n and Python.
Handle Complex Websites
- Extract data from dynamic, JavaScript-rendered pages using headless browsers like Playwright.
- Manage sessions, cookies, pagination, rate limiting, and proxy rotation for reliable large-scale collection.
- Use LLM-powered extraction (structured outputs, tools like Crawl4AI or Firecrawl) to turn messy, unstructured pages into clean data.
Turn Raw Data into Intelligence
- Clean, deduplicate, normalize, and enrich collected data so it's ready for scoring and retrieval.
- Analyze trends in pricing, inventory, and market activity, and surface insights that feed VectorStream's signal detection.
- Prepare data for our RAG pipelines and vector databases.
Monitor and Troubleshoot
- Set up logging, monitoring, and alerts so broken scrapers are caught quickly.
- Track data freshness, coverage, and quality, and fix issues before they reach our agents.
Collect Data Responsibly
- Follow website terms, rate limits, and India's Digital Personal Data Protection (DPDP) Act in every pipeline.
- Keep clear records of data sources, collection methods, and consent where personal data is involved.
Drive Continuous Improvement
- Evaluate new tools and approaches in web data extraction and AI-assisted parsing.
- Propose better ways to collect, validate, and use data across VectorStream.
You Are Likely to Succeed If You Have:
- Strong Python skills and hands-on experience with scraping tools such as Playwright, Selenium, Scrapy, or BeautifulSoup
- A solid understanding of HTTP, REST APIs, HTML/DOM structure, and how browsers render pages
- Experience handling cookies, headers, sessions, and proxies
- Comfort with SQL and data cleaning using pandas or similar tools
- An analytical mindset: you care whether the data is right, not just whether the script ran
- Clear communication in English with technical and non-technical people
Nice to Have
- Experience with n8n, Make, or other workflow automation tools
- Used LLMs for data extraction or structuring
- Exposure to real estate, property tech, or market research data
- Built dashboards or reports (Metabase, Power BI, Looker Studio)
- Familiarity with Docker, cloud deployment, or job schedulers
Tech You'll Touch Python · Playwright / Scrapy / BeautifulSoup · Crawl4AI / Firecrawl · n8n · Claude / OpenAI APIs · pandas · PostgreSQL / Supabase · FAISS / Chroma · Docker · Git
What We Offer
- Real ownership. You'll own VectorStream's entire data acquisition layer, not a small corner of it.
- Work directly with the founder. Fast decisions, no bureaucracy, and a front-row seat to building a product from scratch.
- Growth based on impact. Your progress depends on what you build and deliver, not tenure or office politics.
- Learn the modern stack. Work at the intersection of web data, automation, and agentic AI.
Details
- Location: [Ahmedabad / Gandhinagar / Remote]
- Working hours: [e.g. 10am–7pm IST, flexible]
How to Apply Send your resume, GitHub, and a short note about a scraping or data project you've built to
[email protected] with the subject line "Data Intelligence Analyst – [Your Name]."
Xepmi AI is an equal opportunity employer. We welcome applicants of all backgrounds.
📌 Data Intelligence Analyst (Ahmedabad)
🏢 Xepmi AI
📍 Ahmedabad