⚡ We're building our own AI brain — and we need an AI Infrastructure Engineer to bring it to life.
World Trade Factory (SKXYWTF LLC) is an AI-native financial intelligence company. We run SKORE — an institutional-grade AI equity research pipeline — and a suite of AI advisory services. We're moving our inference off third-party APIs and onto our own GPU infrastructure in Phoenix, AZ.
This is a hands-on systems role. You will physically set up and own the WTFXAI inference server — the hardware that powers our entire AI stack.
We have
✅ A GPU server being procured in Phoenix, AZ
✅ 8+ AI services ready to route inference to local hardware
✅ A transparent cost and strategic reason to own our own compute
✅ A founding team that moves fast and ships real products
You will
— Set up and configure a GPU inference server running vLLM and Ollama
— Quantize open-weight models (Llama 3.2, Mistral, Phi-4, FinBERT) for optimal VRAM efficiency
— Expose a unified OpenAI-compatible inference API via Cloudflare Tunnel
— Wire the inference server into our AI router so all open-weight calls route locally
— Build a monitoring dashboard showing GPU utilization, VRAM, temperature, and cost savings
— Produce a cost savings report showing API cost vs local inference at scale
You should know
— Linux systems administration (Ubuntu Server)
— CUDA and GPU optimization basics
— Python
— Networking fundamentals (ports, tunnels, reverse proxy)
— C++ or parallel computing experience a strong plus
— Verilog / hardware background welcome — we're thinking about FPGAs next
Remote-first. Travel to Phoenix, AZ for server setup (expenses covered). Internship to start with path to full-time for the right person.
DM me or apply at
[email protected] with your GitHub.
- #AIInfrastructure #SystemsEngineering #LLM #GPU #vLLM #FinTech #WorldTradeFactory #AIJobs #Inference
📌 The Future Belongs To You - AI Infrastructure Engineer — Inference & Systems (India)
🏢 SKXYWTF
📍 India