01 Sep
|
fourkites
|
Chennai
As a Senior Data Scientist, you will build and own machine learning models that power core prediction problems across the FourKites platform including ETA/ATA forecasting and message-based status extraction. You will work end-to-end, from data pipeline to production deployment and monitoring, turning noisy real-world logistics data into models that run at scale and directly move the needle on customer outcomes. You will work closely with product, engineering, and operations teams, hands-on building and shipping models yourself while also guiding the technical direction of other data scientists on the team.
What you'll be doing:
- Design, build, and productionize ML models for problems like ETA/ATA prediction, using regression, classification, and time-series forecasting techniques
- Develop NLP/LLM-based extraction pipelines for message-based ETA and status updates (text extraction, entity recognition)
- Own models end-to-end: data pipeline training deployment monitoring retraining
- Work with noisy, real-world logistics and supply chain data (GPS pings, check calls, carrier data) rather than clean, pre-processed datasets
- Diagnose gaps between offline evaluation performance and live production accuracy, and drive fixes
- Build and maintain automated training/retraining pipelines using orchestration tools such as Airflow
- Set up and maintain model monitoring and observability (e.g., Grafana) to catch drift and degradation proactively
- Replace manual or rule-based processes with ML-driven automation (e.g., automating manual check calls)
- Translate model performance improvements into business impact — operational savings, efficiency gains, and deal-relevant outcomes
- Mentor and guide other data scientists/engineers on technical approach and best practices
- Make build-vs-buy and architecture tradeoff decisions independently
About the team:
Our product and engineering teams are dedicated to providing the industry's best-in-class end-to-end supply chain visibility platform. We are committed to building a high-performing, ML-driven team that turns supply chain data into automated, proactive action — and we want you to help lead that effort.
Who you are:
- Robust ML fundamentals across regression, classification, and time-series forecasting
- NLP experience — text extraction, entity recognition, or LLM-based extraction
- Production ML experience — you've shipped models serving real traffic, not just built POCs or notebooks
- Strong Python and SQL skills — pandas, scikit-learn, and comfort querying large datasets (Redshift/Snowflake a plus)
- Experience with cloud and data infrastructure — AWS (S3, EC2), and orchestration tools like Airflow for training/retraining pipelines
- Experience setting up or working with model monitoring and observability tooling (Grafana or similar)
- Comfortable working with noisy, real-world data rather than clean, curated datasets
- Experience diagnosing and closing the gap between offline evaluation results and live production performance
- A track record of replacing manual/rule-based processes with ML solutions
- Ability to translate model output into business value and communicate that impact to non-technical stakeholders
- Experience collaborating cross-functionally with product, engineering, and operations teams
- Experience mentoring or guiding other data scientists or engineers
- Ability to make build-vs-buy and architecture tradeoffs independently
- A track record of reducing manual intervention or turnaround time through automation
- Excellent oral and written communication skills
Nice to have:
- Experience in logistics, supply chain, or transportation
- Familiarity with real-time/streaming data (Kafka)
- Exposure to LLM/GenAI applications in production
📌 Senior Data Scientist (Chennai)
🏢 fourkites
📍 Chennai