Who Works on Your Machine Learning Project
Last updated:
The mistake of hiring a data scientist on their own
Plenty of businesses have tried the obvious approach: hire a talented data scientist, give them access to the data and wait. Six months later there is a notebook with an impressive validation score and no way to run it anywhere. The problem was never the talent. It was that a production machine learning system needs several kinds of work, and most of them are not data science.
So we staff ML projects around the whole path from raw data to a prediction someone acts on. The people change shape with the project, but the principle stays the same: senior engineers who write production code, not experiment scripts that someone else has to rewrite.
The two people who stay throughout
Your lead engineer owns the model, the data pipeline design and every technical decision. They run the workshop, lead the proof of value, present each demo and answer your questions directly in a shared channel. They can explain why the model made a particular prediction without having to go and ask someone.
Your commercial contact handles scope, price, timelines and anything contractual. Both are named in the proposal. If either has to change, which is rare, you hear about it in advance and there is a documented handover. This matches how every SpiderHunts project works, as described in who you actually work with.
Who else appears, and when
| Role | When they are involved | What they do on an ML project |
|---|---|---|
| Lead engineer | Whole project | Problem framing, modelling, evaluation, technical decisions |
| Data engineer | Early weeks, then at retraining | Extracting, joining and cleaning source data into a repeatable pipeline |
| Backend engineer | Integration phase | Serving predictions through an API or writing them back to your systems |
| Designer | Only if a new interface is needed | Review screens and dashboards people will actually use |
| QA | Before shadow mode and go-live | Testing the pipeline, the failure paths and the integration |
| DevOps | Setup and launch | Deployment, monitoring and the retraining schedule |
On a small project, the lead engineer and one colleague may cover all of this between them. On a larger one there are more people, but the number you deal with directly stays at two.
Your side of the team matters as much as ours
The ML projects that go well have a domain expert on the client side who looks at predictions and says “that one is wrong, and here is why.” That feedback finds leaks in the data, mislabelled outcomes and business rules the history does not show. No amount of statistical skill replaces it.
- A decision owner who can approve scope and accept or reject the go-live recommendation
- A domain expert who reviews samples of predictions for an hour or two in key weeks
- Someone with data access who can run exports or grant read access quickly
- A system owner for wherever the predictions will land, when integration begins
These can be the same person in a small business. What they cannot be is nobody.
Why we keep ML teams small
Machine learning work is unusually sequential. You cannot engineer features until the data is understood, cannot evaluate until the features exist and cannot integrate until you know what the prediction looks like. Adding people to a sequential process mostly adds meetings.
A small team also means the person who found an oddity in the data in week one is the person who remembers it when the model behaves strangely in week eight. That continuity is worth more on ML projects than on most software, because so much of the knowledge lives in why the data looks the way it does.
The most expensive handover in machine learning is the one inside the project, when the person who understood the data leaves the person who deploys the model to guess.
Where the team works from
SpiderHunts has a UK head office in London and an engineering office in Lahore, Pakistan. Everyone on your project is on our own team; we do not use subcontractors. We state the arrangement plainly because it explains our pricing and means part of the team starts its day several hours before UK clients.
For UK, European and Gulf clients the working days overlap for most of the day. For US clients the overlap is narrower and we shift part of the team's hours rather than pretending otherwise. The practical detail is in working with us across time zones, and the kinds of models this team builds are listed on our machine learning service page.
Frequently asked questions
Do we get a data scientist or a software engineer?
Will the engineer from the discovery call work on our project?
Do you use subcontractors for machine learning work?
Do we need our own data scientist to work with you?
Want to meet the engineer before you commit?
Book a call about your data and the decision it supports. The engineer you speak to is the one who would lead the work.