Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
  1. Home
  2. Blog
  3. Who Owns the Model, the Data and the Outputs?
Software Strategy

Who Owns the Model, the Data and the Outputs?

Ownership of trained models, training data and derived insights is often unaddressed in contracts. The clauses worth settling before work starts.

Updated 2 min readBy SpiderHunts Technologies

Free estimateNo obligation

Get a free estimate

Tell us what you need. A senior engineer reads every enquiry.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →

Quick answer — TL;DR

Contracts commonly cover code and say nothing about the trained model, the training data or whether the supplier may reuse what it learned. Settle those four questions in writing before work starts - afterwards, positions harden. General guidance, not legal advice.

The gap in a standard development contract

Most software development agreements deal with deliverables, IP in code and confidentiality. A machine learning project produces things those clauses do not cleanly cover.

  • The trained model - parameters learned from your data, not written by anyone
  • Engineered features and the transformation logic, which often encode real business knowledge
  • Labelled datasets created during the project, frequently the most expensive asset produced
  • The supplier's general learning, which they will inevitably carry to other clients

Silence on these is not neutral. It leads to a disagreement later, usually when the relationship is ending and goodwill is lowest.

Four questions to settle

  1. Who owns the trained model? Typically the client where it was trained on client data and paid for by the client, but say so.
  2. Who owns labelled data created in the project? Often overlooked and often the most reusable asset.
  3. May the supplier reuse what they learned? General expertise, yes - that is unavoidable and reasonable. Your specific data, features or model, that needs stating.
  4. What happens at the end? Exactly what is handed over, in what form, and whether it is usable without the supplier.

Handover that actually works

Owning a model is worthless if you cannot run it. A handover clause should list what must be delivered in enough detail that ownership is meaningful.

DeliverableWhy it is needed
Model artefacts and weightsThe model itself
Training and feature codeTo retrain when it degrades
Training data or a reference to itReproducibility and audit
Environment specificationTo run it at all
Documentation and evaluation resultsTo know what it does and how well
A working deployment exampleProof that what was handed over runs

That last line is worth insisting on. A handover that has never been tested by anyone other than the supplier frequently turns out to be incomplete.

Being realistic about supplier learning

A supplier cannot unlearn general techniques. Demanding that they never apply similar methods elsewhere is unenforceable and will either be refused or ignored.

The reasonable line is between general expertise, which they keep, and your specific assets - your data, your labels, your trained model, your particular feature definitions - which they do not reuse. Most competent suppliers will agree to that readily, and reluctance tells you something useful.

Third-party components

Models frequently build on open-source components and pre-trained models with their own licences, some of which restrict commercial use or impose conditions on derived models.

Ask for a list of components and their licences as a deliverable, and check it before launch rather than during due diligence for a funding round or a sale. Discovering a restrictive licence deep in a production system is an expensive surprise.

If the contract does not mention the trained model, nobody has agreed who owns it.

FAQ

Frequently asked questions

The questions readers ask us after this guide.

Still have a question?

Ask us directly — a senior engineer will get back to you.

Ask about your project

Is the trained model automatically ours?

Not automatically. It depends on the contract and jurisdiction, which is exactly why it should be stated explicitly.

Can we stop a supplier working with competitors?

Usually not, beyond specific confidentiality and non-reuse of your assets. Broad restrictions tend to be unenforceable.

What if we used a vendor's pre-trained model?

Then their licence governs what you may do, including whether you may use outputs commercially or fine-tune further. Read it before building on it.

Should we ask for source code as well as the model?

Yes. Without the training and feature code you cannot retrain, which means the model has a limited life.

Keep reading

More on Software Strategy

Software Strategy

Machine Learning Myths That Waste Budgets

Eight beliefs about machine learning that quietly inflate project costs, what is actually true instead, and how to spot each one in a proposal.

Start here

Want machine learning project details from us?

Tell us what you are trying to predict and roughly what data you hold. We will come back with an honest view on whether machine learning is the right tool, what the work would involve and a realistic cost range. If a spreadsheet would do the job, we will say so.

  1. You tell us what you needTwo minutes on the form, or a message on WhatsApp.
  2. A senior engineer reviews itAnd comes back with questions, a realistic range and an honest view on fit.
  3. Free 30-minute scoping callWe talk through scope, options and a realistic estimate — with no obligation.
Free estimateNo obligation

Talk to someone who builds this

Send a short brief and we will come back with an honest view and a realistic range.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →