Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
  1. Home
  2. Blog
  3. Our AI Pilot Works, but It Lives in One Person's ChatGPT Account. How Do We Make It Real?
Problems We Solve

Our AI Pilot Works, but It Lives in One Person's ChatGPT Account. How Do We Make It Real?

An AI pilot that runs on one person's Custom GPT is not a system. We turn it into a proper tool with shared access, logging, tests and a route into production.

Updated 3 min readBy SpiderHunts Technologies

Free estimateNo obligation

Get a free estimate

Tell us what you need. A senior engineer reads every enquiry.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →

Quick answer — TL;DR

A pilot that runs as one person's Custom GPT, or a set of prompts they paste in each morning, proves the idea but cannot be relied on. To move it to production you need the prompts and knowledge moved into a system the business owns, connected to the tools where the work happens, with a test set that shows it still works after every change.

The pilot that everybody quietly depends on

It started as an experiment. Someone keen, often in operations or customer service, built a Custom GPT that drafts replies to supplier queries, or checks quotes against a price list, or turns site notes into a report. It worked. Colleagues started sending them things to run through it. Now a real part of the week depends on it.

And it all lives in one person's account. The instructions are in their Custom GPT settings, the reference files are uploaded under their login, and the process is "send it to Sam and Sam will run it". When Sam is on holiday the work waits. Nobody else is quite sure what the instructions say. Management has been told the pilot was a success and keeps asking when it will go live, and nobody can say what "live" would even involve.

Why it gets stuck at this stage

The pilot proved the model can do the task. It did not prove anything about running the task as a business process, and those are different questions. A pilot in a chat window has no connection to your systems, so someone copies inputs in and outputs out. It has no record of what it produced last Tuesday. There is no way to tell whether a change to the instructions made it better or quietly worse.

The bigger blocker is ownership. Nobody commissioned it formally, so nobody owns the move to production. IT sees it as a personal experiment. The person who built it has a day job. The pilot is stuck because it is sitting in the gap between a successful experiment and a funded project.

What staying in limbo costs

While it stays a pilotThe effect
One person runs itWork queues behind their availability and holidays
Inputs copied by handTime goes on shuffling text between windows, and mistakes creep in
No log of outputsWhen a customer disputes something, there is no record of what the AI wrote
Instructions changed ad hocQuality drifts and nobody can say when or why
Reference files out of dateAnswers are based on last quarter's price list or policy

Then there is the risk that sits on top of all of it. If that person leaves, the pilot leaves with them, along with the knowledge of why each instruction was written the way it was.

How we move it into a system you own

  1. We sit with the person who built it and extract everything: the instructions, the reference files, the examples they use, and the unwritten rules they apply when checking output.
  2. We collect real past cases, including the awkward ones, and turn them into a test set with the answers a competent person would accept.
  3. We rebuild the logic as a small service in your cloud account (Azure or AWS), calling OpenAI or Anthropic Claude through the API, with prompts stored in version control rather than a settings box.
  4. We connect it to where the work actually arrives, for example a shared mailbox in Microsoft 365, a form, or a folder in SharePoint, so nobody has to copy things in.
  5. We add a review step where a person approves or corrects output before it goes anywhere important, and the corrections are recorded.
  6. We run the test set every time the prompt, model or reference data changes, and we keep a log of every input and output.

The person who built the pilot stays involved as the subject expert. Their work is what makes the system useful; the rebuild just stops it depending on their login.

What the team works with afterwards

The task runs whether or not the original builder is in. Inputs arrive on their own, output lands where people already work, and anyone with the right role can review and approve. When someone asks what the AI said to a supplier three weeks ago, there is an answer.

Changes become safe to make. If you want to update the price list or tighten the tone, the test set shows whether anything else broke before the change goes out. That is the difference between a clever trick and something you can build the week around.

Is this your situation?

  • A useful AI process runs from one person's ChatGPT, Claude or Copilot account.
  • Colleagues send that person work to "run through the AI".
  • Management has called the pilot a success but there is no plan for going live.
  • Nobody else could say exactly what instructions the AI is following.
  • You would lose the process if that person left.

FAQ

Frequently asked questions

The questions readers ask us after this guide.

Still have a question?

Ask us directly — a senior engineer will get back to you.

Ask about your project

Can we not just share the Custom GPT with the team?

Sharing helps with access but not with logging, testing, integration or ownership. It is a reasonable stop-gap while the proper version is built.

Will the rebuilt version give the same answers as the pilot?

It should give answers at least as good on the test set, and we compare the two side by side before switching. Where they differ, the test cases show which is right.

Do we have to use the same AI model?

No. Once the task has a test set, we can try other models and pick on quality and running cost rather than habit.

What drives the cost of productionising a pilot?

How many systems it has to read from and write to, how much review workflow is needed, and how messy the reference data is. The AI part is usually the smaller share.

What do you need from the person who built it?

Some of their time to walk us through how they use it, the files and instructions, and a set of real examples with the answers they would accept.

Keep reading

More on Problems We Solve

Start here

Tell us where AI is going wrong for you

Describe what your team is doing with AI today, what is worrying you and which systems are involved. We will tell you what we would build, what we would leave alone, and if a smaller change would fix it, we will say so.

  1. You tell us what you needTwo minutes on the form, or a message on WhatsApp.
  2. A senior engineer reviews itAnd comes back with questions, a realistic range and an honest view on fit.
  3. Free 30-minute scoping callWe talk through scope, options and a realistic estimate — with no obligation.
Free estimateNo obligation

Talk to someone who builds this

Send a short brief and we will come back with an honest view and a realistic range.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →