A buyer's guide, written by a supplier

How to choose a fractional AI officer

Justas Butkus is a fractional AI officer based in Vilnius, Lithuania, working with mid-market companies and scale-ups across the UK, EU and US. He publishes the selection criteria he would want a buyer to apply, including the ones he finds difficult to answer.

Short answer

Ask what they have personally put into production, who is accountable when a system produces a wrong answer, and what they would tell you not to do. The strongest signal is whether someone will disqualify their own engagement; the weakest is a list of impressive logos.

Who wrote this, and why does that matter?

I sell this service, so treat this page accordingly. I have written it as the criteria I would want applied to me, including the questions I find least comfortable, because a selection guide that quietly favours its author is worth nothing to you and is obvious anyway.

The ten questions, in order

  1. What have you personally put into production, and can I see it? The single question that separates the field fastest.
  2. What would you tell us not to do? Judgement shows up as willingness to disqualify a brief.
  3. Who is accountable when it produces a wrong answer? Whether failure was designed for, not just success.
  4. What happens to our systems if we stop working together? Whether you end up owning what was built, or renting it.
  5. Do you sell anything you would also recommend? The conflict is not disqualifying if it is disclosed and managed.
  6. Can you describe a project you misjudged? A harder, more specific version of question two.
  7. How do you measure success, and did you say so before starting? Whether the engagement has a defined shape or drifts.
  8. What is the engagement structure, exactly? Milestones, hours, or something vaguer.
  9. How many other engagements are you running right now? A capacity check on someone claiming to also build and operate systems.
  10. Will you tell us when this role becomes unnecessary? The hardest thing to fake, because saying it costs the supplier the work.

The rest of this page takes each of these in turn.

What have you personally put into production?

The single most useful question, and the one that most cleanly separates the field.

A great deal of AI advice is delivered by people who have never had to operate an AI system after it launched. That is where the difficulty actually lives: models drift, edge cases arrive, an answer is wrong in a way that reaches a customer, and someone has to decide whether to roll back. Advice from people who have not experienced that tends to be confident about the wrong things.

Ask for something specific and verifiable. Not a client name under NDA, which they cannot give you and which proves nothing anyway, but something you can go and look at.

What would you tell us not to do?

A supplier who cannot name something they would decline is telling you they will take any brief you hand them.

Good answers are specific: this use case is not worth automating, your data is not ready for that, this problem is a process problem and AI will make it faster and no better. If every idea you float is met with enthusiasm, you are talking to someone selling capacity rather than judgement.

Who is accountable when it produces a wrong answer?

Every AI system is eventually wrong in a way that matters. The question is whether that was anticipated.

  • Is there a defined path for a customer or an auditor to challenge an automated decision?
  • Is there human oversight where it matters, and is it real oversight rather than a person clicking approve on a queue?
  • Can you reconstruct why the system produced a particular answer on a particular day?
  • Is there a rollback that has actually been tested?

Someone who treats these as compliance overhead will build you something that works in the demo. Someone who treats them as design constraints will build you something that survives contact with a regulator or an angry customer.

What happens if we stop working with you?

Ask it directly and early. You are looking for whether you end up owning the systems or renting them.

  • Are accounts and infrastructure in your name from the start, or theirs?
  • Is documentation a deliverable, or something produced at the end if the relationship is amicable?
  • Could a competent internal engineer pick this up from what is written down?
  • Is handover in the contract?

Do you sell anything you would also recommend?

This one is aimed squarely at people like me, and you should ask it.

Anyone who advises on AI strategy and also sells AI products has a conflict. The conflict is not disqualifying, and in fact the people who build things often give better advice than those who only advise. But it has to be handled explicitly rather than left unmentioned.

What you want to hear is a stated policy: what they will not procure from themselves, what they disclose, and at what point they step out of a decision. If the answer is that there is no conflict, that is either an evasion or they have not thought about it.

Can they describe a project they misjudged?

This is the harder version of "what would you tell us not to do." Anyone can claim, in the abstract, that they would decline the wrong brief. Fewer people can point to a specific occasion where they got the call wrong, either by taking on something they should have declined or by declining something that turned out to matter.

The answer you want is specific: what the situation was, what they misjudged, and what changed afterwards in how they decide. A vague answer about "learning and growing" is not that. An answer that names no mistake at all, from someone who has supposedly been doing this for years, is worth treating with more suspicion than an honest one.

How do they measure success, and did they say so before starting?

Ask this before signing anything, because the honest answer to it should already exist rather than being invented for the question.

  • A good answer names specific, checkable outcomes for the first ninety days, agreed in writing before the engagement starts.
  • A weaker answer talks about "alignment" or "momentum", which cannot be checked and will not be checked.
  • A bad answer is deferred: "we will figure out what success looks like as we go." That usually means it will be redefined afterwards to match whatever happened.

Without an agreed definition, evaluating the engagement six months in becomes an argument about impressions rather than a comparison against a number.

What is the engagement structure, exactly?

Three structures are common in this category, and they carry different incentives.

  • Milestones written into the engagement. What exists by day 30, 60 and 90 is specified in advance. The incentive points toward delivering those things.
  • Hourly or day-rate billing with no defined deliverable. The incentive points toward more hours, not toward finishing.
  • No written structure at all. Common with informal arrangements and the hardest to evaluate, because there is nothing to hold anyone to afterwards.

Ask to see the structure in writing before you sign, not just hear it described on a call.

How many other engagements are they running right now?

A useful capacity check, particularly for anyone claiming to also build and operate production systems as part of the offer.

There is no single right number, but the answer should be specific and should match the days-per-month figure they quoted you. Someone claiming unlimited capacity while also claiming to be hands-on with production systems is describing two things that do not fit together, and one of the two claims is probably softer than it sounds.

Will they tell you when this role becomes unnecessary?

The hardest signal to fake, because it is the one that costs the supplier money to say out loud.

Ask directly: "at what point would you tell us to stop paying for this?" A specific answer, naming a condition under which the engagement should end or convert to something smaller, is the strongest positive signal in this whole list. A general reassurance that "we will always find ways to add value" is the answer of someone optimising for renewal rather than for your outcome.

See how a fractional AI officer engagement usually ends for the three ways this normally plays out, so you have something concrete to compare their answer against.

Which signals are worth less than they look?

Selection signals, ranked by how much they actually predict
SignalWhat it tells you
Systems you can inspect yourselfA great deal. Hard to fake, easy to verify.
A written engagement structureA lot. Shows they have done this more than once.
Willingness to disqualify workA lot. Judgement rather than capacity.
Named clientsLess than it appears. Confirms someone paid, not that it worked.
An impressive prior employerLittle on its own. Working somewhere is not the same as having done the thing.
Certifications in AI strategyVery little. The field is too young for the credential to mean much. I hold one and I would still rank it here.
Confident numbers about your ROI, before diligenceNegative. Nobody can know that yet.

Frequently asked questions

What should I ask a fractional AI officer before hiring them?

What they have personally put into production and can show you, what they would advise you not to do, who is accountable when a system produces a wrong answer, what happens to your systems if the engagement ends, and how they handle any conflict between advising and selling.

How do I check someone is credible without client references?

Look for things you can verify directly rather than take on trust: systems you can inspect or use, published technical or regulatory depth, and a written engagement structure. Then run a working session on a real problem of yours before committing.

Is it a problem if they also sell AI products?

Not inherently, and builders often give better advice than pure advisors. It becomes a problem when it is unmanaged. Ask for a stated policy on what they will not procure from themselves and when they step out of a decision.

How many candidates should we talk to?

Two or three is usually enough, because the differences show up quickly once you ask what they would decline to do. Long processes tend to select for people who are good at long processes.

What is the biggest mistake buyers make?

Choosing on confidence. The category rewards people who sound certain, and certainty at the selection stage is unwarranted, because nobody can know what your data will support before they have looked at it.

Should we ask about a project they got wrong?

Yes, and it is a harder test than asking what they would decline. Anyone can claim they would say no to the wrong brief in the abstract; fewer can name a specific past misjudgement and what changed in how they decide afterwards.

How do we know if the engagement structure is sound?

Ask to see it in writing before signing: milestones for day 30, 60 and 90, agreed in advance. Hourly billing with no defined deliverable, or no written structure at all, both leave you with nothing to hold anyone to later.

What is the strongest positive signal in a selection process?

A specific answer to "at what point would you tell us to stop paying for this." It is the one signal that costs the supplier money to give honestly, which is what makes it hard to fake.

Ask me these questions

The list above is the one I would want applied to me. If the answers are useful, that is a reasonable basis for a conversation.