The AI Tax
The call to the AI model is the small part of an AI feature. Most of the build is the work around it: preparing the input, checking the output, time limits, fair billing, fallbacks and cleanup. Budget for that.
The call to the AI model takes three seconds. The software around it took three months.
An AI demo sends a question to the model and shows the response. It looks like a weekend project. Then real users arrive, and most of the engineering goes on everything around the call: checking responses that come back in unpredictable shapes, handling time-outs, not charging twice when two requests arrive at once, and falling back so the user always gets a result.
We call this the AI tax. It is invisible in demos, and it is the part an AI estimate is most likely to leave out.
What gets budgeted
Someone asks
Text, a file or a question
The AI model answers
Send it to the model, get a response
The answer is shown
Display the output
This is what the budget usually covers.
What ships
In the two products below, an AI feature needed seven layers of software around the model call. Each exists because the product first shipped without it and watched it fail.
Preparing the input
Read dates with a parser, summarize long documents, split text at sentence breaks.
Checking the input
Check every input, fence off the user’s text, add the domain guidance.
Time limits
A limit on every outside call and a budget for the whole job.
Checking the output
Check the whole response, then each item, then the raw text.
Charging fairly
Charge only on success, in one update that cannot overdraw.
Fallbacks
A useful result even when a step fails.
Cleaning up
Delete temporary files, refund failures, record every charge.
Where the engineering time goes
Chronologiq and RecommendMe are two AI products our founder, Shan Peiris, built before neesh Inc. The engineering below is theirs, and the numbers are his rather than a client’s.
Across the two products, the split came out roughly like this:
Illustrative: approximate, from two projects; the proportions are the argument, not a measurement
The gold segment, the work on the model itself, is about an eighth of the effort. Everything else is the tax.
The seven layers, one by one
What to budget for
On the illustrative split above, the whole build is about eight times the work on the model itself. When you scope an AI feature, or read someone else’s quote for one, look for the work that makes it:
- trustworthy: checked inputs and outputs, grounded in evidence;
- forgiving: fallbacks, so a failure still returns something useful;
- fair: charged only on success;
- steady under load: time limits and cleanup.
What does your AI estimate include besides the call to the model?
Book Free Assessment