neesh Inc.
AI StrategyEngineering CostArchitecture

The AI Tax

The call to the AI model is the small part of an AI feature. Most of the build is the work around it: preparing the input, checking the output, time limits, fair billing, fallbacks and cleanup. Budget for that.

The call to the AI model takes three seconds. The software around it took three months.


An AI demo sends a question to the model and shows the response. It looks like a weekend project. Then real users arrive, and most of the engineering goes on everything around the call: checking responses that come back in unpredictable shapes, handling time-outs, not charging twice when two requests arrive at once, and falling back so the user always gets a result.

We call this the AI tax. It is invisible in demos, and it is the part an AI estimate is most likely to leave out.

What gets budgeted

1

Someone asks

Text, a file or a question

2

The AI model answers

Send it to the model, get a response

3

The answer is shown

Display the output

This is what the budget usually covers.

What ships

In the two products below, an AI feature needed seven layers of software around the model call. Each exists because the product first shipped without it and watched it fail.

1

Preparing the input

Read dates with a parser, summarize long documents, split text at sentence breaks.

2

Checking the input

Check every input, fence off the user’s text, add the domain guidance.

3

Time limits

A limit on every outside call and a budget for the whole job.

4

Checking the output

Check the whole response, then each item, then the raw text.

5

Charging fairly

Charge only on success, in one update that cannot overdraw.

6

Fallbacks

A useful result even when a step fails.

7

Cleaning up

Delete temporary files, refund failures, record every charge.

Where the engineering time goes

Chronologiq and RecommendMe are two AI products our founder, Shan Peiris, built before neesh Inc. The engineering below is theirs, and the numbers are his rather than a client’s.

Across the two products, the split came out roughly like this:

Preparing input
Preparing input 18%
Checking input 15%
Time limits 8%
Checking output 15%
Billing 14%
Fallbacks 12%
Cleanup 6%
Prompts and model 12%

Illustrative: approximate, from two projects; the proportions are the argument, not a measurement

The gold segment, the work on the model itself, is about an eighth of the effort. Everything else is the tax.

The seven layers, one by one


What to budget for

On the illustrative split above, the whole build is about eight times the work on the model itself. When you scope an AI feature, or read someone else’s quote for one, look for the work that makes it:

  • trustworthy: checked inputs and outputs, grounded in evidence;
  • forgiving: fallbacks, so a failure still returns something useful;
  • fair: charged only on success;
  • steady under load: time limits and cleanup.

What does your AI estimate include besides the call to the model?

Book Free Assessment