Ship It Twice
The first version of an AI feature shows you what the second one must be, and no requirements document can. Plan and budget for two versions, and learn from real users on the first.
Chronologiq’s first AI insights feature shipped, and users said “this is making things up.” The second version shipped, and they said “how did it know that?”
The model and the data were the same both times. The design around them changed.
Chronologiq and RecommendMe are two AI products our founder, Shan Peiris, built before neesh Inc. The engineering below is theirs, and the numbers are his rather than a client’s.
The instinct with an AI feature is to spend months on the perfect design and release it once. Plan for two versions instead. The first shows you what users need, which no requirements document predicts; the second is the product. The gap between “it works” and “users trust it” only shows once real people use real output.
In both products, the first version taught the second in five places.
Five first versions, and what replaced them
Parsing
From invented dates to dates read before the AI starts
Chronologiq’s first text-to-timeline feature sent pasted text straight to the AI model and asked it to find the events, dates and categories.
The model guessed dates, invented years the text never mentioned and miscategorized events, and the bad data went straight onto the timeline.
A date-parsing library (chrono-node) reads the dates first, and guidance for the user's domain explains ambiguous terms. The model arranges events around dates already checked.
The fix was giving the model less to do.
Insights
From plausible patterns to findings with the numbers behind them
The first insights feature gave the model 200 events across five streams and asked what patterns it saw.
The answers sounded authoritative and cited no evidence. Users could not check them, and some findings were invented.
Six statistical checks run first: clusters in time, correlations between streams, unusual values, gaps, durations and density. Only findings with a confidence of 0.4 or more reach the model, which explains them in plain English.
One line of confidence filtering removed more noise than any rewording of the prompt.
Upload
From blind trust to a preview the user controls
The first upload flow went from pasted text to events on the timeline in one step. Users disliked it.
Parsed events went straight into the database. Users found the errors after the timeline was built and fixed them one at a time.
A preview shows every parsed event with a checkbox. Users deselect what is wrong and approve the rest before anything is saved. The AI parses; the person has the final say.
Trust came from control. Where the second version parsed exactly as the first did, users still trusted it more, because they could override the output before it was saved.
Tone
From one tone for every letter to a tone that matches the rating
RecommendMe drafts recommendation letters. The first version wrote the same glowing letter whatever rating the recommender gave. Real recommenders write a three-star letter as supportive but measured.
Every letter used superlatives. A three-star and a five-star letter were indistinguishable, and the output read as dishonest.
The instructions change with the rating. Five stars: specific achievements. Three stars: fair, encouraging, focused on growth. One star: potential and effort only, never negative.
Any model can write a good letter. The fix was making it write an honest one, and a single input, the star rating, did that.
Billing
From charge first to charge on success
Both products charge credits for AI operations. The first version deducted the credits, then ran the operation, so a timeout or an outage left the user paying for nothing.
Credits were deducted before the AI ran. Refunds for failures were manual and slow, and users learned to distrust anything that spent credits.
Check the balance, run the AI step, then deduct the credits only on success, in one conditional database update that stops two requests from overdrawing the balance.
A single condition on the database update removed a whole class of bugs. Nobody would have found them without watching the first version charge users twice under load.
What the five fixes have in common
None of the five came from a requirements document or a better prompt. Each came from watching real users, and each follows one of four rules:
- Give the model less to do: read whatever ordinary code can read first.
- Give people control: a preview, an override, a way to check.
- Ground the output in evidence: statistics before language.
- Charge on success: never bill for a failure.
Budget for the second version from the start, and put the first in front of real users early.
How much of your AI budget is set aside for the second version?
Book Free Assessment