Who Writes the Hint?
For the last few weeks I've been running an experiment on a puzzle game I'm building. It isn't on the App Store yet. The question is simple. Should an AI model write game hints, or should I keep the hint sentences I wrote by hand?
A hint in my game comes in four steps. The first is a small nudge, and each tap on "Tell me more" gives a bit more, until the last step gives the answer. I took 300 puzzle positions and had two models, Sonnet 5 and Haiku 4.5, write all four steps for each one. Each model did it twice. That's 1,200 AI hints, checked against my templates on the same 300 positions.
The short answer: I'm keeping the templates. Not because they were perfect. They weren't. But their mistakes were a few fixed sentences I could find and fix. The models' mistakes moved around from run to run, and became a headache to chase.
Nothing passed
Before I saw any results, I set the bar: zero misleading hints. A hint step is misleading if it says something false, or if it would send a player the wrong way.
Nothing met that bar. Not Sonnet, not Haiku, and not my templates. So, you know, we could only go up from here!
Here's how many of the 259 everyday puzzle positions had at least one misleading step:
My templates: 91
Sonnet 5: 64 in the first run, 72 in the second
Haiku 4.5: 29 in the first run, 41 in the second
(42 harder positions are left out here. Every source failed almost all of them.)
So my templates did worst. That surprised me. Almost all of their 91 came from one sentence in the first step. It told the player there was one move of a certain kind to find. It sounds fine. But on 90 of those positions there were several options, so the player could go looking for the wrong one.
Measured as I first tested them, nothing met the bar. The templates' failures were two kinds of fixed sentence, and once the game rewrote them, every check came back clean. The models' failures can't be fixed that way, because they come out different every time.
What I predicted, and what happened
On 25 September, before any results, I wrote down six predictions. I've judged each one as I wrote it, not as I wish I'd written it. Hindsight 20/20 and all that.
My templates follow the rules better than the AI. Partly right. They made no serious mistakes (a wrong fact, the answer given too early, or no hint at all), while the models had 2 to 12 per run. But every template hint broke a style rule, and the templates had more misleading hints.
The AI's most common mistakes are about style, not facts. Right. Style mistakes outnumbered factual ones at least nine to one.
The AI gives the answer away too early more often than it gets facts wrong. Wrong. Early giveaways were rare: at most 3 of 300.
The AI only does better on the hardest hints. Wrong on following the rules and on misleading hints. It did better elsewhere, and the hardest hints are where most of its mistakes were. I never measured whether its explanations were clearer.
Two predictions about a follow-up "why?" question. Not tested. I never ran that part.
Some of the biggest results weren't in any prediction:
The models' typical mistake was a vague place. They'd point at a part of the board without saying which part. My game shows one hint step at a time, so the player may not know where to look.
10 of 1,200 hints never arrived. The model kept thinking until it hit its limit, then returned nothing.
The cheaper model followed my style rules far better. Haiku broke them on about a quarter of everyday positions, Sonnet on about three quarters. Haiku still cost a bit more per hint, because it thinks about three times as long.
My automatic checks missed things. When I read 100 hint steps they had passed, without knowing who wrote them, 16 turned out to be misleading. So every number here leans on my own reviews, and on checks I built after that.
The real difference: same mistake, or a new one each time
My templates are predictable. The same puzzle position always gets the same hint, word for word. Every misleading template step came from one of 7 fixed sentences, and one sentence alone was 117 of the 174. They went wrong in places I could see and adjust. So I could fix them once and check the fix.
And that's exactly what I did! The game rewrote those sentences, and I reran the checks on the same 300 positions and on 140 new ones the fix had never seen. Both problems went to zero.
The models are not predictable. I gave each one the same instructions and the same 300 positions, twice:
Sonnet 5: serious mistakes on 16 positions in the first run and 7 in the second. Only 2 failed both times. Vague places on 48 and 62, only 7 in both.
Haiku 4.5: serious mistakes on 11 positions in the first run and 7 in the second. None failed both times. Vague places on 24 and 33, only 3 in both.
The mistakes mostly landed on different positions each time. There's no single sentence to fix.
So a clean test run doesn't promise a clean next one. That's the problem with putting a model in front of players.
One exception: the 42 harder positions. Every source struggled with those, the models included, and the models failed many of the same ones in both runs.
I only ran each model twice. That's enough to see the mistakes move, not to say how often any one comes back.
What it would cost me
The whole experiment cost $16.19 for 1,200 AI hints. But I ran it at a half-price rate where answers come back hours later. A real game can't wait hours.
At the normal rate, an AI hint costs roughly 2.4 to 3 cents. That sounds like nothing until you do the math. If a player asks for 20 hints a day, that's about 50 to 60 cents a day, for one player. That's a guess, not a forecast.
And cost isn't the only catch. Live AI hints would mean:
Players need the internet for every hint. On a plane or in the subway, no hint.
I pay for every hint, every day, for as long as people play.
I need a server of my own. The key that pays for the AI can't live inside the app, because anyone could pull it out and run up my bill.
This is where I drew a line. I don't want a game that costs me money every day it's being played. And I want people to be able to play offline if at all possible.
Live AI hints fail both. The templates are free, work offline and show up instantly.
What I'm doing next
I'm shipping the templates, with the fixed sentences. Their known problems are gone, and they cost me nothing per player.
The AI can still help. Just while I build, not while people play.
My next step is already planned. The hints still use puzzle jargon, the names experienced players give to solving techniques, and I want to rewrite them in plain words. That's a rewrite of the template sentences, which is exactly the kind of job an AI could draft. So I'll have it draft the new sentences, I'll review every one, and I'll run the same checks as in this experiment before anything ships. I pay once, while building. Nothing per player, and it still works offline.
An AI that runs on the phone itself would be free and offline too. I haven't tested one yet.
What this doesn't tell you
I was the only reviewer. Every judgment call was mine.
Two models, one set of instructions, two runs each. Different instructions or settings might do better or worse.
I didn't review every flag in the second run. 45 flagged hints were left unreviewed and not counted as misleading.
Some calls were made after I saw results. The biggest: how strictly to read that first-step sentence. Strictly, it's 90 misleading template steps. A looser reading gives 24, and the loosest gives 0. I chose strict.
The fixed templates were checked by my automatic checks, not by me. "Clean" means none of my checks found a problem. It doesn't mean I've reread every hint.
Things I never measured: whether the AI's explanations were clearer, how players handle a follow-up "why?", and how often a player taps "Tell me more" before finding the move on their own.
I went into this expecting to measure the AI against my templates. I ended up measuring my templates too, and they had the worst score on the board. What made them the right choice wasn't that they were good. It was that their mistakes stayed put long enough for me to fix them.