Instead of improving the model’s accuracy, I used its output as a draft

· Product · 한국어 · 日本語

An AI agent sent me the results of an experiment in extracting features from photos. All 10 test photos passed the success criterion at the time. But for each photo, the extraction result also contained 4 to 8 incorrect features. In one case it suggested moustache for someone with no moustache.

By that criterion, it was a success. But the results weren’t good enough to show on screen.

What I first wanted to build was a small tool that could extract the features needed for character generation from a photo. There was one condition from the start, though. I didn’t want to upload my original photos to an external AI service or a server. But picking features like hair length, hair color, glasses, clothes, and expression one by one from scratch was a hassle too.

So I decided to try having a local model in the browser suggest a limited set of features, without sending the original photo to a server. The AI agent proposed and worked out the candidate models, the evaluation method, and the success criterion, and I reviewed and approved them. The AI agent ran the actual experiments too.

When the first results came in, what struck me before the accuracy was that it really ran in the browser.

“Huh, it actually runs. That’s kind of fun.”

It succeeded, and something was off

The first success criterion was simple.

Pass if a photo yields 5 or more correct features.

The problem was that this criterion, which I had approved, counted only the features that were right. Even if moustache came up for someone with no moustache, the photo passed as long as five other features were correct.

So the next step was to revise the criterion to account for wrong values and missing key features, not just the correct ones. Under the new criterion, none of the first 10 photos passed, and even after several rounds of tuning on 10 newly chosen photos, at most 1 did.

Couldn’t I just use a better model?

Running a stronger model on a server is an option. But that would have meant sending photos to a server. Dropping the condition of not sending original photos anywhere would make it a different tool from the one I wanted to build.

That was when the AI agent suggested treating the model’s output not as a final answer, but as a starting point someone could edit. On separate photos that hadn’t been used for tuning, it got 7 to 8 of the 15 fields right on average. That wasn’t enough to use as the answer, but it looked usable as a way to avoid choosing every field from scratch.

So I accepted the proposal.

I used the model’s output as the form’s starting values

The approach I first had in mind was this.

We accurately extract your features from your photo.

After the change, it became this.

We look at your photo and fill in the feature form first. Fix only what’s wrong.

I made the values the model picked easy to check and correct right away. The same error is a wrong answer if it stays in an extraction result, but in a pre-filled form it is a field to fix.


Try Character Prompt

Your original photo is processed in the browser and is not sent to a server. Some values may be wrong. Make whatever adjustments you need, then give it a try.