All Perspectives

May 15, 2026  ·  6 min read

Why Multi-Subject Image Generation Is So Hard

One photo of you. One of your dog. Four panels that look like you were there together. Here’s what it actually took to build.

A friend of mine was scrolling through her phone a few months ago and got a little sad. Hundreds of pictures of her dog. Hundreds of pictures of herself. Almost none of the two of them together — because she was always the one holding the phone.

That’s the problem we set out to solve. Simple: you upload a photo of yourself, a photo of your dog, and the app generates a vintage photobooth strip of the two of you together. Four sepia cells, a little film grain, the works.

It sounds like a simple build. It absolutely was not.

“Just use AI” doesn’t quite get there

Type “a golden retriever wearing sunglasses on the beach” into Midjourney or Flux and you get something genuinely impressive. The current AI models are great at making plausible images.

But “plausible” and “this is me and my dog” are very different bars. Once the user has a specific reference in their head — their own face, their own dog — the model has to nail not one identity but two, plus the way they relate to each other in the frame. Get the dog right and the human wrong, or vice versa, and the strip feels uncanny in a way users notice immediately even if they can’t quite name it.

You really need to be careful about showing the user a two-person photo or a deformed dog. It totally kills any excitement.

People don’t grade on a curve. One bad panel out of four isn’t a 75% success — it’s a strip they don’t want to keep.

What’s actually in the box

Behind the scenes:

  • Flux Kontext (run on a service called Replicate) is the model that makes the images. It accepts multiple reference photos, so we hand it yours and your dog’s together.
  • Claude (from Anthropic) does two jobs: before generation, it looks at your uploads and writes a detailed description that becomes part of the prompt. After generation, it grades each panel and tells us what went wrong.
  • Then there’s the un-glamorous part: stitching four panels into a vintage strip with sepia tones, dust specks, faint scratches, and a paper texture. No AI involved — just classical image processing. But it’s a huge part of why the result feels like a real photobooth strip and not “AI output.”

Four panels generate in parallel and the whole thing takes 30–60 seconds.

That last number used to be three to six minutes. Most of what we learned came from getting it down.

We mostly built this by deleting things

Our first version was much fancier than today’s. It generated each panel, then ran a face-swap model to paste in the user’s face, then ran a second model to clean up the face, then ran another model to redo the head region in higher fidelity. Each step was a real quality win.

It also took six minutes per strip. And the face-swap step left a faint vertical seam down every panel that you couldn’t quite unsee once you noticed it.

So we started cutting. Face swap, gone. Face restoration, gone. Head re-do, gone. A clever “regenerate failed panels automatically” loop that occasionally made strips take ten minutes — gone.

Each deletion felt like a quality regression. Each one made the product better.

Three minutes feels like the app froze. Sixty seconds feels like a magic trick. We deleted things until it felt like a magic trick.

The thing nobody tells you

About six months in, our most useful realization was that this isn’t really a model problem. Swapping AI models would change which mistakes we make, not the difficulty of the project. The actual engineering went into everything around the model: the photo analysis up front, the parallelization, the post-processing, the way users recover when a panel is off.

A consumer AI product also has to be honest about being wrong sometimes. The whole post-generation flow — tap a single panel to regenerate it, try a different theme for free, discard the result with no charge — exists because the first output is sometimes off. The model is allowed to fail. The product has to be okay with that.

The actual metric isn’t pixels. It’s recognition.

A technically impressive panel where the user doesn’t quite see themselves is a failure. A slightly fuzzy panel that captures the exact crinkle of their dog’s eyebrows is a win.

My friend showed me a strip last week, from when her dad visited in the spring. The third panel caught the dog looking up at him with one ear cocked — the exact way the dog looks when he hears the back door open.

“That’s how he looks when he hears the door,” she said.

That’s the product. Everything else is in service of that one moment.

Originally published on HelixAI · May 15, 2026