Skip to content
Mirai Minds

Case studyComputer vision

AI photo pipeline for theme-park souvenir photos

By Sagar DavaraPublished

Pilot

In short

Mirai Minds builds AI and platform services for Photo Experience, which runs souvenir photography at theme parks, attractions and events. We built generative photo editing with an identity check, face search, the image-processing pipeline, device sign-in and image delivery. The generative stage takes about 18.7 seconds per photo on one NVIDIA A100, down from 60–90 seconds.

Results

Results

  • 18.7 s

    Per generative photo on one NVIDIA A100 80GB

    Before: 60–90 s

    Source · Benchmark, 7 scenarios, Sep 2026; before = previous model version

  • 3.36 images/min

    Throughput with two GPU workers

    Source · Load test at 2–8 simultaneous requests, Sep 2026

  • 22/22

    Quality and load-test requests completed

    Source · Benchmark report, Sep 2026

  • 2.9 s

    Background removal per photo, warm CPU worker

    Source · Staging run, one 800×600 photo, Sep 2026

How it works

System diagram: AI photo pipeline for theme-park souvenir photosINPUTGuest photo from acapture stationMODELBackground removal(matting)MODELGenerative edit,Qwen-Image-Edit-2511MODELIdentity check,SFace at least 0.65SYSTEMVenue templates andwatermarkOUTPUTGallery, kiosk andprint

How it flows

  1. 01 Guest photo from a capture station → Background removal (matting)
  2. 02 Background removal (matting) → Generative edit, Qwen-Image-Edit-2511
  3. 03 Generative edit, Qwen-Image-Edit-2511 → Identity check, SFace at least 0.65
  4. 04 Identity check, SFace at least 0.65 → Venue templates and watermark
  5. 05 Background removal (matting) → Venue templates and watermark
  6. 06 Venue templates and watermark → Gallery, kiosk and print

The problem

Photo Experience runs the souvenir-photo business of a venue: cameras at rides, booths and studios, an AI enhancement pipeline, guest galleries, sales kiosks and printing, all managed from one console. The platform predates us. The client wanted new AI stages, generative themed photos and face search, plus sturdier plumbing between capture and delivery, without slowing a venue's queue.

The generative stage was the hard part. The previous model version took 60–90 seconds per image, drifted faces and sometimes added people who weren't there. A venue can't sell that.

What we built

  • Generative editing service. Qwen-Image-Edit-2511 with an 8-step Lightning acceleration adapter, served from one NVIDIA A100 80GB with two resident workers and a bounded queue. The guest photo, an optional destination scene and a face close-up (when the face is small) go in together, and an SFace identity check runs before anything is returned.
  • Face search. YOLOv11 finds faces in each new photo, InsightFace and ArcFace turn them into embeddings, and a FAISS index matches them to known people.
  • Image pipeline. A Redis-driven flow runs background removal, generative effects, and venue templates and watermarks in order, tracking each stage and retrying failures.
  • Device sign-in and delivery. An OAuth 2.0 server gives every kiosk and capture station its own credentials, and services check its signed tokens locally, so a device can be cut off from the admin console. An image-delivery API sends each device the photos it hasn't downloaded yet and records every acknowledgment.
  • Console and booth screens. Work on the admin console and the self-service booth interface.

How it works in production

A photo from a capture station enters the queue and loses its background through deep-learning matting, about 2.9 s on a warm CPU worker in staging. If the generative stage is switched on for that capture, it goes to the GPU service; blank frames are rejected before they reach the model. In the September 2026 benchmark the service averaged 18.7 s per image across seven scenarios, held 3.36 images per minute with 2 to 8 simultaneous requests, and completed all 22 test requests. Outputs below the identity threshold return an error instead of a photo, and a review mode lets operators inspect them. Branded and watermarked results flow to galleries, kiosks and printers.

What we'd change

  • Expressions still fail the identity check. A "frightened on a roller coaster, eyes closed" prompt was rejected at 0.65. Lowering the threshold would hide the problem, not fix it; the next step is a stronger face reference or an identity-safe refinement pass.
  • The test set is too narrow. Seven scenarios with one reference person reproduce quality; they don't prove it. We want diverse faces, ages and harder poses before calling identity solved.
  • Timeouts should come from measurements. Early on, a cold background-removal worker took longer than its caller's timeout. Stage timeouts now come from measured cold starts, which is where they should have started.

What it runs on

Generation
Qwen-Image-Edit-2511LightX2V Lightning adapterDiffusersPyTorch
Faces
YuNetSFaceYOLOv11 faceInsightFaceArcFaceFAISS
Pipeline
Redis job queueFlaskFastAPIMySQLAlembic
Security
OAuth 2.0 client credentialsRS256 JWTJWKS
Infrastructure
NVIDIA A100 80GBDocker

Asked on the first call

Why reject a generated photo instead of fixing the face?

Because a souvenir photo of someone who isn't quite you is worse than no photo. The service compares the output face with the original using SFace and returns an error when similarity is under 0.65. We don't mask, restore or lower the threshold to push borderline images through; the check accepts or rejects the model's output without changing its pixels.

Why two GPU workers and not four?

We tested four on the same A100. One of four requests ran out of GPU memory, and throughput fell to 2.43 images per minute, against 3.29 with two workers in the same test. Two workers with a bounded queue handled 8 simultaneous requests without errors.

Does face search need the internet?

No. It runs on CPU in Docker with its detection and recognition models built into the image, so it works on a machine without internet access. Its face index has to be kept and backed up: losing it would give every known face a new ID.

How did you choose the image model?

License first, then quality, then speed. Several strong image-editing models were ruled out because their weights are licensed for non-commercial use only, or because they are only available as an API. Qwen-Image-Edit-2511 with an 8-step acceleration adapter passed all three.

Have a system in mind? Let's scope it.

A 30-minute call with an engineer who has shipped this before. You leave with a plan, a rough timeline and what it would take — whether or not we build it.