# AI photo pipeline for theme-park souvenir photos

> Mirai Minds builds AI and platform services for Photo Experience, which runs souvenir photography at theme parks, attractions and events. We built generative photo editing with an identity check, face search, the image-processing pipeline, device sign-in and image delivery. The generative stage takes about 18.7 seconds per photo on one NVIDIA A100, down from 60–90 seconds.

Client: Photo Experience (Souvenir photography for theme parks and events) · Status: pilot · Year: 2023–present · Written by Sagar Davara

## Results

- **Per generative photo on one NVIDIA A100 80GB:** 18.7 s (before: 60–90 s) — source: Benchmark, 7 scenarios, Sep 2026; before = previous model version
- **Throughput with two GPU workers:** 3.36 images/min — source: Load test at 2–8 simultaneous requests, Sep 2026
- **Quality and load-test requests completed:** 22/22 — source: Benchmark report, Sep 2026
- **Background removal per photo, warm CPU worker:** 2.9 s — source: Staging run, one 800×600 photo, Sep 2026

## Key facts

- **Generative model:** Qwen-Image-Edit-2511 with an 8-step Lightning adapter
- **GPU:** One NVIDIA A100 80GB, two workers, bounded queue
- **Identity check:** SFace cosine similarity of at least 0.65
- **Face search:** YOLOv11 face detection, ArcFace embeddings, FAISS
- **Background removal:** MattingRefine deep-learning matting
- **Device security:** OAuth 2.0 credentials per device, tokens verified locally


## The problem

Photo Experience runs the souvenir-photo business of a venue: cameras at rides, booths and studios, an AI enhancement pipeline, guest galleries, sales kiosks and printing, all managed from one console. The platform predates us. The client wanted new AI stages, generative themed photos and face search, plus sturdier plumbing between capture and delivery, without slowing a venue's queue.

The generative stage was the hard part. The previous model version took 60–90 seconds per image, drifted faces and sometimes added people who weren't there. A venue can't sell that.

## What we built

- **Generative editing service.** Qwen-Image-Edit-2511 with an 8-step Lightning acceleration adapter, served from one NVIDIA A100 80GB with two resident workers and a bounded queue. The guest photo, an optional destination scene and a face close-up (when the face is small) go in together, and an SFace identity check runs before anything is returned.
- **Face search.** YOLOv11 finds faces in each new photo, InsightFace and ArcFace turn them into embeddings, and a FAISS index matches them to known people.
- **Image pipeline.** A Redis-driven flow runs background removal, generative effects, and venue templates and watermarks in order, tracking each stage and retrying failures.
- **Device sign-in and delivery.** An OAuth 2.0 server gives every kiosk and capture station its own credentials, and services check its signed tokens locally, so a device can be cut off from the admin console. An image-delivery API sends each device the photos it hasn't downloaded yet and records every acknowledgment.
- **Console and booth screens.** Work on the admin console and the self-service booth interface.

## How it works in production

A photo from a capture station enters the queue and loses its background through deep-learning matting, about 2.9 s on a warm CPU worker in staging. If the generative stage is switched on for that capture, it goes to the GPU service; blank frames are rejected before they reach the model. In the September 2026 benchmark the service averaged 18.7 s per image across seven scenarios, held 3.36 images per minute with 2 to 8 simultaneous requests, and completed all 22 test requests. Outputs below the identity threshold return an error instead of a photo, and a review mode lets operators inspect them. Branded and watermarked results flow to galleries, kiosks and printers.

## What we'd change

- **Expressions still fail the identity check.** A "frightened on a roller coaster, eyes closed" prompt was rejected at 0.65. Lowering the threshold would hide the problem, not fix it; the next step is a stronger face reference or an identity-safe refinement pass.
- **The test set is too narrow.** Seven scenarios with one reference person reproduce quality; they don't prove it. We want diverse faces, ages and harder poses before calling identity solved.
- **Timeouts should come from measurements.** Early on, a cold background-removal worker took longer than its caller's timeout. Stage timeouts now come from measured cold starts, which is where they should have started.

## Stack

- **Generation:** Qwen-Image-Edit-2511, LightX2V Lightning adapter, Diffusers, PyTorch
- **Faces:** YuNet, SFace, YOLOv11 face, InsightFace, ArcFace, FAISS
- **Pipeline:** Redis job queue, Flask, FastAPI, MySQL, Alembic
- **Security:** OAuth 2.0 client credentials, RS256 JWT, JWKS
- **Infrastructure:** NVIDIA A100 80GB, Docker

## Frequently asked questions

### Why reject a generated photo instead of fixing the face?

Because a souvenir photo of someone who isn't quite you is worse than no photo. The service compares the output face with the original using SFace and returns an error when similarity is under 0.65. We don't mask, restore or lower the threshold to push borderline images through; the check accepts or rejects the model's output without changing its pixels.

### Why two GPU workers and not four?

We tested four on the same A100. One of four requests ran out of GPU memory, and throughput fell to 2.43 images per minute, against 3.29 with two workers in the same test. Two workers with a bounded queue handled 8 simultaneous requests without errors.

### Does face search need the internet?

No. It runs on CPU in Docker with its detection and recognition models built into the image, so it works on a machine without internet access. Its face index has to be kept and backed up: losing it would give every known face a new ID.

### How did you choose the image model?

License first, then quality, then speed. Several strong image-editing models were ruled out because their weights are licensed for non-commercial use only, or because they are only available as an API. Qwen-Image-Edit-2511 with an 8-step acceleration adapter passed all three.


---

Canonical: https://www.miraiminds.co/work/photo-experience-ai-photo-platform
Last updated: 2026-09-23
Publisher: Mirai Minds LLP, 906 Sarthana Business Hub, Nana Varachha, Surat, Gujarat 395013, India. hello@miraiminds.co
