Case studyComputer vision
AI photo pipeline for theme-park souvenir photos
By Sagar DavaraPublished
PilotIn short
Mirai Minds builds AI and platform services for Photo Experience, which runs souvenir photography at theme parks, attractions and events. We built generative photo editing with an identity check, face search, the image-processing pipeline, device sign-in and image delivery. The generative stage takes about 18.7 seconds per photo on one NVIDIA A100, down from 60–90 seconds.
Results
Results
18.7 s
Per generative photo on one NVIDIA A100 80GB
Before: 60–90 s
Source · Benchmark, 7 scenarios, Sep 2026; before = previous model version
3.36 images/min
Throughput with two GPU workers
Source · Load test at 2–8 simultaneous requests, Sep 2026
22/22
Quality and load-test requests completed
Source · Benchmark report, Sep 2026
2.9 s
Background removal per photo, warm CPU worker
Source · Staging run, one 800×600 photo, Sep 2026
The system
How it works
How it flows
- 01 Guest photo from a capture station → Background removal (matting)
- 02 Background removal (matting) → Generative edit, Qwen-Image-Edit-2511
- 03 Generative edit, Qwen-Image-Edit-2511 → Identity check, SFace at least 0.65
- 04 Identity check, SFace at least 0.65 → Venue templates and watermark
- 05 Background removal (matting) → Venue templates and watermark
- 06 Venue templates and watermark → Gallery, kiosk and print
The story
The problem
Photo Experience runs the souvenir-photo business of a venue: cameras at rides, booths and studios, an AI enhancement pipeline, guest galleries, sales kiosks and printing, all managed from one console. The platform predates us. The client wanted new AI stages, generative themed photos and face search, plus sturdier plumbing between capture and delivery, without slowing a venue's queue.
The generative stage was the hard part. The previous model version took 60–90 seconds per image, drifted faces and sometimes added people who weren't there. A venue can't sell that.
What we built
- Generative editing service. Qwen-Image-Edit-2511 with an 8-step Lightning acceleration adapter, served from one NVIDIA A100 80GB with two resident workers and a bounded queue. The guest photo, an optional destination scene and a face close-up (when the face is small) go in together, and an SFace identity check runs before anything is returned.
- Face search. YOLOv11 finds faces in each new photo, InsightFace and ArcFace turn them into embeddings, and a FAISS index matches them to known people.
- Image pipeline. A Redis-driven flow runs background removal, generative effects, and venue templates and watermarks in order, tracking each stage and retrying failures.
- Device sign-in and delivery. An OAuth 2.0 server gives every kiosk and capture station its own credentials, and services check its signed tokens locally, so a device can be cut off from the admin console. An image-delivery API sends each device the photos it hasn't downloaded yet and records every acknowledgment.
- Console and booth screens. Work on the admin console and the self-service booth interface.
How it works in production
A photo from a capture station enters the queue and loses its background through deep-learning matting, about 2.9 s on a warm CPU worker in staging. If the generative stage is switched on for that capture, it goes to the GPU service; blank frames are rejected before they reach the model. In the September 2026 benchmark the service averaged 18.7 s per image across seven scenarios, held 3.36 images per minute with 2 to 8 simultaneous requests, and completed all 22 test requests. Outputs below the identity threshold return an error instead of a photo, and a review mode lets operators inspect them. Branded and watermarked results flow to galleries, kiosks and printers.
What we'd change
- Expressions still fail the identity check. A "frightened on a roller coaster, eyes closed" prompt was rejected at 0.65. Lowering the threshold would hide the problem, not fix it; the next step is a stronger face reference or an identity-safe refinement pass.
- The test set is too narrow. Seven scenarios with one reference person reproduce quality; they don't prove it. We want diverse faces, ages and harder poses before calling identity solved.
- Timeouts should come from measurements. Early on, a cold background-removal worker took longer than its caller's timeout. Stage timeouts now come from measured cold starts, which is where they should have started.
Stack
What it runs on
- Generation
- Qwen-Image-Edit-2511LightX2V Lightning adapterDiffusersPyTorch
- Faces
- YuNetSFaceYOLOv11 faceInsightFaceArcFaceFAISS
- Pipeline
- Redis job queueFlaskFastAPIMySQLAlembic
- Security
- OAuth 2.0 client credentialsRS256 JWTJWKS
- Infrastructure
- NVIDIA A100 80GBDocker