# Computer vision for photos, products and faces

> Mirai Minds builds computer vision systems that run in production: face search across event photos, background removal, generative photo editing with an identity check, and visual product search from a phone photo. We choose models by license and measured quality, serve them on GPUs or CPUs, and publish latency numbers with the method next to them.

## Key facts

- **Generative photo latency:** 18.7 s per image on one NVIDIA A100 80GB
- **Identity check:** SFace similarity of at least 0.65, or the image is rejected
- **Face search:** YOLOv11 face detection, ArcFace embeddings, FAISS
- **Licenses:** Commercial-use terms checked before any benchmark
- **Serving:** GPUs for generation; CPUs for detection and matching


## What we build

### Generative photo editing that keeps the face

- Problem: Guests want themed souvenir photos, but image models drift faces, and the previous model version took 60–90 s per image.
- System: Qwen-Image-Edit-2511 with an 8-step Lightning adapter on one NVIDIA A100 80GB, two GPU workers and a bounded queue. SFace compares the output face with the original and rejects anything under 0.65 cosine similarity.
- Result: 18.7 s per image in a seven-scenario benchmark and 3.36 images per minute under load. Outputs that fail the identity check are rejected, not delivered.

### Face search across event photos

- Problem: Guests shouldn't have to scroll through a whole day's photos to find their own.
- System: YOLOv11 detects faces, InsightFace and ArcFace turn each face into an embedding, and a FAISS index matches it to known people. It runs on CPU in Docker with its models built into the image, so it needs no internet.
- Result: Photos are tagged by person as they arrive.

### Background removal on every photo

- Problem: Green screens are impractical at rides and booths.
- System: A deep-learning matting model cuts the guest out with hair-level edges as the first stage of the photo pipeline, with queued jobs and retries.
- Result: About 2.9 s per photo on a warm CPU worker in staging.

### Visual product search from a photo or a reel

- Problem: Store staff need to find a product from a photo, or from a marketing video playing on a screen.
- System: An image-embedding index scores catalog photos and frames from marketing reels. Reel matches map to the products tagged in that reel, and results are de-duplicated to the top 10.
- Result: Live for an ethnic-wear brand's staff since July 2026.

## What do we build with computer vision?

Systems that sit in a real workflow, with a queue of people waiting on the result. For [Photo Experience](/work/photo-experience-ai-photo-platform), which runs souvenir photography at theme parks and events, we built generative photo editing, face search and the image pipeline around background removal. For an [ethnic-wear e-commerce brand](/work/ecommerce-visual-search-and-cart-recovery) we built visual product search that works on photos of products and on photos of video reels playing on a screen.

## How do you pick a model?

In three passes. First the license: several strong image-editing models were out for a commercial photo business because their weights are non-commercial. Then quality on the client's own images, not on demo pictures. Then speed and memory on the hardware the client will actually pay for.

The last pass often decides things that benchmarks online don't. On one A100 we tested four GPU workers for the generative service: one of four requests ran out of memory and throughput fell to 2.43 images per minute, against 3.29 with two workers in the same test. We shipped two.

## How do you measure speed and quality?

With the method written next to the number. The generative service was measured on the GPU it is deployed on, across seven scenarios: 18.7 s per image on average, 3.36 images per minute with 2 to 8 simultaneous requests, and all 22 quality and load requests completed. With that sample size, the "P95" is simply the slowest request, and we say so.

Quality gets the same treatment. The identity check is a gate, not a guarantee; seven scenarios with one reference person are not enough to call identity solved, so the next test set adds more faces, ages and harder poses. When an edit fails the check, the service returns an error, and no face is quietly repaired to pass.

## Where does it run?

Wherever the work is. Generative editing needs a data-center GPU; we run it with the model resident in memory, two workers and a bounded queue, so there is no per-request model loading. Face search runs on CPU in Docker on a machine without internet. Visual product search runs in a separate serverless service, kept apart from the store's production database. If the images arrive through a chat, the [catalog assistant](/work/resin-catalogue-assistant) shows the lighter pattern: a vision model describes the photo, then search takes over.

## Typical stack

- **Faces:** YOLOv11 face, InsightFace, ArcFace, YuNet, SFace
- **Generation:** Qwen-Image-Edit-2511, LightX2V Lightning adapter, Diffusers
- **Segmentation:** MattingRefine background matting
- **Search:** FAISS, Image-embedding index on AWS Lambda
- **Serving:** PyTorch, FastAPI, NVIDIA A100 80GB, Docker CPU builds
- **Vision-language:** OpenAI vision models

## Related work

- [Catalog assistant with photo search for a resin manufacturer](https://www.miraiminds.co/work/resin-catalogue-assistant.md) — Pilot on a 15-product catalog sample
- [AI photo pipeline for theme-park souvenir photos](https://www.miraiminds.co/work/photo-experience-ai-photo-platform.md) — 18.7 s per generative photo on one A100 (was 60–90 s)
- [Visual product search and cart-recovery calls for an ethnic-wear brand](https://www.miraiminds.co/work/ecommerce-visual-search-and-cart-recovery.md) — Visual search live for store staff since July 2026

## Frequently asked questions

### Can we use any open model commercially?

No, check the license first. For one photo project we ruled out several strong image-editing models because their weights don't allow commercial use, and others because they were only available as an API. We check licenses before we benchmark, so nobody gets attached to a model you can't ship.

### How fast is generative image editing?

It depends on the model, steps, resolution and GPU. In our September 2026 benchmark, the photo service averaged 18.7 s per 1088×1632 image on one NVIDIA A100 80GB and held about 3.36 images per minute with two workers. The previous model version took 60–90 s per image.

### How do you stop a generated photo from changing someone's face?

An identity check. SFace compares the face in the output with the original, and anything under 0.65 cosine similarity is rejected with an error instead of delivered. Strong expression changes fail more often; we report that honestly rather than lowering the threshold. Operators can inspect rejected images in a separate review mode.

### Does face search need a GPU or the cloud?

No. Our face search runs on CPU in Docker, with its detection and recognition models built into the image, so it can run on a machine without internet access. Generative editing is different: it needs a data-center GPU.

### How do you handle face data?

Face embeddings are numbers rather than photos, but they are still personal data. We keep them inside the client's infrastructure, store only what matching needs, and agree retention with the client before launch. Consent at capture is the operator's responsibility; we can build the consent step into their flow.

### What do you need from us?

Sample images or video from real conditions (lighting, angles, crowds), the hardware you can run on, the throughput you need at peak, and examples of right and wrong results. Studio photos alone make any model look better than it will be in production.


---

Canonical: https://www.miraiminds.co/services/computer-vision
Last updated: 2026-09-23
Publisher: Mirai Minds LLP, 906 Sarthana Business Hub, Nana Varachha, Surat, Gujarat 395013, India. hello@miraiminds.co
