What do we build with computer vision?
Systems that sit in a real workflow, with a queue of people waiting on the result. For Photo Experience, which runs souvenir photography at theme parks and events, we built generative photo editing, face search and the image pipeline around background removal. For an ethnic-wear e-commerce brand we built visual product search that works on photos of products and on photos of video reels playing on a screen.
How do you pick a model?
In three passes. First the license: several strong image-editing models were out for a commercial photo business because their weights are non-commercial. Then quality on the client's own images, not on demo pictures. Then speed and memory on the hardware the client will actually pay for.
The last pass often decides things that benchmarks online don't. On one A100 we tested four GPU workers for the generative service: one of four requests ran out of memory and throughput fell to 2.43 images per minute, against 3.29 with two workers in the same test. We shipped two.
How do you measure speed and quality?
With the method written next to the number. The generative service was measured on the GPU it is deployed on, across seven scenarios: 18.7 s per image on average, 3.36 images per minute with 2 to 8 simultaneous requests, and all 22 quality and load requests completed. With that sample size, the "P95" is simply the slowest request, and we say so.
Quality gets the same treatment. The identity check is a gate, not a guarantee; seven scenarios with one reference person are not enough to call identity solved, so the next test set adds more faces, ages and harder poses. When an edit fails the check, the service returns an error, and no face is quietly repaired to pass.
Where does it run?
Wherever the work is. Generative editing needs a data-center GPU; we run it with the model resident in memory, two workers and a bounded queue, so there is no per-request model loading. Face search runs on CPU in Docker on a machine without internet. Visual product search runs in a separate serverless service, kept apart from the store's production database. If the images arrive through a chat, the catalog assistant shows the lighter pattern: a vision model describes the photo, then search takes over.