How We Built an AI Vehicle Inspection System From 10 Photos
A production-minded approach to vehicle condition capture: ten photos, a constrained vision pipeline, and the booking workflow a dealership can actually run.
· Updated · 4 min read · AI
Dealership operations do not fail because someone forgot a neural network. They fail because condition is still a WhatsApp album, a clipboard, and an argument after the car has left the lot.
ZarkiTech Insights publishes from that kind of work. We have already shipped Carovenue — a vehicle marketplace and dealership operations platform — so inspection is not a thought experiment. It is the missing layer between listings, bookings, and the floor.
This is how we would (and do) design an inspection system that starts from ten photos.
The product constraint
Ten photos is a product decision, not a model decision.
A porter or salesperson will not walk a 40-point shoot list at 6pm. They will take a front three-quarter, a rear three-quarter, both sides, the dash, the odometer, the interior, the tyres, and two damage close-ups. If the system needs more than that, it will be bypassed.
So the contract is:
- a fixed capture order
- a live completeness check
- no “upload later”
- a result a human can override in under a minute
The model is allowed to be wrong. The workflow is not allowed to be vague.
What the ten frames are for
We treat the set as a structured form, not a gallery.
- Identity and listing match (plate, trim cues, colour)
- Exterior panels (left, right, front, rear)
- Glass and lights
- Wheels and tyre condition
- Interior / odometer
- Directed close-ups for damage the human already noticed
Computer vision here is closer to document capture than to autonomous driving. The job is to classify, locate, and describe — then write a record the booking system can store.
Pipeline, not a demo
A demo overlays boxes on a pretty car. A production inspection system has to survive fluorescent lighting, wet paint, and a phone held in a hurry.
The pipeline we use looks like this:
- Capture quality gate — blur, exposure, and “is this even a car panel?”
- View classification — which of the ten slots did we actually get?
- Damage / cleanliness / missing-part hints — localised, with confidence
- Human confirm — the operator ticks, edits, or rejects
- Immutable record — photos + labels + who confirmed + timestamp, attached to the vehicle and the booking
If step 1 fails, we do not invent damage. We ask for a retake. That single rule removes most of the embarrassing false positives.
Why we keep the model small
Large general vision models are useful for research and for odd edge cases. They are a poor default for a lot workflow.
We prefer a constrained stack:
- on-device or edge prechecks so the porter is not waiting on a round trip
- a specialist damage/view model for the common classes
- a slower fallback only when confidence is low
- every accepted label stored with model version
That last point matters more than architecture diagrams. Six months later someone will ask why a bumper was marked “scratched”. If you cannot answer with a model version and a photo, you do not have an inspection system. You have a slideshow.
Where it sits in the product
Inspection is useless as a standalone app. It has to write into the same record as listings and bookings.
On a dealership floor that means:
- the listing cannot go live until the set is complete or explicitly waived
- a booking handover can require a matching outbound inspection
- inbound vs outbound diffs become the dispute file
This is the same instinct as the rest of our AI product work: intelligence has to earn its keep inside an existing operation. The vehicle marketplace already has the entities. Inspection is a new event on those entities.
What we refused to build
- A public “AI score” on the listing. Buyers do not trust a mystery number, and sellers will game it.
- Fully automatic payouts from detected damage. That is a legal process, not a softmax.
- Open-ended chat (“tell me about this car”) as the primary UI. The operator needs a checklist.
We would rather ship a boring, complete capture flow than a clever assistant that cannot be audited.
What to steal from this
If you are adding vision to an operations product:
- freeze the capture set before you pick a model
- fail closed on quality, not on damage
- make a human the publisher of the record
- version the model next to the photo
- attach the result to a business object that already exists
That is how ten photos become an inspection system instead of a folder of JPEGs.
Further reading: Building scalable booking platforms and how we design production NestJS backends. If you are scoping a similar operations product, ask for a build plan.