Explainer: foundation models are eating classical machine vision
Zero-shot inspection models are replacing hand-tuned vision pipelines that took decades to build. Here's what changes, what doesn't, and what to ask your vendor.
Mei Tanaka — Contributing Analyst
Classical machine vision is a discipline of careful engineering: controlled lighting, calibrated optics, and algorithms tuned per part, per line, per defect. Vision foundation models upend the economics of that craft by learning general visual competence from internet-scale data, then adapting to a new inspection task from a handful of examples.
What actually changes
Deployment time is the headline. Tasks that took a vision integrator six weeks of pipeline tuning now reach acceptable accuracy in days, sometimes from a dozen labelled images. Changeover cost collapses too — retraining for a new product variant becomes a same-day operation rather than a new engineering project.
What doesn't change
Physics. A foundation model cannot see a defect the optics did not resolve or the lighting did not reveal. The unglamorous fundamentals of illumination and lens selection still separate working systems from demos — which is why experienced integrators are thriving, not disappearing, in the transition.
Buyers should press vendors on three points: false-accept rates on their actual parts rather than benchmark sets, on-premise retraining capability, and what happens to accuracy when the line changes lighting, speed or upstream suppliers.
Cite this article
Mei Tanaka. "Explainer: foundation models are eating classical machine vision." Autonomous Systems Review, 10 Jul 2026. https://autonomoussystemsreview.com/articles/explainer-machine-vision-foundation-models.