FOXNUT

Definition

Quality inspection: the software decision above the camera

Automating inspection is four decisions above the hardware: what gets looked at, what a defect class is worth, who rules on the borderline unit, and what the record has to show.

By the Foxnut team · Updated

Quality control asks two questions, and only one of them is software

The first question is what the line needs in order to see a part properly: sensors, optics, lighting, fixturing, the cell that presents the unit, the speed it moves at. That is capital equipment, it is quoted by machine-vision integrators, and nothing on this page rules on it. The second question sits above the hardware and is the same on every line whatever equipment is under it. What gets looked at, and how much of the output. What a defect classification means once something is assigned to it. Who rules on the unit that falls between accept and reject, and how fast that ruling has to arrive. And what the record has to be able to show, months later, about a unit that shipped and a unit that did not. Those four are software decisions, they are where automated inspection succeeds or fails as a project, and they are almost never the subject of the proposal a manufacturer receives. A studio can rule on them honestly because they turn on process design, published standards and measured human performance rather than on which camera to buy. That is also why what an inspection system has to leave behind to be trusted is worth settling before any equipment is quoted: the artefacts the four decisions produce are the part of the system a manufacturer still owns after everyone has left.

The four decisions the hardware does not make

Each one changes shape the moment a machine rather than a person is doing the looking, and each one becomes a written policy that somebody has to own.

DecisionWhat changes when a machine does the lookingWhat it costs to get wrong
What gets looked atSampling was a cost-control device. When the marginal cost of an inspection falls to nearly nothing, the coverage question collapses and the statistical risk moves from the lot to the classifierA sampling plan retained out of habit, so the machine inspects everything and the paperwork still describes a lot
What a classification is worthThe class becomes a stored label with a downstream consequence: scrap, rework, quarantine, an investigation. The classes have to be defined and priced before the model is trained on themClasses inherited from a spreadsheet, so the model learns a taxonomy nobody has defended and the reject reason is unusable
Who rules on the borderline unitA confidence score exists and is logged, so the borderline unit is now a routing decision with a threshold, made thousands of times a shiftA threshold set to make the demonstration look good, which either buries a person in referrals or quietly passes the ambiguous units
What the record showsThe evidence stops being an inspector’s sign-off and becomes the model, its test data, its thresholds and its change historyA system that can say a unit was rejected and cannot say what it was rejected for, or what the rule was on the day

The sampling row is the one that surprises buyers. Acceptance sampling exists because looking at things is expensive: a sampling plan is a rule for accepting or rejecting a lot from a count of defects in a sample, and its behaviour is described by an operating characteristic curve, which plots the probability of accepting the lot against the lot fraction defective. The plan is a negotiated split of two risks, the producer’s risk of rejecting a lot at the acceptable quality level and the consumer’s risk of accepting one at the lot tolerance percent defective. Automate the looking and that trade largely dissolves, because there is no longer a reason to inspect a sample. What does not dissolve is the risk itself. It moves inside the classifier, where it is no longer a published curve agreed with a customer but a threshold in a configuration file. In the one industry where 100% inspection has always been mandatory, this is already understood: every filled container of a parenteral product must be inspected individually, so the whole of the remaining argument there is about classification, qualification and records.

What the decision turns on

Six structural dimensions decide whether a process is worth automating. Quality inspection reads unusually on three of them: its exceptions are engineered rather than encountered, its error costs run in opposite directions at once, and its regulatory exposure is either extremely specific or entirely absent depending on what is being inspected.

Quality inspection: process profile
DimensionWhat it reads on quality inspectionSource
Exception varianceThe exception is the borderline unit, and how often it appears is set by a defect rate that is usually very low. Reviewed production figures run from 0.01% for atomic weapon components through 0.2% for coins to 4% for jam tarts, with 1% to 10% described as typical, which means an inspector may examine thousands of parts between defects. Detection gets worse as the rate falls: in the earliest controlled study, 80 inspectors detected significantly fewer defects and raised significantly more false alarms as the defect rate dropped from 16% to 4%, 1% and 0.25%. The variance is also partly a design choice rather than an observation, because the draft AI annex requires the input space to be split into subgroups by characteristics including the types and severity of defects, with acceptance criteria allowed to differ between them.Sandia SAND2012-8590; draft EU GMP Annex 22, sections 3.2 and 4.2
VolumeNot units inspected. A machine looks at everything at no marginal cost, so what scales is the number of outputs a person still has to rule on: false rejects, plus every unit the model declines to call. The human baseline shows the shape. Eighty-two qualified inspectors examining 140 parts for eight defect types correctly rejected 85% of defective items and incorrectly rejected 35% of acceptable ones, against an industry average detection level of about 80%. As arithmetic rather than a finding, transplanting those two rates onto a line where one unit in a hundred is defective would flag roughly a third of the output for a second look, in a study whose own defect rate was 30%. The payback condition follows as a condition rather than a threshold: the automation pays when the adjudication burden it creates is smaller than the inspection burden it removes, and that comparison needs the current false-reject rate, which most lines have never measured.See, Human Factors 57(8), 2015
Cost of an errorAsymmetric and paid in two currencies at once. A false reject destroys good product and is absorbed immediately. A miss leaves the site, and in regulated manufacture it is defined as an event rather than a loss: critical defects should not be identified during any subsequent sampling of containers that have already been accepted, and any critical defect found later triggers an investigation, because it indicates a possible failure of the original inspection process. For anyone selling inspection software rather than running it, the exposure moved in 2024: the recast European product liability rules make software a product in its own right and list a product's ability to continue to learn after being placed on the market among the circumstances relevant to whether it is defective, applying to products placed on the market from 9 December 2026.EU GMP Annex 1, paragraph 8.30; Directive (EU) 2024/2853, Articles 4(1) and 7(2)
ReversibilityThe unit is one-directional and the system is not, which is the reverse of the usual reading. A rejected part can be re-examined and reinstated; a shipped part can only be recovered through the market. But the system's own settings are the thing that is hard to reverse quietly, because a tested model, the system it runs in and the whole process it automates go under change control before deployment, any change to the model, the system or even the physical objects it reads has to be documented and assessed for whether a retest is needed, and a decision not to retest has to be fully justified. Adjusting a threshold on a Friday afternoon is a change to a validated process, not a tuning session.Draft EU GMP Annex 22, sections 10.1 and 10.2
Regulatory exposureBimodal, and the two modes have almost nothing in common. In sterile manufacture the inspection step is named by process: 100% individual inspection, defect classification and criticality determined at qualification and based on risk, a maintained defect library used to train personnel, annual qualification of manual inspectors against that library, automated methods validated to be equal to or better than manual inspection and challenged with representative defects before start-up and at intervals through the batch, and results trended with market impact assessed on an adverse trend. Medical devices in the United States moved to the same family of expectations when the amended quality system rule incorporating ISO 13485:2016 took effect on 2 February 2026. Outside regulated manufacture, no rule names the inspection step at all, and the exposure is contractual and lands in the product liability regime instead. Note where the exposure does not come from: the European AI Act's own list of high-risk uses runs from biometrics and critical infrastructure through employment, creditworthiness and law enforcement, and manufacturing quality control is not on it, so the binding requirements here are the sector's own.EU GMP Annex 1, paragraphs 8.30 to 8.33; FDA final rule 2024-01709; Regulation (EU) 2024/1689, Annex III
Vendor market maturityMature and crowded where the equipment is, and structurally unable to supply the part that costs. No public measure of concentration in inspection software was retrieved, so none is quoted here. What is documented is the allocation of work: the validation obligation sits with the user, who must show the automated method is equal to or better than the manual one, and the documentation of training, validation and testing must be available and reviewed by the regulated user whether the model was built in-house or supplied by a vendor. The draft annex goes further and requires that staff who have had access to the test data are kept out of training the same model, or work in pairs with a colleague who has not. A demonstration cannot discharge any of that, which is why the buyable part of this category is the smallest part of the project.EU GMP Annex 1, paragraph 8.32; draft EU GMP Annex 22, sections 2.2 and 6.5

Two of those rows carry the argument. The volume row says the automation does not remove the human decision, it relocates and multiplies it: the machine converts an inspection workload into an adjudication workload, and whether that trade is good depends on a number about the current process that almost nobody holds. The maturity row says the work that decides the outcome is not for sale. A vendor can supply detection; the intended use, the subgroup definitions, the acceptance criteria, the labelled test set and its independence, and the review of what the model attended to are all obligations of the operator, and they are exactly the artefacts that make the system defensible after handover.

The regulatory row is the quiet one and it sets the design. Where the process is named by a standard, the design is largely dictated and the project is an evidence exercise with a detector inside it. Where nothing names it, the same structure is still the right one, because the record is what turns a rejection into something a customer, an insurer or a court can be shown. The absence of a rule is not the absence of a requirement; it just means nobody will tell the manufacturer what the record has to contain until it is needed.

What this usually gets wrong

The first error is treating a detection score as the deliverable. The draft annex sets the standard bluntly: the acceptance criteria for a model should be at least as high as the performance of the process it replaces, which implies that the performance of the replaced process must be known. Almost no line knows it. There is no measured hit rate for the current inspectors, no false-reject rate, and no agreement study, so a vendor’s number is being compared against a number that does not exist. The single cheapest thing a manufacturer can do before buying anything is to construct a labelled set of its own parts, run its own inspectors against it blind, and find out what the incumbent process actually achieves. That set is also the asset the project will need later, because a test set has to be representative of the full sample space, stratified across subgroups, and verified to a very high degree of label correctness.

The second error is trusting the labels. In the study of 82 inspectors, the accept-or-reject decision was correct 85% of the time, but the rate at which the inspector identified the exact defect present was 35%, meaning correct rejections were frequently reached for the wrong reasons. A model trained on labels of that quality learns a class structure that does not exist, and the failure is invisible in headline accuracy because the accept-or-reject column still looks right. This is why the standard asks for labelling verified through independent review by multiple experts, validated equipment or laboratory tests, rather than for more data.

The third error is treating a confidence threshold as the answer to the borderline case. The draft annex is right that a model should log a confidence score per outcome and should have a threshold below which it declines and flags the outcome as undecided rather than making an unreliable call. What a threshold does not do is decide who then rules, in what time, with what authority. The human analogue is a warning: when inspectors were asked to rate confidence and the low-confidence parts were re-inspected, the phased approach was neither effective nor efficient at improving overall accuracy. A confidence threshold is a routing policy that creates a queue, and a queue with no owner and no service time is where automated inspection projects quietly fail.

The fourth error is filing retraining under maintenance. A retrained model is a change to a process under change control, and so is a change to the system around it or to the physical objects it reads. Performance has to be monitored against the metrics it was accepted on, and the input data has to be monitored for drift out of the sample space the model was tested against. The draft annex gives the example of a lighting condition deteriorating, which is worth noticing: this is the one point where the hardware comes back into the software conversation, and it comes back as a variable the system must watch rather than as a purchase. The three-line implication for a maintenance plan is that somebody owns the metric, somebody owns the drift check, and any change to either the model or the cell reopens the question of whether the acceptance test still stands.

There is one prior question this page does not answer and it belongs before all of the above. Whether the historical inspection records, the defect taxonomy and the images or measurements are complete, consistently coded and current enough for any of this is a readiness question with its own answer, and a classifier trained on records that only ever say pass or fail will produce a system that can also only say pass or fail.

The verdict

The evidence supports building the decision layer first and buying the detection last. The four decisions above the hardware carry the project: what gets looked at, what each class is worth, who rules on the borderline unit and inside what time, and what the record has to show. Where those exist in writing, an inspection system is a straightforward build and the equipment question is a procurement exercise. Where they do not, buying equipment first produces a machine that generates thousands of unowned decisions a shift and a record that cannot explain any of them.

The reason is structural rather than a comment on how good detection currently is. Three of the six dimensions point the same way. The volume row says the machine converts inspection work into adjudication work, so the return depends on a false-reject rate the manufacturer has probably never measured. The maturity row says the artefacts that make the system defensible are obligations of the operator and cannot be bought from the vendor supplying the detector. And the reversibility row says the settings are under change control the moment the system is validated, so the thresholds and classes chosen at the start are expensive to revisit and deserve more argument than they usually get.

For a team weighing this, the useful first move is a measurement rather than a purchase. Take a few hundred of the plant’s own parts, establish ground truth on them properly, and run the current inspectors against them blind. The output is three numbers nobody has: what share of defects the process catches, what share of good product it destroys, and how often two people agree on which defect is present. Those numbers decide whether automation is worth anything here, they are the acceptance criteria the standard already requires, and the labelled set built to produce them is the test data any later project would have had to build anyway. Foxnut Studios builds systems of this kind and hands them over, which is a reason to be direct about the boundary: the camera, the lighting and the cell are capital equipment questions for an integrator, and no answer on this page should be read as advice on any of them.

If you have run the blind test above and have real numbers in hand, bring them to us and we will help you read what they mean for the decision layer.

Sources

  1. European Commission, EudraLex Volume 4, Annex 1 'Manufacture of Sterile Medicinal Products', C(2022) 5938 final of 22 August 2022, in operation from 25 August 2023 - paragraphs 8.30 to 8.33 on individual inspection of all filled containers, defect classification and criticality set at qualification, the defect library, annual qualification of manual inspectors, validation and periodic challenge of automated inspection, and trending of reject levels Retrieved
  2. European Commission, EudraLex Volume 4, draft Annex 22 'Artificial Intelligence', released for stakeholder consultation 7 July 2025 - the scope exclusions, and sections 3 (intended use and subgroups), 4 (test metrics, acceptance criteria, no decrease against the process replaced), 5 and 6 (test data selection, labelling and independence), 8 (feature attribution), 9 (confidence score and threshold) and 10 (change control, performance and input drift monitoring, human review) Retrieved
  3. Judi E. See, 'Visual Inspection: A Review of the Literature', Sandia National Laboratories, SAND2012-8590, October 2012 - typical production defect rates, the effect of defect rate on detection and false alarms, and the documented pressure on inspectors toward accepting borderline items. Read as text extracted from the report PDF Retrieved
  4. Judi E. See, 'Visual Inspection Reliability for Precision Manufactured Parts', Human Factors 57(8), 2015 - eighty-two qualified inspectors, 140 parts, eight defect types: 85% of defective items correctly rejected, 35% of acceptable parts incorrectly rejected, 77% valid hits against 35% exact hits, and the failure of confidence-guided re-inspection. Read in the accepted manuscript hosted by OSTI Retrieved
  5. NIST/SEMATECH e-Handbook of Statistical Methods, section 6.2.2 'What kinds of Lot Acceptance Sampling Plans (LASPs) are there?' - the definitions of acceptable quality level, lot tolerance percent defective, producer's risk and consumer's risk, and what an operating characteristic curve plots Retrieved
  6. Directive (EU) 2024/2853 of 23 October 2024 on liability for defective products, Official Journal 18 November 2024 - Article 4(1) including software in the definition of a product, Article 7(2) on the circumstances relevant to defectiveness including a product's ability to continue to learn after being placed on the market, and recital 63 on non-application to products placed on the market before 9 December 2026. Read in the Official Journal full text on EUR-Lex Retrieved
  7. US Food and Drug Administration, 'Medical Devices; Quality System Regulation Amendments', final rule, Federal Register 2 February 2024, document 2024-01709 - the amendment of 21 CFR part 820 primarily by incorporating by reference ISO 13485:2016, effective 2 February 2026. Read in the Federal Register text hosted on govinfo.gov Retrieved
  8. Regulation (EU) 2024/1689 (the AI Act), Official Journal 12 July 2024 - Annex III, the list of high-risk AI systems referred to in Article 6(2), read in full to establish that manufacturing quality control is not among the listed areas. Read in the Official Journal full text on EUR-Lex Retrieved

Foxnut Studios works on briefs like this one from Bengaluru and Paris. If you want the shape of that before you talk to anyone, here is how an AI engagement is scoped and priced.