The useful figure is pixels across the smallest feature that must be detected, not megapixels. A defect needs several pixels across it to be distinguishable from noise — three at absolute minimum, and five or more for anything that has to be measured rather than merely found.
So the calculation runs backwards from the defect: smallest feature size, times the pixels needed across it, divided into the field of view. That gives the sensor resolution. Choosing a camera first and hoping the defect resolves is the most common and most expensive sequence error in this category.
Field of view should be as tight as the part tolerances allow. Every millimetre of extra view spends resolution on space where a defect cannot occur, and part positioning variation is what forces the margin. Better fixturing frequently buys more than a bigger sensor.