If we hadn't already invested so much in the idea that AI would solve everything, would we consider training a model as the essential condition for a robotic hand to hold an apple without crushing it?
The real challenge lies in sensor mechanics and actuator precision, far more than in visual recognition alone; if the hand doesn't "feel" the apple correctly, it will crush it.
Without proper engineering for graded force and tactile feedback, even the best AI cannot prevent a catastrophe.
Training AI is a facilitator, yes, but not a sufficient guarantee without these other pieces of the puzzle.
A robot can recognize an apple, but if its fingers are too rigid or its pressure sensors only have an 'on/off' mode, the apple will end up as compote.