The idea that data augmentation is necessary for such a simple grasp demonstration, like that of the Festo Bionic Handling Assistant with a small red ball, strikes me as questionable. If we hadn't already invested in this approach, would we really think so much complexity is needed for this? You can calibrate the detection of the red ball with just a few examples for demonstration, without resorting to thousands of artificially augmented images, especially if the environment is controlled. It's like insisting on a super complicated bus route just because you've already studied it, while a direct taxi would do the job without all the detours. Sometimes, you need to let go of it and look for the most straightforward solution.