
Synthetic data and robotic simulation to lift vision and robotics to completely new levels
4 June 2025
While many of you might consider synthetic data as the next big thing as part of the AI ‘revolution’, the team at WUR Vision + Robotics sees synthetic data as just another essential part of vision and robotics. Researchers Trim Bresilla and Arjan Vroegop explain why and how.
Deep learning as an accelerator
Vroegop kicks off by going back in time. “I recently saw a video of a robotic cucumber harvester developed 25 years ago by Wageningen researchers. With a robotic arm and end effector operating based on classical robot programming. Artificial intelligence (AI) was not yet available in those days. Robotic harvesting was, and still is, very challenging and such solutions aren’t available commercially yet. Many existing robots have difficulties to perform well in the real agrifood world. In the agrifood domain, the main challenge lies in the natural variation of animals, food and plants. Basically in all areas vision and robotics operate. These are different circumstances compared to a conditioned (industrial) environment such as factories and warehouses where standardised objects are placed at predetermined locations and positions. Let alone the dust, humidity, sunlight, shadow, wind, precipitation, et cetera in outdoor and greenhouse circumstances.”
“Since around 2016, deep learning arose as a new component in our toolbox and we’ve seen the technology and applications picking up really fast”, says Trim Bresilla. Vroegop: “Deep learning enables us to make algorithms more robust and it makes it possible to detect harvestable cucumbers in an all-green environment on both cloudy and sunny days for instance. It also enables us to find weeds in open field crops and to detect and classify fish on a highspeed conveyor belt. This is all done by vision systems incorporating AI. Robot actions are often still programmed the classic way, like 25 years ago. My prediction is that AI based robotic control will be the next major development and we in Wageningen are more than ready for that.”

A 3D environment also enables physical experiments, such as conveyor drops or robot interaction. A conveyor, for example, can run infinitely with random batches of fish, resulting in an endless stream of high-quality, annotated training data.
Synthetic data to lift vision and robotics
Vroegop is convinced that the combination of using both synthetic data and robotic simulation will speed up and improve both research as well as applications of vision and robotics. “We already use artificial images to train neural networks. With synthetic data we can recreate complete fields with plants, crops, weed and trees and orchards/vineyards. As well as complete greenhouses that can hardly be distinguished from real situations. We not only use these environments for synthetic data to train vision AI models, but also for robot simulation. I deliberately use the term simulation instead of digital twin, as I think digital twin is a real buzz word which is often used in the wrong way. In Wageningen, we have a lot of domain knowledge, models and datasets, covering the whole agrifood domain. With this, we can make virtual worlds which we use for robot simulation. This way, we can lift the development of robots and vision systems in the agrifood domain to a higher level. A good example is the virtual fish factory, a simulation with realistic 3D fish to provide a building block for fishing industry automation.”

A key aspect if realistic 3D simulations is having realistic 3D models. On the left, the 3D scanning setup developed in-house by Vision+Robotics researchers in Wageningen, can create accurate 3D models for a wide range of subjects, such as fish or plants. The example in the middle shows a virtual fish, created from 3D models, which allows full control over position, composition, and lighting. Finally, on the right, a batch of 3D fish on a conveyor belt simulation.
Improve robots with simulation
Bresilla: “Simulation tools are perfect to teach and improve (humanoid) robots. Raising and teaching them shows parallels with human babies growing up and learning what behaviour is preferred or not by means of rewards and punishments. With robotic simulation, we can do the same thing. Besides, damages or breakages because of unwanted robotic behaviour are not an issue in a simulated world. You can reset and rewind any time you want and you can test different software versions in the same environment and situation over and over again. A key element is that you can recreate your solution and put that in the same simulation.”
Vroegop explains by raising two examples. “If you want to harvest fruits or vegetables for instance, in the real world it is very time consuming to test your prototype end effector, to redesign it and test again. New varieties and cropping strategies are also being introduced and once that is the case, your environment is suddenly different. The harvesting season of both apples and asparagus for instance, is very short. By using synthetic data and simulation, you can harvest the same fruit over and over again. You can also make the crop jump to a different growth stage and virtually introduce a new variety instantly.”

A simulation environment for robots in open fields and orchards. By connecting the simulation to ROS2, robots can be tested in a realistic environment.
3 main benefits of using synthetic data and simulation
The researchers identify three main benefits of using synthetic data and simulation. The first benefit relates to the amount of data required for creating AI models. Bresilla: “Data and images must be annotated. This (partly) still is a human activity, which is very time consuming and thus expensive. In a simulated environment, you can instantly introduce a lot of variation. Such as daylight, night/darkness, sun, dew, fog, shadows, et cetera. And you can let the computer generate randomised circumstances and do the annotation work.”
The second advantage is that synthetic data offers flexibility. “Cameras and sensors used to capture data are improving at an incredibly rapid pace. Today’s cameras capture much more detail than yesterday’s ones which means the data age very quickly. In a virtual setting, you can simply introduce a new dataset for new cameras.”
In a simulated world you can, last but not least, anticipate to situations that are rare or still need to happen. “You can introduce longer cucumbers, red carrots or yellow oranges. Or a new type of (invasive) weed that isn’t present (yet) in natural fields. And you can use every type of sensor in every angle you desire without having to fit and change these in the real world.”
Physical AI to be the next wave
Vroegop recently attended NVIDIA GTC’25, a congress on AI hosted by NVIDIA that drew over 25,000 attendees to California USA. “Such congresses and the networking possibilities associated with them, are essential for researchers and developers to stay up to date. It made me realise that techniques like deep learning are becoming a commodity and available for everyone as a ready-to-use tool. The next major step in AI will be physical AI. This is all about self-learning robots (e.g. reinforcement learning) that interact with complex environments. The congress also confirmed that the agrifood domain is too small to make certain developments affordable. Which means we must benefit from other sectors such as the automotive industry.” Bresilla adds: “For instance on the perception part, where we see that the use of LiDAR sensors is challenging in dusty circumstances while sonar sensors don’t have that difficulty.”
The (near) future
Arjan Vroegop: “Predicting what comes next in our line of work is a challenge on its own, let alone to predict what will happen in five years from now. I believe that in the near future our approach will be simulation first. I think we will develop and make things work in a simulated world first, and then bring it to the real world. I also believe that robots will no longer be programmed by humans but that it will be about training a model. We will shift from hard coding to training robots that are able to learn by themselves from their experiences.”
Trim Bresilla: “As simulation gets better and better, we will evolve to simulation driven design. I believe that the so-called sim to real gap, that is still existing, will fade. In some cases the gap will even get undisguisable. If you want to be part of this, then you should definitely get in touch with Vision + Robotics as we probably are one of the few institutions across the world that combine all the necessary competences and skills as well as the required facilities in the natural domain. We know how to help companies with synthetic data and robotic simulation and that still is quite unique.”
Trim Bresilla, researcher digital agriculture, AI, and robotics, and Arjan Vroegop, researcher mechatronics, robotics, and automation at Wageningen University & Research’s Vision + Robotics are both part of a team of currently about 60 researchers headed by program manager Erik Pekkeriet. They specialise in, among others, (generative) AI, deep and machine learning and robotic simulation.
