Dierks: If you look at the way software interfaces are made today, you see that AI is changing things drastically. Today the standard is a GUI interface. Everything you can do with a software, you can see in the GUI and you have tons and tons of menus and stuff. The problem with that approach is that software is becoming so feature rich that the GUI is getting too complex and it has too many features, so you can´t really deal with it any more. The next step is that you just get a chat interface, with a little GUI to see the status, but in the end, you tell your software what you want it to do, and then it just do it. Something similar might happen to the way cameras are configured. In the old days, we had a fixed register layout. Today it’s GenIcam. But even with GenIcam, you have the problem that there are too many features today. My vision is that you get a kind of MCP (Model Context Protocol) style interface based on a SDK. You just take your AI, connect it to the camera and the camera not only exposes the features, but also certain traits on how the image will be treated. You will be able to tell the camera: Can you please increase the contrast, because is looks awful. Then the camera will just look at a histogram, find out what to do, change it, and show a live image. You say OK and then everything is done. We already have done prototypes for that. Maybe some of these ideas may even make it to GenIcam one day.
M. Schmidt: We built an interface ready for modern AI twenty years ago. SFNC works great as an input to large language models like ChatGPT, because it’s a collection of the common wisdom of the machine vision industry. It’s a perfect instruction, how to use a machine vision camera. The XML district description can provide most of the necessary information for a LLM to understand and use the custom features of a camera that are not in the SFNC because the XML description has documentation fields which are a good guide for an AI to configure and to compose systems. Generative AI is nothing I see in machine vision devices, in the future. It will have its place, when it comes to creating and optimising machine vision systems. LLMs should be able to propose the right components and tell you how to connect them and then write the application for it. But I don’t see it in the image processing loop because it’s just too slow and too unpredictable for most applications. GenIcam has a good future in combination with AI systems and will look very similar to what we have today in ten years.

















