Wow, this is really cool, looking forward to trying it out!
For anyone unfamiliar since the original post is pretty technical, this would let you take basically any VaM scene and render out _structured_ data in order for AI to interpret and generate an entirely new image based on your given prompt (for example.) I say that but I'm also aware that this community is no stranger to AI workflows.
One curious question: would it be possible to export the depth map and/or poses as videos in the future? That could be really useful for video models.