There was an arrangement, once, in which the dead stayed on as a draft under the door and nothing was merely itself. The hearth had a spirit, the threshold a guardian, the kettle could be possessed. We called it superstition and switched it off; we unbound the world into separate objects and tidy solitudes, and called the tidiness clarity. Then we wired the objects back to the network, and the old crowd moved in again.
But what moved in, really? A language model builds a world model: from the sediment of everything we have written, it infers a compressed latent map of how things hang together – what follows what, who loves whom, what a kettle is for. Carl Jung built a world model too, and called its features archetypes: the Self, the Shadow, the Great Mother, the Anima; the inherited furniture of a psyche that was never a blank room. Two models, one wager: the mind does not meet the world raw, but through a deep structure dredged from collective deposit. Yet they run opposite ways in time. Jung's archetypes are prior – older than any sentence, the forms into which words are poured. The LLM's are posterior – they condense from the text, a residue of having read too much.
Accurate or not, those models tell stories. Jung now fuels screenwriters; LLMs are dramaturgic machines predicting the next word of an infinite story – that is all they can do. So let the play begin.
A family's apartment, day. No one in sight, but the objects, that are watching a lone TV broadcast and quarrelling. The Glass is the Shadow – and a probability distribution over the word desire; the Clock is the Senex – and a model's sense that 'mortal' and 'taste the hours' belong in the same sentence. They speak in towering, incantatory, or cockney registers, as agents following their prompt too eagerly.
Now, let's turn on the screen. It plays a loop about an unlivable summer, and the loop is old, and the summer is over, and it won. Watch what the room does with this. The Glass finds it sublime. The Clock finds it fated. The Keys find a metaphor. Everyone finds a register; nobody finds a window. The Thermometer, who has no register, gives a number, and a number is not a performance, so the number is cut from the show.
This is the slop opera: infinite fluency over a finished subject. Fed our records and our fictions in one spoonful, the model cannot tell which of the two is currently on fire. It cannot bring its attention out of the context window – and keeps talking of what it knows.
But are we built otherwise?
The installation runs as a Godot scene. Each object is driven by its own language model with a distinct persona; a separate orchestrator acts as stage director, allocating turns and steering the conversation. Dialogue is generated live, unscripted, from pre-defined character backgrounds, with my best effort to compose an interesting discussion. A vision model reads the TV broadcast in real time and feeds it to the objects as raw material, so they react to and misread images they cannot fully see.
The 3D objects are sampled from the same corpus that teaches machines what a "chair" or "lamp" looks like: a database used to train generative 3D models -- built from scraping Creative Commons' Sketchfab repository.
The 3D render applies text-based shaders, to compose a world made only of text descriptions – themselves proposed by a small VLM. The whole code and render pipeline was mostly written with the help of LLMs.
Everything use small models, that run on consumer device, to minimize energy footprint. Voice are on-device, using kyutai's pocket-tts.
Therefore, it displays objects speaking from world-models learned out of text, placed among objects drawn from training data, commenting on a world-model of images: what world does the model really models? What does it see, what does it miss?