A couple of years ago, after reading the book, I sketched some of them on Procreate with an isometric grid, but not all 55. I only arrived at number 4, since each took multiple hour...
Since this project (both yours and OP's) seem more for personal fulfillment and enjoyment, I wonder who got more enjoyment out of the process, and also who will remember the results better in ten years.
It’s interesting, with a human made piece of art, I’m always drawn to « zoom in », to see how you did little details, if you have hidden something here and there, how you connected different parts.
With AI generated stuff, sure it looks amazing at first glance but I definitely don’t want to zoom in, since I know it’s all a cardboard facade with no passion. Look a bit too closely and you’ll see the bridges that make no sense, the bunch of nonsensical threejs cylinders etc.
> Leaving there and proceeding for three days toward the east, you reach Diomira, a city with sixty silver domes, bronze statues of all the gods, streets paved with lead, a crystal theatre, a golden cock that crowns each morning on a tower...
Text in the second person is causing my my D&D instincts to kick in.
I understand why that’s impressive. But I don’t feel much. I guess, I was just used to envy how much people can be absorbed by some idea to spend so much time implementing it. Now it feels like a single shotted nice-typography-and-all presentation at work that someone did 10 min before the meeting.
There's also an eerie feeling of emptiness to the outputs. They lack real intent. Glossy and technically impressive, but they don't tickle your brain. I think it's because they don't really communicate or give you that buzz of new insights. They are pretty but not _helpful_.
For example, I didn't learn about lenses or earthquakes or fusion messing with the controls on these.[1] I didn't know what to look at.
However, there were interactive exhibits at the Children's Museum I spent a long time exploring and loved. That's because they were carefully designed to really teach the concept?
Yeah, the demos look really cool but you're thrown into them and there's nothing explained as to what you're looking at or why you should care. Funny, an article explaining something like how a camera focus works with a few illustrations would be far more useful than a fancy 3D interactive.
It turns out the more technology you throw at something does not equate to ease of learning.
Another thing to keep in mind is that if you asked twelve different artists to render these cities from their descriptions, you would likely end up with twelve very different renditions.
Now imagine the same thing but each person is using Opus 5.5 - I’d wager there would be a lot of commonality among each of the outputs.
> If something that previously took weeks for a human to design, now takes 10 min, that means a whole week of salary went up in smoke.
To me, this is the depressing sentiment. The idea that, somehow, the goal isn't to produce useful things but to spread work out over a period of time in order to maximize payment.
At a high level,I conceptualize my job as "doing useful things." Whether that's design, writing code, debugging, sysadmin, CI/CD, writing doc, whatever. I just want to be useful. If there's some tool that makes that easier, I am happy about it. I don't think that tool will result in my getting paid less. But if it does, oh well.
Weird way to look at it, from another view we can now print a week salary in 10 minutes. Might devalue the output slightly but we still have the product.
If you haven’t read the book yet, I’m tempted to warn you against opening the visualization. Part of the joy of reading this book is letting your mind’s eye run wild, and I fear if I read it for the first time after toying with this site, I’d be remembering the cities rather than conjuring them.
That is the exact problem with the city of Dinipro.
Citizens of this city high in the sky imagine it so fully in their minds that no one there needs to open their eyes to walk the streets and live their normal lives.
But if one makes the mistake of opening their eyes and tries to see the city directly, they would immediately fall through the clouds to the ground deep below.
I heartily recommend the audiobook, narrated by John Lee.
The book is a trip and the dreamy sort of narration matched it perfectly. My wife and I would spend evenings listening to one city at a time, on a walk around our little park, one headphone in each ear. I don't know if it was just the environment or the time in our lives, but we both really look back fondly on this book.
I read the book decades ago, then I lost it in one of the many relocations. I always thought I should buy it again but never acted on the idea... A few years ago, I was walking by an antiquarian, and noticed a thin unmarked gray spine. I entered, picked it up, paid, and walked away.
PS I didn't open the gallery. Nice idea but I don't care about visualizing something I should feel.
How Opus 5.5 imagines these cities is a question for Opus. The real experience is how we imagine - and I believe there are no two people who have exactly the same mental image of an invisible city.
> I’d be remembering the cities rather than conjuring them
Having looked at some of them (and it being one of my favorite books) ... it won't do that. These visualizations are as terrible as they are unnecessary.
Maybe the interactive link in the article for the optic focus can lead you on a fun journey instead. And the bottom-right of every visualization is a link to yet another fun diorama to explore.
It's a techie creating the torment nexus but the imagination version. No need to picture in your mind the invisible cities! Time to make them visible!!!
Talk about missing the entire point of the piece of literature.
Invisible Cities is my favorite piece of literature hands down. The audiobook narrated by Richard Higgins is also done so beautifully and can't recommend it more.
As neat as this project is, this does no justice to whatsoever to the book. On the surface Invisible Cities seems to be a book about random cities' descriptions. In reality though, this book is about semiotics, meaning, language, and it's limits.
The first city I clicked on at random had the opening phrase:
> Arriving, you rejoice at its bridges, each different
The model city had like five bridges total, three of which were didn't bridge anything at all but instead sat in the middle of the river and went along its flow, the "bridges" for some reason pretending to be small islands.
And shame on all the people shitting on it because AI made it. If you valued art for art's sake, you wouldn't be so bothered when somebody makes something art-adjacent using some technique you don't approve of. It in no way detracts from the kind of art you admire. Don't be type who just has to let others know that you find something they like to be beneath you.
While reading the book years ago, I did not visualize Isidora as a city with 9 houses, like a children’s book, or a little video game level intro, but so unlike Kublai Khan or Italo Calvino.
One of my favorite things about the National Treasure Cinematic Universe (two movies and a TV show) is the existence of characters that I like to call "treasure grumps". Their role in the story is to warn the main characters that searching for treasure is a terrible way to spend your time, and can only lead to personal and familial ruin.
wow that's awesome! super cool :), fyi so inspiring that I did something similar for d&d forgotten realms, including a timeline: https://narfman0.github.io/realms-atlas/
fable drove it, opus agents did the work, took 30-45 minutes. (intend on extending to planescape, elaborating on significant historical events, and some text to speech narrating cool events)
This is probably one amazing positive these models have brought forward - the ability to break out of a single modality, and utilise others to help us learn and visualise. While the visuals here are amazing, my son for example prefers to learn by listening and talking, so we convert lot of his study materials into audio and real-time voice roleplay.
I was finishing Invisible Cities yesterday, because the book was due at my local library today. I was wondering whether the cities could be adapted into some visual art form. And obviously I thought about feeding it to Opus. 24 hours later, I see this post…
Citizens there gather their tools at the bottom of a giant staircase and climb up countless stairs, only to find out that what they were planning to do has already been done by the time they reach the top.
Dissapointed, they descend to grab a new set of tools and make new plans for tomorrow, only for the same thing to happen again.
both the linked camera lens example as well as the last one the author linked runs at ~5fps on vivaldi and pegs CPU 100%. i have never seen any 3d visualization website have performance this bad, i assume they're not using webgl given my GPU is 0% utilization.
Six hours of continuous runtime is the underrated part here. Most one-shot demos fall apart well before that, so this doubles as an endurance benchmark.
Why, I once encountered a running session longer than a day!
The agent had spun up a backward shell script to watch for the shutdown of another process, but wrote a bug in the script that would have left it running indefinitely until I got home and noticed it.
This was with Fable, no less! And it happened a few more times, though I caught them sooner.
I’m not sure runtime is an important metric at all. Shouldn’t we aim for 0 runtime with maximal results?
The supervisor agent will keep the session in a loop until 6 hours have passed and eventually the agent will decide to use up the remaining time rather than fighting with it
I wonder how much co2 is being thrown into the atmosphere everyday through the steady stream of “look what I made this LLM do” and endless “benchmarking”?
One useful hint here is the API token cost - in this case "about $74 in API tokens" for Opus 5.5.
We know Anthropic run inference at a margin, so that $74 means that the cost of the electricity involved is substantially less than $74. I'd love to know the actual cost there.
Approximately nothing - the carbon cost of electricity is already very low, and most of the inference costs are GPU and datacenter amortization. You could require all AI inference to be carbon-neutral and that'd barely raise the API prices.
I wonder how much co2 you use to get through the day, and how much you use to remark on other people's creations in a way that suggests you disapprove of their utilization of their available resources. How much co2 do your projects emit, since you seem to be interested in those metrics?
I strongly suspect one HN post (or two now) is not even in the same ballpark. I could ask Claude to figure it out but that would be a complete waste of time and energy. ;)
Intuitively I'd say that just-for-fun AI projects emit orders of magnitude less CO2 than what just-for-fun road trips or air travel emit.
So unless you can show that it approaches all the other frivolous consumption that humans love to waste resources on, maybe we can get back to experimenting with cool technology and talking about it, on this website called "Hacker News"?
$74 vs $ 10 vs $25 for the same prompt,
The interesting number is not actually quality, its what we can get per dollar like 6 subagents runs in parallel.
It reminds me of Total War. The soundtrack and way that the map is shown. Is Total War an inspiration? Anyways, I liked the visualization. But I never read the book, so I followed (partly) droidjj recommendation
What is this for? It doesn't deepen the understanding of the book. The illustrations aren't attractive on their own. They're not a very good representation of the descriptions in the book.
I'm a little torn by this sort of demonstration, because while many of the scenes have obvious markers of slop (impossible intersecting geometry, bridges to nowhere, and so on), the scenes largely do work to convey the intended concept/emotion, and the low-poly aesthetic is executed decently well.
I guess I shouldn't be surprised that another visual medium is starting to fall to LLMs after the success of diffusion models for image generation. Nevertheless it's wild that a mostly text-focussed model is building little worlds like these.
> The days of techies gatekeeping these experiences are over.
Why do people keep using the word "gatekeeping"? "Techies" made open source. Anyone could have taken the time to learn what they wanted, the tools and the resources were available.
The disdain for tech will always baffle me. Yes, this stuff is hard. A lot of it requires a certain level of skill. Even now, you still require some level of skill to get good results.
Also, the "techies" are naturally going to be a lot better at pushing AI to its limits. This doesn't change the playing field - it means more people can do more things, but the upper end of the skill range is going to be just as far ahead, if not further ahead, of everyone else as it always was.
> Anyone could have taken the time to learn what they wanted, the tools and the resources were available.
False in reality. Just ask any dev who has had to listen to family and friends' app ideas over the last decade. The good news for them is soon they will never be asked again.
Having a concept is completely different, their unwillingness to dive into implementation cannot be blamed on the "techies". I would wager that the same people who were pitching unrealistic app ideas then are still not trying to make apps using AI.
What AI has largely changed is how the communication barrier (or perception thereof) is dealt with between two sides: one which consists of those who perceive gatekeeping or attribute their lack of substantive progress to it, and the other side which consists of the time, resources, and skills the former feels entitled to or indebted towards.
I’ve had a running theory for a few years now that some tech folks are more open to receiving an error message from a compiler or other tool, than they are to critical feedback from a peer. If they’re blocked and finally have exhausted their other avenues (pre-LLM: web searches, forums, etc), they’ll reluctantly engage with a peer, and do so with extremely limited patience.
The first paragraph of this reply is extending that theory to non-tech —-> tech, but replace compilers with LLMs that _appear_ to be more useful (or at least busy) working with what limited context you provide them, especially when it’s insufficient.
It won’t quite tell you your app idea is unrealistic or that the implementation is nearly impossible, but after burning a bunch of time and tokens, those people will eventually get the point.
This is more comparable to a mass produced plastic replica than the handcrafted work of an artist. Experience and skill are not orthogonal to creativity or taste.
And scarcity gives things value. No one bats an eye at a mass produced cotton t-shirt with a detailed custom print on it. There's services online where you can order a thousand with whatever print on it for quite a cheap price.
But take one of those shirts 80 years back in time and it'll likely impress everyone.
Perhaps that's what's going to happen to software too. If you're an idea guy you might get a monkey paw situation, at least if you planned on making some sort of serious income from your software idea.
Now, you just need to use someone else's words and someone else's AI built on everyone else's stolen art! But sure, yes, thats you being creative... Don't use your imagination at all.
There are three main categories of AI ethical dilemmas for me:
1. IP infringement
2. Climate impact
3. Cultural impact
I think that 2 & 3 apply to all AI right now - artistic, coding, writing, whatever. However, I don't think that #1 applies to everything. Code has been openly shared for as long as it's been possible to do so. There have probably been cases where AI trained on closed-source code, but I hardly believe that was ever necessary at any point to get where we are.
My point in saying this is that your critique of AI is valid but I think a little less so here.
That's fine, I'm only making the point that it's not really relevant to coding. Programmers have been outright copying each other's code for decades to the chagrin of nobody.
I totally understand. I'm a developer with almost 20 years experience myself and it has taken me a while to come to terms with this too. Fortunately, I made enough during the good times to see me out once my job ceases to exist. I'm sorry for everyone who didn't.
https://camillovisini.com/drawing/fc4jc6-le-citta-invisibili...
https://camillovisini.com/drawing/p2yrdj-le-citta-invisibili...
With AI generated stuff, sure it looks amazing at first glance but I definitely don’t want to zoom in, since I know it’s all a cardboard facade with no passion. Look a bit too closely and you’ll see the bridges that make no sense, the bunch of nonsensical threejs cylinders etc.
> Leaving there and proceeding for three days toward the east, you reach Diomira, a city with sixty silver domes, bronze statues of all the gods, streets paved with lead, a crystal theatre, a golden cock that crowns each morning on a tower...
Text in the second person is causing my my D&D instincts to kick in.
For example, I didn't learn about lenses or earthquakes or fusion messing with the controls on these.[1] I didn't know what to look at.
However, there were interactive exhibits at the Children's Museum I spent a long time exploring and loved. That's because they were carefully designed to really teach the concept?
[1] https://sael.net/fusion-pulse/?ref=earthquake-tower
It turns out the more technology you throw at something does not equate to ease of learning.
Now imagine the same thing but each person is using Opus 5.5 - I’d wager there would be a lot of commonality among each of the outputs.
If something that previously took weeks for a human to design, now takes 10 min, that means a whole week of salary went up in smoke.
When have we automated enough?
To me, this is the depressing sentiment. The idea that, somehow, the goal isn't to produce useful things but to spread work out over a period of time in order to maximize payment.
At a high level,I conceptualize my job as "doing useful things." Whether that's design, writing code, debugging, sysadmin, CI/CD, writing doc, whatever. I just want to be useful. If there's some tool that makes that easier, I am happy about it. I don't think that tool will result in my getting paid less. But if it does, oh well.
The problem also is, if a human took weeks to design this, it would be much better. This is quick but also bad.
Citizens of this city high in the sky imagine it so fully in their minds that no one there needs to open their eyes to walk the streets and live their normal lives. But if one makes the mistake of opening their eyes and tries to see the city directly, they would immediately fall through the clouds to the ground deep below.
The book is a trip and the dreamy sort of narration matched it perfectly. My wife and I would spend evenings listening to one city at a time, on a walk around our little park, one headphone in each ear. I don't know if it was just the environment or the time in our lives, but we both really look back fondly on this book.
PS I didn't open the gallery. Nice idea but I don't care about visualizing something I should feel.
How Opus 5.5 imagines these cities is a question for Opus. The real experience is how we imagine - and I believe there are no two people who have exactly the same mental image of an invisible city.
The website gives you a good hint before opening up the visuals. So, I got an idea.
I don't really know about Marco Polo. My only knowledge comes from the Netflix series, which doesn't talk about these invisible cities.
Having looked at some of them (and it being one of my favorite books) ... it won't do that. These visualizations are as terrible as they are unnecessary.
Talk about missing the entire point of the piece of literature.
As neat as this project is, this does no justice to whatsoever to the book. On the surface Invisible Cities seems to be a book about random cities' descriptions. In reality though, this book is about semiotics, meaning, language, and it's limits.
> Arriving, you rejoice at its bridges, each different
The model city had like five bridges total, three of which were didn't bridge anything at all but instead sat in the middle of the river and went along its flow, the "bridges" for some reason pretending to be small islands.
Opus Eutropia has a few isekai circle cities with one of them... ON the river, giving no space for ship traffic.
Slop has never been this beautiful before!
And shame on all the people shitting on it because AI made it. If you valued art for art's sake, you wouldn't be so bothered when somebody makes something art-adjacent using some technique you don't approve of. It in no way detracts from the kind of art you admire. Don't be type who just has to let others know that you find something they like to be beneath you.
I once instructed it to spawn at most 20 subagents and it spawned 40, with an adversarial reviewer for each of the 20.
It doesn't follow these instructions very well.
rock on
Source code: https://github.com/stared/invisible-cities-opus-5.5
Citizens there gather their tools at the bottom of a giant staircase and climb up countless stairs, only to find out that what they were planning to do has already been done by the time they reach the top.
Dissapointed, they descend to grab a new set of tools and make new plans for tomorrow, only for the same thing to happen again.
The agent had spun up a backward shell script to watch for the shutdown of another process, but wrote a bug in the script that would have left it running indefinitely until I got home and noticed it.
This was with Fable, no less! And it happened a few more times, though I caught them sooner.
I’m not sure runtime is an important metric at all. Shouldn’t we aim for 0 runtime with maximal results?
The supervisor agent will keep the session in a loop until 6 hours have passed and eventually the agent will decide to use up the remaining time rather than fighting with it
Sigh
We know Anthropic run inference at a margin, so that $74 means that the cost of the electricity involved is substantially less than $74. I'd love to know the actual cost there.
What a terrible thing to say to human! Being (presumably) human, it is their inalienable, natural-born right to do so.
Shame on you.
So unless you can show that it approaches all the other frivolous consumption that humans love to waste resources on, maybe we can get back to experimenting with cool technology and talking about it, on this website called "Hacker News"?
Vanadium (chrome)
https://imgur.com/a/nGLyJf6
I guess I shouldn't be surprised that another visual medium is starting to fall to LLMs after the success of diffusion models for image generation. Nevertheless it's wild that a mostly text-focussed model is building little worlds like these.
I'm so glad I'm able to witness this Cambrian explosion of software :).
Why do people keep using the word "gatekeeping"? "Techies" made open source. Anyone could have taken the time to learn what they wanted, the tools and the resources were available.
The disdain for tech will always baffle me. Yes, this stuff is hard. A lot of it requires a certain level of skill. Even now, you still require some level of skill to get good results.
The top-level comment feels a lot like the perspective of somebody who never cared to try.
False in reality. Just ask any dev who has had to listen to family and friends' app ideas over the last decade. The good news for them is soon they will never be asked again.
I’ve had a running theory for a few years now that some tech folks are more open to receiving an error message from a compiler or other tool, than they are to critical feedback from a peer. If they’re blocked and finally have exhausted their other avenues (pre-LLM: web searches, forums, etc), they’ll reluctantly engage with a peer, and do so with extremely limited patience.
The first paragraph of this reply is extending that theory to non-tech —-> tech, but replace compilers with LLMs that _appear_ to be more useful (or at least busy) working with what limited context you provide them, especially when it’s insufficient.
It won’t quite tell you your app idea is unrealistic or that the implementation is nearly impossible, but after burning a bunch of time and tokens, those people will eventually get the point.
Or reluctantly reach out to a tech person anyway.
But take one of those shirts 80 years back in time and it'll likely impress everyone.
Perhaps that's what's going to happen to software too. If you're an idea guy you might get a monkey paw situation, at least if you planned on making some sort of serious income from your software idea.
1. IP infringement
2. Climate impact
3. Cultural impact
I think that 2 & 3 apply to all AI right now - artistic, coding, writing, whatever. However, I don't think that #1 applies to everything. Code has been openly shared for as long as it's been possible to do so. There have probably been cases where AI trained on closed-source code, but I hardly believe that was ever necessary at any point to get where we are.
My point in saying this is that your critique of AI is valid but I think a little less so here.