An experiment documented in a Quesma blog post tasked Opus 5.5 with visualizing Invisible Cities from a single prompt, then allowed the system to run for six hours.
The setup strips the interaction down to a basic question: what can an AI system produce when it receives a broad creative brief without repeated human steering? The exercise shifts attention from short-answer capability toward sustained execution on an open-ended task.
For startup teams building with generative AI, such tests are more relevant than polished demonstrations alone. Long-running workflows can expose questions about instruction quality, reliability, review requirements and the practical limits of autonomous output.
The account, shared through Hacker News, offers a narrow case study rather than a general benchmark. Still, it reflects growing interest in evaluating AI systems through complete creative or technical tasks instead of isolated prompts.
