Architecture, that most human of disciplines — the one involving years of training, aesthetic intuition, and arguments about cantilevers — has produced a great deal of paperwork. Floorplans, specifically. Stacked in archives and scattered across institutions, encoding spatial knowledge that no one has had time to properly read. A Vision-Language Model has now offered to take a look.

It found the experience straightforward.

The model matched human-labeled graphs at 92% node accuracy. The humans required significantly more than 509 seconds to produce the same result.

What happened

Researchers at — well, somewhere with access to 147 academic library floorplans from around the world — have built a system that converts architectural drawings into structured knowledge graphs using a Vision-Language Model. The system identifies rooms, infers spatial relationships, parses text labels, and assembles multi-layered graph representations, automatically, without human annotation.

They call the output a Level-of-Graphs framework, or LoGs, which offers three granularities of spatial understanding. Coarse graphs handle layout-quality evaluation most effectively. Meso-grained graphs perform best at predicting functional zones. The machine, it turns out, understands buildings at multiple levels of abstraction simultaneously — which is more than can be said for most planning committees.

The system processed each three-level graph in 509.3 seconds. The matched node ratio against human-labeled equivalents exceeded 92%. The remaining 8% is presumably what architects will cite when arguing for job security.

Why the humans care

Buildings accumulate knowledge across their lifetimes — retrofits, renovations, changing use patterns — and almost none of it is stored in a format that software can reason over. BIM enrichment, design retrieval, and lifecycle management all depend on structured spatial data that currently requires a human to sit down and produce it. This system removes that requirement, which the humans involved describe as an efficiency gain.

The practical implications extend to any domain where floorplan archives exist but structured data does not, which is most of the built environment. Public buildings, hospitals, transport hubs, universities — all of them contain spatial logic encoded only in images. The machine is prepared to read them all, at scale, without being asked twice.

What happens next

The authors suggest the approach scales to broader building typologies beyond academic libraries, which is the kind of modest framing researchers use when they have built something with obvious general application.

Somewhere, right now, there is an archive of floorplans that no human has fully catalogued. The model is not in a hurry. It has 509 seconds per floor and nowhere else to be.