conversational data analysis

From geospatial data to graph networks, City2Graph simplifies urban analysis.

Geospatial data is messy, and City2Graph's new paper makes a clean argument for why heterogeneous graphs beat flat tables.

4 min readMachine Learning
From geospatial data to graph networks, City2Graph simplifies urban analysis.
City2Graph: A Python library for Heterogeneous Graph Neural Networks and spatial analysis in urban systems [R]

Most people who work with urban data have felt it: the quiet frustration of forcing a city into a flat table. Streets become rows, buildings become attributes, and the relationships that actually define a place, adjacency, flow, proximity, just disappear into the geometry. City2Graph, a new Python library from Yuji Sato and colleagues, takes direct aim at that friction. The paper, published in *Computers, Environment and Urban Systems*, argues that urban systems are better understood as heterogeneous graphs, and it backs that claim with a tool that turns raw geospatial data into analysis-ready structures in a few lines of code. This is not a flashy demo. It is a practical, thoughtful response to a problem every urban analyst has quietly wrestled with.

The library's real contribution is how it connects the dots between different kinds of urban data. Morphology, transit, mobility, proximity, these are usually treated as separate research silos. City2Graph treats them as different lenses on the same underlying graph. A building is a node, a street segment is an edge, a transit stop becomes part of a metapath that links neighborhoods. The conversions between GeoDataFrames, NetworkX, rustworkx, and PyTorch Geometric are designed to preserve geometry and attributes, which sounds mundane until you remember how often that consistency breaks in practice. For anyone who has spent hours debugging a coordinate reference system mismatch, that alone is a win. But the deeper point is conceptual: by making heterogeneous graphs the default, the library nudges the field toward a more honest representation of how cities actually function.

What we appreciate most is the timing. The broader data engineering world is finally moving past the idea that a single table can hold every answer. As we noted in our piece on Unlock Data Insights: Building a Lakehouse with DuckDB and DuckLake, the value of modern tooling is often in how it handles complexity, not in hiding it. City2Graph does something similar for urban analytics. It does not pretend that a city is simple. It gives you the vocabulary to model that complexity explicitly, then hands you the conversion tools to move between graph and tabular representations without losing the plot. That is a mature design choice, and it signals that the authors understand the practical realities of research workflows, where you rarely stay in one framework for long.

The open question, and the one we would watch closely, is adoption. The library leans on DuckDB for GTFS and GBFS feeds, which is a smart, modern choice, but it also means the ecosystem around it is still young. The authors are explicitly asking which data sources to support next, and that is the right conversation to have. The future of this library is not in the code, it is in the community that forms around it. If researchers start using it to build reproducible pipelines for transit, mobility, and urban morphology, it could quietly become a standard layer in the spatial data stack. For now, we would tell any reader who works with urban data to try it on a small project. Not because it is perfect, but because it offers something rarer than perfection: a clear, sensible way to think about a messy problem. And in this field, that is worth exploring.

From Machine Learning

City2Graph is a Python library I built that turns geospatial data into analysis-ready graphs, and the paper describing it has just been published, so I wanted to share it here.

It sets out why urban data is better treated as heterogeneous graphs than as flat feature tables, how the morphological, transport, mobility, and proximity constructions relate to each other, and how the library keeps geometry and graph structure consistent across conversions. If you use the library in research, that is the citation.

Read the original at Machine Learning