data compatibility

Self-contained schemas make your data readable across past, present, and future tools.

Data formats usually die because they can't evolve.

3 min readInfoQ
Self-contained schemas make your data readable across past, present, and future tools.

Most of our data lives in formats that will outlive the tools we use to read them. That is the quiet crisis behind every migration project, every archived CSV, every spreadsheet that breaks when the software that created it disappears. Seph Gentle's proposal takes a different route: instead of asking us to trust external schemas or central authorities, he suggests embedding self-contained schemas directly into file headers. The idea draws from the same adaptability that kept HTML and HTTP alive for decades. It is not a new product or a plugin. It is a bet that data can carry its own interpretation with it, like a letter that includes its own dictionary.

That bet deserves serious attention, especially when we consider how much of our collaborative work depends on formats that refuse to evolve. We have all felt the friction of a file that opens in one tool but not another, or a column that loses meaning because the schema lived in someone's head. Gentle's approach addresses that by making the schema a first-class citizen of the file itself. Forward compatibility means older software can skip what it does not understand. Backwards compatibility means newer tools can still read legacy data. Sideways compatibility means two different systems can share a file without prior coordination. That is not a minor technical convenience. It is a structural shift in how we think about data ownership, one that aligns with the broader push toward AI-powered knowledge graphs for seamless development, where meaning is extracted and preserved rather than assumed.

But let us be honest about what this is not. This is not a magic bullet that solves every interoperability problem overnight. The format is experimental, and its success depends on adoption across tools, libraries, and teams. That is the hard part. We have seen how quickly promising formats fail when they rely on everyone agreeing to a standard. What makes Gentle's approach different is that it removes the need for agreement up front. You do not have to convince the world to adopt a new schema. You only have to make sure your file knows what it is. That is a lower bar, and it is also a more honest one. It accepts that data will outlive any single system, and it plans for that reality.

Our take is straightforward: this is the kind of thinking that turns a clever idea into a durable practice. It does not ask users to trust a platform or a vendor. It asks them to trust that a well-formed file can stand on its own. For teams already struggling with the gap between how work is imagined and how it actually happens, as explored in presentations on marathon incidents, the ability to preserve meaning across time and tools is not a luxury. It is a risk-control measure. And for those of us who have watched promising initiatives stall because of data lock-in, this is worth watching closely. The specific thing to track is whether toolmakers embrace it as a default, not a feature. If they do, we may finally stop treating data migration as a project and start treating it as a property of the file itself.

From InfoQ

Drawing from the enduring adaptability of HTML and HTTP, Seph Gentle proposes embedding self-contained schemas directly into file headers, ensuring data remains readable without external definitions. His experimental format prioritises forward, backwards, and sideways compatibility, enabling data format evolution without central coordination or data loss

Read the original at InfoQ