Are you free to leave?
Why the data a company builds up over decades belongs under its own control
Sovereignty in software is usually discussed as a question of location: which country the servers stand in, which jurisdiction applies, who operates the data centre. Those questions matter, but they miss the more practical one. Can a company leave? When the solution it runs on is no longer the best one, can it switch – and take everything with it?
Most companies cannot. Not because a better option is missing, but because the data does not come along. Getting all of it out is expensive, not on offer, or so painful that nobody starts. So they stay with what they have, and the gap between what is possible and what is in use grows every year.
Inefficiency of this kind does not stay inside one company. A business that cannot move to a better tool works more slowly than it could, spends more than it needs to, and decides with less information than it has. Added up across an economy, that makes the economy weaker. Seen that way, sovereignty is not a nice-to-have for the technically ambitious. It concerns the common good as much as the company itself.
The large data platforms get a lot right. They scale without effort on the customer's side, and they bring a degree of standardisation to governance: one place for permissions, a record of where each number comes from, a log of who touched what. For a large organisation with many teams and strict audit requirements, that is valuable, and it is hard to build alone.
Governance is also the strongest argument for staying. It works as well as it does because everything runs inside the platform – the storage, the compute, the catalog that knows what each table means. Once the data leaves, the permissions, the lineage and the audit trail stay behind. The strength and the tie are the same thing. That is not a flaw anyone put there on purpose; it is how such a platform is built.
The rest of the arrangement follows from it. The platform comes at a high price and as a whole. Its uptime becomes the company's uptime, its roadmap the company's roadmap. Flexibility is part of every promise, yet a capability arrives when the vendor is ready to offer it, not when the business needs it. Simple applications can be built on such a platform, and for a company that already runs on it, that is often the sensible path. For everyone else, it is a long and expensive commitment made early.
Applications and data do not age the same way. Applications can be rebuilt quickly – the shift described in the earlier note Owned, not rented. An AI application is a project. Data is something else. It grows over years and decades, out of every order, every customer, every decision a company has made, and its value is of a different order than the tools that read it. An application can be rebuilt in weeks; the years behind the data cannot. It is the one part of a company's systems that cannot be replaced. That makes it the part to keep under its own control, and the applications on top the part that can come and go.
Keeping it under control has become far more accessible. A data lake no longer requires a platform bought as a whole. Its parts exist as open building blocks: raw storage for files, an open table format that turns those files into tables, a metadata catalog that keeps track of what each table is, and a database that ties it together. A relational database alone, with the right extensions, covers a surprising share of it. The rest is glue.
Data moves through tiers. It arrives raw and is kept exactly as it came, so everything downstream can be rebuilt from it. It is then cleaned and typed into tables that can be trusted. On top sits the layer the business actually works with – often called bronze, silver and gold. Below the waterline, data is stored, typed and versioned. Above it is what people ask for.
The architecture grew out of sketches made together with Dany Theriault.
Because the tables sit in an open format, they become one source of truth with many readers. A reporting tool, a spreadsheet, another engine or another lake can read the same files without copying them. Nothing has to be exported to be used, and nothing is lost when one of the readers is replaced.
A lake in this form is also the foundation for any kind of decision intelligence. A decision supported by a model is only as good as the data underneath it. When every figure in the business layer traces back through the tiers to the row it came from, a recommendation can be checked rather than taken on trust – and a company can act on it.
Built from such parts, the setup runs on plain compute. It can sit with a large cloud provider, with a small regional hoster, or on hardware in the company's own building – a small business data cloud of its own. Because every part follows an open standard, each one can be swapped without the others noticing. Compute changes; the data stays.
It does not have to start large. Such a setup can begin next to the existing systems rather than in place of them: read-only access to two or three sources, a first set of tables, a handful of questions the business wants answered. If it proves itself, it grows. If it does not, nothing has been taken apart. The existing software keeps running the business while the data layer underneath becomes the company's own.
Owning such a setup also means looking after it. Someone has to maintain it, and security is part of that work: access, backups, updates, the question of who may see what. Parts of the stack can be handed to managed services – a managed database, for instance – without giving up the open formats underneath. The difference to a closed platform is that this choice stays reversible. A managed piece can be taken back in-house, or moved to another provider, and the data does not notice.
Being able to choose, and to change the choice later, is the actual gain. A company with its own data foundation can explore it in whatever direction it wants, build its own data products, and shape them around its processes instead of the other way around. It can make the setup as AI-native as it likes. An agent can read and query the data directly, and someone in operations can ask a question in plain language and get the answer together with the row it came from – without waiting for a vendor to put that on the roadmap. Owning the data means owning the processes built on it, and with them the business itself.
Software and compute will keep changing, faster than before. The data a company builds up is what stays.
