Enterprise AI
Sovereign compute costs a third more. Can you use a third less?
Sovereign compute costs a third more. Can you use a third less?

Dr. Anoj Winston Gladius
·
14
Min. Lesezeit

aufsatz
Image: AI generated with neuland.ai HUB
Four numbers published in 2026 tell a story that none of them tells alone. European data centres run 20 to 30 percent above American operating costs, driven by energy prices, building permits and grid access, with connection lead times of up to ten years in primary markets - and when Cedrik Neike, who runs Digital Industries at Siemens, asked the audience at VivaTech how much more they would pay for European compute, the answer was five to ten percent. Electricity for energy-intensive industry in the EU has averaged more than twice the American price, according to the International Energy Agency, and CBRE expects the cost of data centre capacity in Frankfurt, London, Amsterdam, Dublin and Paris to rise by around 12 percent in 2026. And on 22 September, Anthropic released Claude Opus 5.5 at a price 20 percent below its predecessor. On 28 September the EU Institute for Security Studies put the pieces together in a brief arguing that energy has become a precondition for AI power: the United States, China and the Gulf states combine technological ambition with abundant, affordable energy, while Europe faces high prices, constrained grids and import exposure. Read together, these are not four unrelated stories. The cheapest intelligence in the world is getting cheaper, the sovereign version is getting more expensive, and buyers have said plainly that they will not pay the difference. That leaves European enterprises with one question that matters more than the others: if sovereign compute costs a third more, can you use a third less?
This is the twenty-fifth piece in a series I have been writing for neuland.ai. Each one has pushed the same underlying argument forward: the value, the risk and the moat in enterprise AI are not in the model. They are in the layer above and around it. [¹] Earlier pieces in this series argued that sovereignty of location is about to become a commodity while sovereignty of control is not for sale, and that compute - not model quality - has become the binding constraint. This piece follows both arguments to the place where they meet the budget.
Sovereignty is an efficiency problem before it is a procurement problem. And these four numbers are the clearest demonstration of why that I have seen.
Number one: European compute costs more, and will keep costing more
The cost gap is structural rather than cyclical. European data centres pay more for power, wait longer for grid connections and spend longer in permitting than their American counterparts, and the difference adds up to 20 to 30 percent on operating costs. [²] EU industrial electricity has averaged more than twice the US price, and the energy price shock of 2026, tied to the conflict in the Gulf, has widened the gap rather than closed it. CBRE's European data centre research team expects the cost of securing capacity in the five largest European markets to rise by around 12 percent on average in 2026, and is blunt about where that goes: providers are increasingly passing these costs on to customers. [³]
Demand is not waiting for prices to settle. The International Energy Agency reports that global data centre electricity demand grew 17 percent in 2025, to 485 TWh, and that AI-focused facilities grew by half in a single year. [⁴] Not all of Europe is equal - the Nordic markets and France have materially cheaper power than the European average, which is exactly why Mistral has put a large part of its data centre investment into Sweden. But for most European enterprises, the compute they would like to call sovereign sits in markets where the price is moving up.
Number two: buyers have already said what they will pay
At VivaTech 2026, Siemens' Cedrik Neike asked a room of technology buyers how large a premium they would accept for European compute. Very few hands went up for anything above ten percent. His summary was precise: "5 to 10%, that's the maximum somebody's willing to pay for European compute." [²]
That is the number every sovereignty strategy has to be measured against, and it is a long way short of 20 to 30. It also arrives on top of budgets that are already under pressure. McKinsey's State of AI survey, published in August, found that a fifth of organisations say AI operating costs, including token costs, already constrain their use of AI. [⁵] No procurement committee approves a thirty percent premium on a growing line item because a strategy paper says sovereignty matters. In practice, the cheaper option wins - quietly, workload by workload.
Number three: the frontier keeps getting cheaper
On the other side of the Atlantic the price is moving in the opposite direction. When Anthropic released Claude Opus 5.5, it priced the model 20 percent below Claude Opus 5. [⁶] OpenAI and Google have followed the same pattern across their recent releases. Frontier tokens get cheaper because they are produced at enormous scale, in places where power is cheap, on hardware amortised across the largest customer bases in the world.
Europe is not going to out-invest that. Stanford's AI Index puts private AI investment in 2025 at roughly 286 billion dollars in the United States against about 21 billion in Europe. [⁴] Whatever European compute is built over the next five years will be built at a cost disadvantage, by a market with a fraction of the capital.
What ties the numbers together
It would be tempting to read these as an energy story, a procurement story and a pricing story. They are one story. The cheapest intelligence available to a European enterprise is getting cheaper and it is not sovereign. The sovereign version is getting more expensive. The spread between them is the price of control, and it is widening.
The EUISS brief is right that energy geography, grid reform and location incentives matter, and that Europe should match AI demand to where and when surplus renewable power is available. [⁷] All of that should happen. None of it will close a 20 to 30 percent gap inside the planning horizon of a CIO who has to sign a renewal next quarter. A ten-year grid queue is not a procurement option.
So if the premium will not be paid, and will not be subsidised away in time, there is one variable left that a European enterprise actually controls.
The point that matters more than the price of energy
Here is the part that I have not seen written down clearly enough in the European sovereignty debate. You cannot change the price of a kilowatt-hour in Frankfurt. You can change how many of them you spend per answer, per document reconciled, per engineering check completed.
The arithmetic is worth doing once. Z.ai's GLM-5.3-Flash, released under an MIT licence in August, is priced at roughly a tenth of its larger sibling GLM-5.3, with comparable capability across a wide range of enterprise work. [⁸] Routing a workload from one to the other changes the cost of that workload by around ninety percent. A thirty percent energy premium is small next to a tenfold difference in how much compute the work actually consumes.
And most enterprise AI stacks are a long way from efficient, because they were designed while tokens were effectively subsidised and nobody was measuring. That is where the real room is.
Where the compute actually goes
In our projects we see the same seven sources of waste again and again. None of them needs a research breakthrough to fix.
Every request goes to the biggest model
Most enterprise traffic is extraction, classification, drafting and question answering over the company's own documents. For that work, a domain-tuned open-weight model running on hardware the customer controls frequently matches a frontier model and sometimes beats it, because it has seen the company's vocabulary and documents. [⁹] Frontier capacity should be a routing decision for the work that needs it, not the default for everything.
The whole corpus goes into the prompt
Very large context windows are expensive in a way that grows faster than the context itself, and answer quality drops as the window fills. [¹⁰] The first piece in this series argued that you should always reduce before you reason. Retrieve narrowly, let the system query the corpus during reasoning rather than loading all of it up front, and a task that looked as if it needed a million tokens often needs a few thousand relevant ones.
Every session gets a machine
A common agent design provisions an execution environment for every session, in case the agent needs to run code. Most agent turns never execute anything - they retrieve, read, plan and write. Provisioning on the first genuine need, rather than on arrival, removes a large share of infrastructure that otherwise sits idle.
Idle agents hold on to everything
If every conversation is its own process holding its state in memory, an idle agent costs nearly as much as a busy one - and most agents are idle most of the time. If instead a session keeps its state in a durable log and holds nothing between turns, it can be rebuilt on demand, many sessions can share one process, and idleness becomes almost free.
All work is treated as urgent
This is where the EUISS recommendation becomes practical. Matching AI demand to cheap and renewable power is only possible if the architecture separates work that has to answer in seconds from work that can wait. Bulk ingestion, embedding, re-indexing, evaluation runs and batch analysis can move to cheaper sites and cheaper hours. Interactive reasoning stays close to the user. A platform that treats everything as interactive cannot use cheap power even when it is available.
The computer on every desk does nothing
Apple, Qualcomm, Intel and AMD now ship neural processors in essentially every business laptop, and almost none of that capacity is used for enterprise AI. An earlier piece in this series argued that the endpoint is the compute tier nobody prices. [¹¹] It cannot hold the corpus or the permission model, so retrieval stays on the server - but reasoning steps can run on hardware that is already paid for, already inside the perimeter and already inside European jurisdiction.
Nobody measures cost per result
Cost per token and cost per turn both mislead, because reasoning length varies enormously and a cheap turn that fails costs more than an expensive one that succeeds. The number that shows whether a stack is efficient is cost per completed unit of work: per document reconciled, per finding verified, per case closed. I would encourage every CIO reading this to ask their vendors for that number before the next renewal. If a platform cannot report it, nobody can tell whether it is efficient - including the vendor.
What efficiency cannot do
Three limits belong in the argument. Efficiency does not bring frontier training to Europe; the largest models will keep being trained where energy is cheap, and some tasks will genuinely need them. That is fine when reaching them is a deliberate routing decision with jurisdiction as one of its variables, and not fine when it is the default. Efficiency is also not free: routing, distillation, retrieval discipline and a runtime that provisions only on demand are real engineering work. And it gives Europe no edge over anyone, because an American enterprise can make the same improvements. It removes a handicap. Given the size of the handicap, that is enough.
Where neuland.ai stands on this
The neuland.ai HUB is a sovereign AI management and orchestration platform, and efficiency is one of the reasons it is built the way it is. Every workload is routed against capability, cost, residency and policy: open-weight models we host - GLM, Mistral Large 3, Qwen, DeepSeek and others - carry the bulk of the work, and proprietary endpoints such as Claude and GPT are used where the task justifies them and the customer's policy allows it. Because we do not sell a model, the cheapest adequate answer is a normal outcome rather than a threat to our margin. Retrieval narrows the context before the model reasons, with the user's permissions applied first. The HUB runs on the infrastructure the customer chooses - on-premises, a European sovereign cloud such as STACKIT, or a hyperscaler region where the workload allows. And the agent runtime underneath the platform is being built around two decisions: execution is provisioned only on first need, and sessions hold nothing between turns. Endpoint execution is a stated direction, not yet a shipped capability. [¹²]
This is also the answer to a question I get from CFOs more often than from CIOs: isn't sovereign AI simply more expensive? Per token, honestly, yes - and it will stay that way for some time. Per completed piece of work, it does not have to be. In our experience, the customers who end up paying more for sovereignty are almost always paying for waste they would have had anywhere.
Personal take
I want to be careful with the framing here. This piece is not an argument against European data centres, and it is not an argument against the energy and grid reforms the EUISS brief recommends. Both are necessary, and Europe will need its own compute regardless of how efficient its software becomes.
What I am pushing back against is the idea that sovereignty is something a European enterprise buys - a certificate, a region, a premium on an invoice. Europe's AI debate has so far been conducted almost entirely in capital expenditure: how many gigawatts, how many accelerators, which gigafactory. That is a race Europe loses by construction, and the investment figures say so. The race it can win is output per kilowatt-hour. European industry has competed on efficiency rather than cheap energy for decades, and there is no reason enterprise AI should be the exception.
The regulatory direction points the same way. The EU AI Act has applied broadly since 2 August 2026, with the Digital Omnibus pushing the high-risk Annex III obligations to 2 December 2027 and Annex I to 2 August 2028. [¹³] Under the Union's rating scheme for data centre energy performance, efficiency is also becoming something operators report rather than something they claim. Enterprises that start measuring cost per unit of work now will find both conversations considerably easier.
We at neuland.ai would rather spend our customers' money on fewer, better-chosen tokens than on a premium for tokens they never needed. Sovereign compute costs a third more. Most enterprises can cut far more than a third of what they use today.
Series articles available at neuland.ai/en/resources/insights.
Cedrik Neike (Member of the Managing Board, Siemens AG; CEO Digital Industries), remarks at VivaTech 2026, as reported by DIGITIMES, "Europe's AI infrastructure: the cost gap that policy cannot paper over", June 2026. European data centre operating costs 20 to 30 percent above US levels, driven by energy prices, building permits and grid constraints; lead times of up to ten years in primary markets.
International Energy Agency, Electricity 2026: EU electricity prices for energy-intensive industries averaging more than twice US levels. CBRE European data centre research, 2026: average 12 percent rise in the cost of securing data centre capacity across Frankfurt, London, Amsterdam, Dublin and Paris; quotation attributed to Kevin Restivo, Director, European Data Centre Research, CBRE, as reported by OilPrice.com, May 2026. On the 2026 energy price shock and the relative cost advantage of the Nordic markets and France, see CNBC, "High energy prices could derail Europe's AI race with U.S. and China", 18 May 2026.
International Energy Agency, Key Questions on Energy and AI, 2026: global data centre electricity demand up 17 percent in 2025 to 485 TWh, AI-focused facilities up 50 percent. Stanford HAI, AI Index 2026: private AI investment in 2025 of approximately US$285.9 billion in the United States and approximately US$20.9 billion in Europe.
McKinsey, The State of AI, published 25 August 2026: 20 percent of respondents reporting that AI operating costs, including token costs, constrained their use of AI.
Anthropic, Claude Opus 5.5, released 22 September 2026, priced 20 percent below Claude Opus 5.
European Union Institute for Security Studies, "Feeding the beast: Energy security as a precondition for the EU's AI power", brief, 28 September 2026, edited by Clotilde Bômont and Caspar Hobhouse.
Z.ai, GLM-5.3-Flash, released 26 August 2026 under the MIT licence; list pricing of US$0.15 per million input tokens and US$0.50 per million output tokens, against US$1.40 and US$4.40 for GLM-5.3. See also the earlier piece in this series on blind evaluation.
See the earlier piece in this series on distillation and customer-specific small models.
Hong, Troynikov and Huber, "Context rot: how context degradation affects LLM performance", Chroma technical report, July 2025. On reducing before reasoning, see the first piece in this series, "Control Panels, Execution Surfaces und das Ende der Prompt-First-Automatisierung".
See the earlier piece in this series on the compute constraint and the endpoint as an unpriced tier.
neuland.ai HUB: routing per workload against capability, cost, residency and customer-set policy across self-hosted open-weight and proprietary models; permission-aware retrieval enforced ahead of reasoning; deployment on-premises, in a European sovereign cloud or in a hyperscaler region, including air-gapped operation. The runtime principles described reflect current engineering direction; endpoint-tier inference is a stated architectural direction rather than a shipped capability. neuland.ai AG retains responsibility for content quality and clean delivery of results across all customer engagements.
Regulation (EU) 2024/1689 (AI Act), broadly applicable from 2 August 2026. Council of the EU and European Parliament provisional political agreement on the Digital Omnibus on AI, 7 May 2026: Annex III high-risk obligations postponed to 2 December 2027; Annex I obligations postponed to 2 August 2028. Commission Delegated Regulation (EU) 2024/1364 on the first phase of a common Union rating scheme for data centres under the Energy Efficiency Directive.
Four numbers published in 2026 tell a story that none of them tells alone. European data centres run 20 to 30 percent above American operating costs, driven by energy prices, building permits and grid access, with connection lead times of up to ten years in primary markets - and when Cedrik Neike, who runs Digital Industries at Siemens, asked the audience at VivaTech how much more they would pay for European compute, the answer was five to ten percent. Electricity for energy-intensive industry in the EU has averaged more than twice the American price, according to the International Energy Agency, and CBRE expects the cost of data centre capacity in Frankfurt, London, Amsterdam, Dublin and Paris to rise by around 12 percent in 2026. And on 22 September, Anthropic released Claude Opus 5.5 at a price 20 percent below its predecessor. On 28 September the EU Institute for Security Studies put the pieces together in a brief arguing that energy has become a precondition for AI power: the United States, China and the Gulf states combine technological ambition with abundant, affordable energy, while Europe faces high prices, constrained grids and import exposure. Read together, these are not four unrelated stories. The cheapest intelligence in the world is getting cheaper, the sovereign version is getting more expensive, and buyers have said plainly that they will not pay the difference. That leaves European enterprises with one question that matters more than the others: if sovereign compute costs a third more, can you use a third less?
This is the twenty-fifth piece in a series I have been writing for neuland.ai. Each one has pushed the same underlying argument forward: the value, the risk and the moat in enterprise AI are not in the model. They are in the layer above and around it. [¹] Earlier pieces in this series argued that sovereignty of location is about to become a commodity while sovereignty of control is not for sale, and that compute - not model quality - has become the binding constraint. This piece follows both arguments to the place where they meet the budget.
Sovereignty is an efficiency problem before it is a procurement problem. And these four numbers are the clearest demonstration of why that I have seen.
Number one: European compute costs more, and will keep costing more
The cost gap is structural rather than cyclical. European data centres pay more for power, wait longer for grid connections and spend longer in permitting than their American counterparts, and the difference adds up to 20 to 30 percent on operating costs. [²] EU industrial electricity has averaged more than twice the US price, and the energy price shock of 2026, tied to the conflict in the Gulf, has widened the gap rather than closed it. CBRE's European data centre research team expects the cost of securing capacity in the five largest European markets to rise by around 12 percent on average in 2026, and is blunt about where that goes: providers are increasingly passing these costs on to customers. [³]
Demand is not waiting for prices to settle. The International Energy Agency reports that global data centre electricity demand grew 17 percent in 2025, to 485 TWh, and that AI-focused facilities grew by half in a single year. [⁴] Not all of Europe is equal - the Nordic markets and France have materially cheaper power than the European average, which is exactly why Mistral has put a large part of its data centre investment into Sweden. But for most European enterprises, the compute they would like to call sovereign sits in markets where the price is moving up.
Number two: buyers have already said what they will pay
At VivaTech 2026, Siemens' Cedrik Neike asked a room of technology buyers how large a premium they would accept for European compute. Very few hands went up for anything above ten percent. His summary was precise: "5 to 10%, that's the maximum somebody's willing to pay for European compute." [²]
That is the number every sovereignty strategy has to be measured against, and it is a long way short of 20 to 30. It also arrives on top of budgets that are already under pressure. McKinsey's State of AI survey, published in August, found that a fifth of organisations say AI operating costs, including token costs, already constrain their use of AI. [⁵] No procurement committee approves a thirty percent premium on a growing line item because a strategy paper says sovereignty matters. In practice, the cheaper option wins - quietly, workload by workload.
Number three: the frontier keeps getting cheaper
On the other side of the Atlantic the price is moving in the opposite direction. When Anthropic released Claude Opus 5.5, it priced the model 20 percent below Claude Opus 5. [⁶] OpenAI and Google have followed the same pattern across their recent releases. Frontier tokens get cheaper because they are produced at enormous scale, in places where power is cheap, on hardware amortised across the largest customer bases in the world.
Europe is not going to out-invest that. Stanford's AI Index puts private AI investment in 2025 at roughly 286 billion dollars in the United States against about 21 billion in Europe. [⁴] Whatever European compute is built over the next five years will be built at a cost disadvantage, by a market with a fraction of the capital.
What ties the numbers together
It would be tempting to read these as an energy story, a procurement story and a pricing story. They are one story. The cheapest intelligence available to a European enterprise is getting cheaper and it is not sovereign. The sovereign version is getting more expensive. The spread between them is the price of control, and it is widening.
The EUISS brief is right that energy geography, grid reform and location incentives matter, and that Europe should match AI demand to where and when surplus renewable power is available. [⁷] All of that should happen. None of it will close a 20 to 30 percent gap inside the planning horizon of a CIO who has to sign a renewal next quarter. A ten-year grid queue is not a procurement option.
So if the premium will not be paid, and will not be subsidised away in time, there is one variable left that a European enterprise actually controls.
The point that matters more than the price of energy
Here is the part that I have not seen written down clearly enough in the European sovereignty debate. You cannot change the price of a kilowatt-hour in Frankfurt. You can change how many of them you spend per answer, per document reconciled, per engineering check completed.
The arithmetic is worth doing once. Z.ai's GLM-5.3-Flash, released under an MIT licence in August, is priced at roughly a tenth of its larger sibling GLM-5.3, with comparable capability across a wide range of enterprise work. [⁸] Routing a workload from one to the other changes the cost of that workload by around ninety percent. A thirty percent energy premium is small next to a tenfold difference in how much compute the work actually consumes.
And most enterprise AI stacks are a long way from efficient, because they were designed while tokens were effectively subsidised and nobody was measuring. That is where the real room is.
Where the compute actually goes
In our projects we see the same seven sources of waste again and again. None of them needs a research breakthrough to fix.
Every request goes to the biggest model
Most enterprise traffic is extraction, classification, drafting and question answering over the company's own documents. For that work, a domain-tuned open-weight model running on hardware the customer controls frequently matches a frontier model and sometimes beats it, because it has seen the company's vocabulary and documents. [⁹] Frontier capacity should be a routing decision for the work that needs it, not the default for everything.
The whole corpus goes into the prompt
Very large context windows are expensive in a way that grows faster than the context itself, and answer quality drops as the window fills. [¹⁰] The first piece in this series argued that you should always reduce before you reason. Retrieve narrowly, let the system query the corpus during reasoning rather than loading all of it up front, and a task that looked as if it needed a million tokens often needs a few thousand relevant ones.
Every session gets a machine
A common agent design provisions an execution environment for every session, in case the agent needs to run code. Most agent turns never execute anything - they retrieve, read, plan and write. Provisioning on the first genuine need, rather than on arrival, removes a large share of infrastructure that otherwise sits idle.
Idle agents hold on to everything
If every conversation is its own process holding its state in memory, an idle agent costs nearly as much as a busy one - and most agents are idle most of the time. If instead a session keeps its state in a durable log and holds nothing between turns, it can be rebuilt on demand, many sessions can share one process, and idleness becomes almost free.
All work is treated as urgent
This is where the EUISS recommendation becomes practical. Matching AI demand to cheap and renewable power is only possible if the architecture separates work that has to answer in seconds from work that can wait. Bulk ingestion, embedding, re-indexing, evaluation runs and batch analysis can move to cheaper sites and cheaper hours. Interactive reasoning stays close to the user. A platform that treats everything as interactive cannot use cheap power even when it is available.
The computer on every desk does nothing
Apple, Qualcomm, Intel and AMD now ship neural processors in essentially every business laptop, and almost none of that capacity is used for enterprise AI. An earlier piece in this series argued that the endpoint is the compute tier nobody prices. [¹¹] It cannot hold the corpus or the permission model, so retrieval stays on the server - but reasoning steps can run on hardware that is already paid for, already inside the perimeter and already inside European jurisdiction.
Nobody measures cost per result
Cost per token and cost per turn both mislead, because reasoning length varies enormously and a cheap turn that fails costs more than an expensive one that succeeds. The number that shows whether a stack is efficient is cost per completed unit of work: per document reconciled, per finding verified, per case closed. I would encourage every CIO reading this to ask their vendors for that number before the next renewal. If a platform cannot report it, nobody can tell whether it is efficient - including the vendor.
What efficiency cannot do
Three limits belong in the argument. Efficiency does not bring frontier training to Europe; the largest models will keep being trained where energy is cheap, and some tasks will genuinely need them. That is fine when reaching them is a deliberate routing decision with jurisdiction as one of its variables, and not fine when it is the default. Efficiency is also not free: routing, distillation, retrieval discipline and a runtime that provisions only on demand are real engineering work. And it gives Europe no edge over anyone, because an American enterprise can make the same improvements. It removes a handicap. Given the size of the handicap, that is enough.
Where neuland.ai stands on this
The neuland.ai HUB is a sovereign AI management and orchestration platform, and efficiency is one of the reasons it is built the way it is. Every workload is routed against capability, cost, residency and policy: open-weight models we host - GLM, Mistral Large 3, Qwen, DeepSeek and others - carry the bulk of the work, and proprietary endpoints such as Claude and GPT are used where the task justifies them and the customer's policy allows it. Because we do not sell a model, the cheapest adequate answer is a normal outcome rather than a threat to our margin. Retrieval narrows the context before the model reasons, with the user's permissions applied first. The HUB runs on the infrastructure the customer chooses - on-premises, a European sovereign cloud such as STACKIT, or a hyperscaler region where the workload allows. And the agent runtime underneath the platform is being built around two decisions: execution is provisioned only on first need, and sessions hold nothing between turns. Endpoint execution is a stated direction, not yet a shipped capability. [¹²]
This is also the answer to a question I get from CFOs more often than from CIOs: isn't sovereign AI simply more expensive? Per token, honestly, yes - and it will stay that way for some time. Per completed piece of work, it does not have to be. In our experience, the customers who end up paying more for sovereignty are almost always paying for waste they would have had anywhere.
Personal take
I want to be careful with the framing here. This piece is not an argument against European data centres, and it is not an argument against the energy and grid reforms the EUISS brief recommends. Both are necessary, and Europe will need its own compute regardless of how efficient its software becomes.
What I am pushing back against is the idea that sovereignty is something a European enterprise buys - a certificate, a region, a premium on an invoice. Europe's AI debate has so far been conducted almost entirely in capital expenditure: how many gigawatts, how many accelerators, which gigafactory. That is a race Europe loses by construction, and the investment figures say so. The race it can win is output per kilowatt-hour. European industry has competed on efficiency rather than cheap energy for decades, and there is no reason enterprise AI should be the exception.
The regulatory direction points the same way. The EU AI Act has applied broadly since 2 August 2026, with the Digital Omnibus pushing the high-risk Annex III obligations to 2 December 2027 and Annex I to 2 August 2028. [¹³] Under the Union's rating scheme for data centre energy performance, efficiency is also becoming something operators report rather than something they claim. Enterprises that start measuring cost per unit of work now will find both conversations considerably easier.
We at neuland.ai would rather spend our customers' money on fewer, better-chosen tokens than on a premium for tokens they never needed. Sovereign compute costs a third more. Most enterprises can cut far more than a third of what they use today.
Series articles available at neuland.ai/en/resources/insights.
Cedrik Neike (Member of the Managing Board, Siemens AG; CEO Digital Industries), remarks at VivaTech 2026, as reported by DIGITIMES, "Europe's AI infrastructure: the cost gap that policy cannot paper over", June 2026. European data centre operating costs 20 to 30 percent above US levels, driven by energy prices, building permits and grid constraints; lead times of up to ten years in primary markets.
International Energy Agency, Electricity 2026: EU electricity prices for energy-intensive industries averaging more than twice US levels. CBRE European data centre research, 2026: average 12 percent rise in the cost of securing data centre capacity across Frankfurt, London, Amsterdam, Dublin and Paris; quotation attributed to Kevin Restivo, Director, European Data Centre Research, CBRE, as reported by OilPrice.com, May 2026. On the 2026 energy price shock and the relative cost advantage of the Nordic markets and France, see CNBC, "High energy prices could derail Europe's AI race with U.S. and China", 18 May 2026.
International Energy Agency, Key Questions on Energy and AI, 2026: global data centre electricity demand up 17 percent in 2025 to 485 TWh, AI-focused facilities up 50 percent. Stanford HAI, AI Index 2026: private AI investment in 2025 of approximately US$285.9 billion in the United States and approximately US$20.9 billion in Europe.
McKinsey, The State of AI, published 25 August 2026: 20 percent of respondents reporting that AI operating costs, including token costs, constrained their use of AI.
Anthropic, Claude Opus 5.5, released 22 September 2026, priced 20 percent below Claude Opus 5.
European Union Institute for Security Studies, "Feeding the beast: Energy security as a precondition for the EU's AI power", brief, 28 September 2026, edited by Clotilde Bômont and Caspar Hobhouse.
Z.ai, GLM-5.3-Flash, released 26 August 2026 under the MIT licence; list pricing of US$0.15 per million input tokens and US$0.50 per million output tokens, against US$1.40 and US$4.40 for GLM-5.3. See also the earlier piece in this series on blind evaluation.
See the earlier piece in this series on distillation and customer-specific small models.
Hong, Troynikov and Huber, "Context rot: how context degradation affects LLM performance", Chroma technical report, July 2025. On reducing before reasoning, see the first piece in this series, "Control Panels, Execution Surfaces und das Ende der Prompt-First-Automatisierung".
See the earlier piece in this series on the compute constraint and the endpoint as an unpriced tier.
neuland.ai HUB: routing per workload against capability, cost, residency and customer-set policy across self-hosted open-weight and proprietary models; permission-aware retrieval enforced ahead of reasoning; deployment on-premises, in a European sovereign cloud or in a hyperscaler region, including air-gapped operation. The runtime principles described reflect current engineering direction; endpoint-tier inference is a stated architectural direction rather than a shipped capability. neuland.ai AG retains responsibility for content quality and clean delivery of results across all customer engagements.
Regulation (EU) 2024/1689 (AI Act), broadly applicable from 2 August 2026. Council of the EU and European Parliament provisional political agreement on the Digital Omnibus on AI, 7 May 2026: Annex III high-risk obligations postponed to 2 December 2027; Annex I obligations postponed to 2 August 2028. Commission Delegated Regulation (EU) 2024/1364 on the first phase of a common Union rating scheme for data centres under the Energy Efficiency Directive.