Does efficiency beat scale?
On the assumption that bigger is always better, the moment that assumption gave way, and what it means for anyone choosing a model.
Last reviewed: 4 September 2026
Between 2023 and 2025, one rule of thumb held in the AI sector: whoever has the most computing power wins. Larger models, more chips, more data centres. The investment decisions of the entire sector rested on that assumption, as did a considerable part of the market valuations attached to it.
In January 2025, that assumption came under visible pressure for the first time.
What happened
The Chinese firm DeepSeek published its V3 and R1 models. What stood out was not the performance itself, but the ratio between performance and resources.
According to DeepSeek's own technical report, V3 was trained in 2.788 million H800 GPU hours, which at an assumed rate of two dollars an hour amounts to approximately 5.58 million dollars. Estimates for comparable Western models are an order of magnitude higher.
That figure comes with a qualification that drops out of much of the reporting: it covers the final training run. Preceding research and discarded experiments are not included. DeepSeek states this itself. Anyone counting the full development cost arrives at a higher number. Even without that correction, the difference remains considerable.
The market's response was immediate. The announcement removed approximately 589 billion dollars from Nvidia's market capitalisation in a single trading day.
Why the response was so sharp
The valuation of a chip manufacturer in this market rests on the expectation that models will continue to become larger, more expensive and more hardware-hungry. On that assumption, computing power is scarce, and scarcity sets the price.
What DeepSeek showed is that a considerable part of the progress can also come from architecture and training method rather than from more equipment. If the same performance is achievable with fewer chips, then the scarcity is smaller than assumed.
That explains the scale of the response. It was not one competitor being repriced; it was an assumption being revised on which an entire investment logic rested.
What it does not prove
Three corrections to the story as it is often told.
- The curve peaked and fell back.
- DeepSeek today ranks fourth on monthly active users. Its share of global chatbot traffic peaked at approximately 12.1 per cent in February 2025 and settled at around 4.1 per cent in May 2026. A technical breakthrough therefore does not automatically translate into a lasting market position.
- Investment has not stopped.
- On the contrary: the large providers raised their budgets in 2026. Efficiency gains in this market lead to more use, not less. Whoever can compute more cheaply, computes more.
- One case is not a rule.
- That efficiency beat scale here does not mean scale is irrelevant. The largest models still perform best on the hardest general tasks. The correct conclusion is narrower: for a growing number of specific tasks, the largest model is no longer the only workable answer.
What this means for anyone choosing a model
It is precisely that narrower conclusion that is useful to a business.
The question is no longer which model is best in general, but which model is sufficient for your task. That difference has practical consequences.
- A smaller model can run on your own equipment.
- The largest models require infrastructure you do not put in place yourself. A smaller model that performs adequately on your task can be brought in-house, with everything that follows from that for jurisdiction and control.
- A smaller model is more predictable in cost.
You pay for the equipment that stands there, not a rate per use set by a third party. - The measure shifts from benchmark to task.
A model that scores higher on general tests is not necessarily the model that checks your invoices better. That is a measurement problem, and it is solvable: you establish in advance what you measure and against what.
The other side belongs with it. A smaller model performs less well on general and creative tasks. You manage infrastructure yourself. And you carry responsibility for its operation yourself, because there is no supplier standing behind it.
Why we choose this
Thor 1.0 builds on Mistral Small 3, a small open-weight model, rather than on the largest model available.
That is the same trade-off as above. The model has to be able to run on isolated equipment in a Belgian data centre, which is not feasible with the largest models. And it has to be good enough at one defined task: setting a purchase invoice, the corresponding purchase order and the recorded goods receipt alongside one another.
Whether that will work, we do not know. That is why we set the standard in advance, together with a baseline: an error rate below 8 per cent on the test set, with no more than 3 percentage points of regression on general benchmarks. That second threshold is there because tuning to one task generally comes at the expense of others.
That is the substantive point behind this whole article. The question is not whether your model is the largest, but whether you can demonstrate that it is good enough for what you use it for. That requires a measurement point, and most organisations do not have one today.