A year after its release, Apertus is the largest fully open language model available today. It’s trained—transparently—on clean energy, cooled with water from Lake Lugano and built to comply with the EU AI Act. And, it’s already running inside real institutions.
It’s also smaller than the frontier commercial models. It still makes basic mistakes in some of the very languages it was built to serve, and, according to the first independent benchmark to test LLM sustainability claims, not as green in practice as its training story suggests.
But, do the current limitations of Apertus mean that we should throw the baby out with the clean, Swiss lakewater? And what can the story of smaller AI models tell us about the complexity of labelling AI as “good” or “bad”?
An open model built for scrutiny, not for consumers
In a nutshell, Apertus was developed by EPFL, ETH Zurich and the Swiss National Supercomputing Centre and released in September 2025 under the Apache 2.0 licence. Its weights, architecture and training process were all made public, a level of transparency that’s almost unheard of among today’s frontier systems. It comes in 8-billion and 70-billion parameter versions, trained on text in more than 1,800 languages, including minority languages like Romansh and Swiss German that mainstream models typically skip.
Read more: How AI’s Failure on Linguistic Diversity is Deepening Global Inequality
It isn’t a consumer product. There’s no flagship app, and it wasn’t designed to compete head-to-head with mainstream tools like ChatGPT or Gemini. Instead, it’s a foundation model meant to be built into other tools, for example, to perform calculations for the AI functions of commercial tools or to assist researchers with their projects. In particular, its target market is businesses and institutions that need to know exactly what went into the system they’re deploying.
Some current use cases are the translation of sensitive government documents in the Swiss Canton of Ticino, and in a Basel-based online news outlet to prepare local political news for a daily newsletter.
At 70 billion parameters, it’s significantly smaller than GPT-3’s 175 billion (let alone whatever today’s undisclosed frontier models have grown to). Independent testers have flagged awkward phrasing and factual slips in languages native to Switzerland such as Romansh, but also, according to recent online reviews, relatively larger languages such as Dutch.
A work in progress
It’s worth remembering, however, that this kind of rough start isn’t unique to Apertus. ChatGPT’s own 2022 launch was notoriously bumpy. Even now, three years and several versions later, it still hallucinates, frequently generating answers that sound realistic but are infactual. Early clumsiness hasn’t historically been a reliable predictor of how successful a model will be.
Likewise, the consortium behind Apertus has argued that training-data quality and provenance matter more than raw scale, particularly for institutions that can’t risk their AI vendor being caught scraping copyrighted material or leaning on underpaid data labour, as several major commercial labs have been.
Apertus’s training data comes exclusively from public, legally licensed sources, and it’s the first large-scale model (despite coming from non-EU Switzerland) built to meet the EU AI Act’s requirements on transparency and data traceability.
For a European institution choosing a vendor to trust with their most valuable data, that compliance can matter more than a few extra IQ points on a benchmark, particularly if that institution is dealing mostly with English or other dominant languages.
The sustainability picture is more complicated than “clean training run”
At RESET, we’re particularly interested in the sustainability standpoint of new LLMs, particularly if they have any claim to being “green”. Apertus was trained on Alps, CSCS’s supercomputer in Lugano. Its operators describe it as running on carbon-neutral electricity and famously cooling itself with water drawn from Lake Lugano, as opposed to traditionally water- and energy-hungry methods.
This itself is a genuinely unusual data point in an industry where training energy use is normally undisclosed. Those responsible have put a concrete number on it, too. They say the electricity needed to train Apertus is roughly equivalent to what an SBB train would use running continuously for three months. Independent reporting by the NZZ put the actual figure higher, calculating that Alps drew more than 5 gigawatt-hours training Apertus, comparable to the annual electricity use of about 1,500 Swiss households, at a cost of roughly CHF 1.5 million.
That’s a genuinely different posture from the industry norm. Although OpenAI don’t release transparent data, running ChatGPT-4 through a single 100-word email has been estimated to consume roughly the equivalent of a small bottle of water or powering 14 LED light bulbs for an hour. A 2025 study across 22 public broadcasters found that leading AI assistants misrepresented news content in nearly half of their answers. Transparency about the process, even when the numbers themselves are large, is surely worth something.
How to actually use AI as a reader or an institution
None of this is only about Apertus. It’s a useful prompt to think about how anyone, be that a person, a company, a public agency, decides to use AI at all.
The first question is usually the one skipped in the rush to adopt: does this task actually need a large model, or would a simple search or a smaller tool do the job just as well?
Environmental and social impact scale with model size, so reaching for the biggest available system by default is rarely the sustainable choice. Where AI genuinely is the right tool, it’s worth looking past training claims alone to the fuller picture: how the model was built, who labelled its data and under what conditions, where it runs and what happens to the hardware once it’s retired.
Read more: Looking at the Entire Life Cycle: Tips for Sustainable AI Development and Use
A disappointing sustainability score
But a well-documented training run only tells part of the sustainability story. It took a new independent benchmark to show where the gaps are. In mid-2026, Bundesdruckerei, the German federal printing and identity-technology group, published MÖVE (Modelle für die öffentliche Verwaltung evaluieren — “evaluating models for public administration”), the first holistic benchmark built specifically for the public sector.
Rather than the usual leaderboards focused on language test scores, MÖVE scores more than 50 language models across a range of metrics including performance, sustainability and alignment with democratic values, allowing agencies to weight the criteria that actually matter to them.
On MÖVE’s overall ranking, Apertus 70B scores a respectable 69.9. However, its sustainability sub-score comes in at just 53.8, notably weaker than its overall standing. By contrast, Google’s Gemma 4 26B (A4B) posts a sustainability score of 76.2 on the same scale. Although, perhaps we would be remiss to trust Google’s transparent reporting here.
That’s not a contradiction of the clean-energy training claim; it’s a different question entirely. MÖVE’s sustainability score doesn’t factor in how a model was trained at all. It only measures inference: the energy a model uses to generate each answer, converted to CO₂-equivalents assuming standard grid electricity, regardless of what actually powers the model in production. Alps’ hydropower and lake cooling simply never enter that calculation. So a model can be trained as cleanly as Apertus and still score worse here, if it’s less efficient at answering everyday queries afterwards.
MÖVE describes itself as a living benchmark under active development, so this is a snapshot rather than a final verdict. But it’s a useful, humbling one for a project that’s still finding its footing.
Still a work in progress, on every axis
None of this is a reason to write Apertus off. If anything, it’s a fair description of where the project actually is. Joshua Tan, lead maintainer of the nonprofit Public AI Inference Utility, has called Apertus the strongest evidence yet that AI can be run as “public infrastructure, like highways, water pipes, or power grids”, built by public institutions in the public interest rather than by a handful of commercial labs.
That’s an ambitious claim, and it’s also aspirational: infrastructure has to be reliable and well-understood, and Apertus is still working toward both. The rough edges—awkward translations, a sustainability score that undercuts its own marketing—aren’t objectively disqualifying. They’re what you’d expect from a first release built in the open, where the flaws are visible precisely because nothing was hidden.
What’s next
Apertus’s next training round will draw on CHF 20 million ($25 million) in federal funding, again run on the Alps supercomputer. Its developers have called for heavier public investment, arguing that Switzerland currently spends less on this than comparable countries, at the risk of ceding “digital self-determination” to a handful of foreign labs.
Whether that investment pays off may now be easier to judge than it was a year ago. Apertus doesn’t have to beat GPT-4o or Gemini on raw capability to be worth funding, but as tools like MÖVE mature, it will increasingly have to prove its sustainability and governance claims with numbers, not just intentions. That’s a higher bar than “open and renewable-powered,” and it’s one every AI vendor, not just Apertus, is about to start facing.
An AI that Says Less?
While we’re still waiting for a commercial platform to implement Apertus as a foundation, there are different ways to use AI more sustainably and ethically. German start-up Bergwaldprojekt hosts a mixture of LLMs on local servers. RESET recently talked to an expert on sustainable prompting about how to use AI more energy-efficiently.
—


