The UK-LLM sovereign AI initiative, established in 2023 as BritLLM and led by University College London, has built a model based on NVIDIA Nemotron that can reason in both English and Welsh, a language spoken by about 850,000 people in Wales today. Enabling high-quality AI reasoning in Welsh supports the delivery of public services including healthcare, education, and legal resources in the language.
"I want every corner of the U.K. to be able to harness the benefits of artificial intelligence. By enabling AI to reason in Welsh, we're making sure that public services — from healthcare to education — are accessible to everyone, in the language they live by," said U.K. Prime Minister Keir Starmer. "This is a powerful example of how the latest AI technology, trained on the U.K.'s most advanced AI supercomputer in Bristol, can serve the public good, protect cultural heritage and unlock opportunity across the country."
The Welsh model was developed in collaboration with Wales' Bangor University and NVIDIA, aligning with the Welsh government's Cymraeg 2050 strategy, which targets a million Welsh speakers by 2050. U.K.-based AI cloud provider Nscale will make the model available to developers through its API. "The aim is to ensure that Welsh remains a living, breathing language that continues to develop with the times," said Gruffudd Prys, senior terminologist and head of the Language Technologies Unit at Canolfan Bedwyr. "AI shows enormous potential to help with second-language acquisition of Welsh as well as for enabling native speakers to improve their language skills."
The new model could also boost accessibility of Welsh resources by enabling public institutions and businesses operating in Wales to translate content or provide bilingual chatbot services — helping healthcare providers, educators, broadcasters, retailers, and restaurant owners keep written content as readily available in Welsh as in English. Beyond Welsh, the UK-LLM team aims to apply the same methodology to other U.K. languages such as Cornish, Irish, Scots, and Scottish Gaelic, and to collaborate internationally on models for languages from Africa and Southeast Asia.
The model is based on NVIDIA Nemotron, an open-source family featuring open weights, datasets, and recipes. The team tapped the 49-billion-parameter Llama Nemotron Super and the 9-billion-parameter Nemotron Nano, post-training them on Welsh-language data. Because far less source data exists in Welsh than in English or Spanish, the team used NVIDIA NIM microservices for gpt-oss-120b and DeepSeek-R1 to translate NVIDIA Nemotron open datasets — over 30 million entries — from English to Welsh, training on hundreds of NVIDIA GH200 Grace Hopper Superchips on Isambard-AI, the U.K.'s most powerful supercomputer, backed by £225 million in government investment and based at University of Bristol.
Bangor University supplied linguistic expertise, with Prys bringing about two decades of language-technology experience to verify machine-translated training data and the model's handling of nuances such as Welsh initial consonant mutation. "This collaboration with NVIDIA and Bangor University enabled us to create new training data and train a new model in record time, accelerating our goal to build the best-ever language model for Welsh," said Pontus Stenetorp, professor of natural language processing and deputy director of the Centre for Artificial Intelligence at UCL. "Our aim is to take the insights gained from the Welsh model and apply them to other minority languages, in the U.K. and across the globe." The model and its Welsh datasets are expected to be made available for enterprise and public sector use.