The new paper is now published and explains how we constructed our 671B R1T Chimera daughter model in less than an hour of CPU time from known base models.
Our new paper “Assembly of Experts: Linear-time construction of the Chimera LLM variants with emergent and adaptable behaviors” has been published on arXiv and Hugging Face.
Serving LLMs for over 50 applications, thereby consuming more than 100 million tokens while generating over ten millions tokens per day, requires us to carefully tune our request...
Serving LLMs for over 50 applications, thereby consuming more than 100 million tokens while generating over ten millions tokens per day, requires us to carefully tune our request...
Our new article shows you how to modernize outdated legacy code with artificial intelligence and migrate it to a new Java version.
Our article “AI-assisted Java Migration” will show you how to modernize outdated legacy code and migrate it to a new Java version using Artificial Intelligence.
Our 24-hour Follow the Sun Hackathon brought together 19 colleagues from Australia, Germany, Hungary, and the UK. For a whole day, we worked continuously to create an AI...
Our 24-hour Follow the Sun Hackathon brought together 19 colleagues from Australia, Germany, Hungary, and the UK. For a whole day, we worked continuously to create an AI...
Our robot “G1PO” has successfully been taught to walk with the help of reinforcement learning.
Over the past few weeks, our Innovation Hacking Team has successfully trained our Unitree G1 Robot 'G1PO' to walk using Reinforcement Learning techniques. We achieved this...
Eight new AMD MI325X GPUs joined our compute cluster of 24 H100. The new Supermicro server is AI beast of a machine having 2 Terabytes GPU memory. It gives us more capacity and...
Eight new AMD MI325X GPUs joined our compute cluster of 24 H100. The new Supermicro server is AI beast of a machine having 2 Terabytes GPU memory. It gives us more capacity and...
DeepSeek-R1T-Chimera, an open-weights model that adds R1's reasoning capabilities to DeepSeek AI V3-0324, has been released.
On the weekend, we released DeepSeek-R1T-Chimera, an open weights model adding R1 reasoning to DeepSeek AI V3-0324. In benchmarks, it appears to be as smart as R1 but much faster...
We recently carried out AI-supported fine-tuning to automate internal workflows.
We recently created a fine-tune of an Optical Character Recognition (OCR) AI model based on olmOCR to help us automate our internal document processing workflows. In our new...
At our recent third AI & Prompt Engineering Meetup, we welcomed 60 guests to our Munich office for an evening full of experiments with Generative AI. In a special edition of...
At our recent third AI & Prompt Engineering Meetup, we welcomed 60 guests to our Munich office for an evening full of experiments with Generative AI. In a special edition of...
Collaborative editing of documents has become an essential requirement for successful remote work. But setting up these collaborative features and maintaining a shared state in...
Collaborative editing of documents has become an essential requirement for successful remote work. But setting up these collaborative features and maintaining a shared state in...
At TNG, we are self-hosting numerous Large Language Models on our cluster of 24 H100 GPUs. It supports 50 different applications, handles over 5,000 inferences per hour, and...
At TNG, we are self-hosting numerous Large Language Models on our cluster of 24 H100 GPUs. It supports 50 different applications, handles over 5,000 inferences per hour, and...