Models

JetBrains debuts Mellum2.1 open model for coding agents

JetBrains has released Mellum2.1, an open-source 12B mixture-of-experts model that uses reinforcement learning to deliver high-speed, self-hosted reasoning for software engineering agents.

MarkTechPost22 hrs agoModels
Image: MarkTechPost

JetBrains has launched Mellum2.1, an open-source 12B mixture-of-experts reasoning model designed for coding agents. Released under the Apache 2.0 license on Hugging Face as Mellum2.1-12B-A2.5B-Thinking, the model retains the architecture of Mellum2 but gains its capabilities from reinforcement learning. It features 28 layers and 64 experts, activating 2.5B parameters per token via a router that selects 8 experts. The model utilizes grouped-query attention with 32 query heads and 4 KV heads, a 1,024-token sliding window on three out of four layers, a 131,072-token context window, and a 98,304-token vocabulary.

The upgrade focuses on post-training, shifting reinforcement learning to the primary training phase. JetBrains trained the model across millions of sandboxed environments using real code repositories. This improved its agentic capabilities: on SWE-bench Verified, Mellum2.1 jumped to 47.0 from Mellum2's score of 2.0, while SWE-bench Pro rose from 0.0 to 28.0. It also scored 17.4 on Terminal-Bench 2.1, up from 0.6. In testing using the Pi v0.73.1 harness, Mellum2.1 led on LiveCodeBench v6 with a score of 82.0, beating Qwen3.5-9B at 75.4 and Gemma 4 E4B at 69.4. It also achieved 91.5 on HumanEval+, 79.4 on MBPP+, and 62.3 on BFCL v4. However, Qwen3.5-9B still leads on SWE-bench Verified at 50.0, SWE-bench Pro at 38.0, GPQA Diamond at 77.8, and AIME 25/26 at 86.7, where Mellum2.1 scored 64.6 and 83.3. Safety improved, with HarmBench scores dropping from 21.5 to 8.5.

For practitioners, Mellum2.1 offers an efficient, self-hostable option for local agentic workflows. Running on an NVIDIA H200 GPU via vLLM or SGLang, the model serves nearly double the tokens of Qwen3.5-9B under heavy load, and multi-token prediction makes single requests 1.6 times faster. Developers can run the model locally using GGUF builds starting at 7.0 GB, with the recommended Q4_K_M build at 8.1 GB offering an 88.0% top-token match. To deploy on vLLM, practitioners should use the reasoning parser for Qwen3 and enable Hermes tool-calling, with recommended settings of 0.6 temperature, 0.95 top_p, and 20 top_k.

This is our own summary of reporting by MarkTechPost

More in Models