<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Mistral AI on Bitsy Wiki</title>
    <link>https://wiki.bitsy.services/wiki/ai/models/mistral/</link>
    <description>Recent content in Mistral AI on Bitsy Wiki</description>
    <generator>Hugo</generator>
    <language>en</language>
    <atom:link href="https://wiki.bitsy.services/wiki/ai/models/mistral/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Mistral Large 3</title>
      <link>https://wiki.bitsy.services/wiki/ai/models/mistral/mistral-large-3/</link>
      <pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate>
      <guid>https://wiki.bitsy.services/wiki/ai/models/mistral/mistral-large-3/</guid>
      <description>&lt;p&gt;Mistral Large 3 is &lt;a href=&#34;https://wiki.bitsy.services/wiki/ai/models/mistral&#34;&gt;Mistral AI&lt;/a&gt;&amp;rsquo;s flagship, released 2 December 2025: 675 billion total parameters, 41 billion active, a 256,000-token context, and weights published under &lt;strong&gt;Apache 2.0&lt;/strong&gt;.&lt;/p&gt;&#xA;&lt;p&gt;That last fact is the reason for this page. Mistral&amp;rsquo;s documentation lists its most capable model under &lt;em&gt;open weight models&lt;/em&gt; rather than in a premier tier — and among the eight builders in this section, no other flagship is Apache 2.0. &lt;a href=&#34;https://wiki.bitsy.services/wiki/ai/models/deepseek&#34;&gt;DeepSeek&lt;/a&gt; publishes its current generation under MIT, which is comparably permissive; every other lab either withholds its best weights entirely or attaches conditions to them.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Mixtral 8x7B</title>
      <link>https://wiki.bitsy.services/wiki/ai/models/mistral/mixtral/</link>
      <pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate>
      <guid>https://wiki.bitsy.services/wiki/ai/models/mistral/mixtral/</guid>
      <description>&lt;p&gt;Mixtral 8x7B, released by &lt;a href=&#34;https://wiki.bitsy.services/wiki/ai/models/mistral&#34;&gt;Mistral AI&lt;/a&gt; on 11 December 2023 under Apache 2.0, was the first sparse &lt;a href=&#34;https://wiki.bitsy.services/wiki/ai/llm/mixture-of-experts&#34;&gt;mixture of experts&lt;/a&gt; that large numbers of people actually ran. Sparse routing had been published years earlier; Mixtral is where it stopped being a research idea and became something a developer downloaded on a Monday.&lt;/p&gt;&#xA;&lt;p&gt;It is on this wiki for a second reason, which is that &lt;strong&gt;its name is wrong in an instructive way&lt;/strong&gt;. &lt;code&gt;8x7B&lt;/code&gt; reads as eight 7-billion-parameter models stapled together — 56 billion. The model has 46.7 billion parameters and uses 12.9 billion per token. Working out where those numbers come from teaches most of what a mixture of experts is.&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
