ExplainedMistral Large 4Mistral AIopen-weight modelsmixture of expertsLLM pricingself-hosting LLMs
Mistral Large 4, explained: a 1 trillion parameter model with open weights promised, and what it takes to run
Mistral Large 4 is a very large AI model that companies will soon be able to download and run on their own servers, if they can afford the machine. Mistral, a French AI company, opened an early preview on 6 October 2026 and promises the files by the end of October. I checked what I could reach today, ran the same code on a small Mistral model, and worked out the hardware and the bill.
Every lesson in both courses has an AI tutor beside it. It reads the same lesson you are reading and answers your questions from it. Try the tutor
The API is open now. The weights are promised for the end of the month.
Where Mistral Large 4 stands on 7 October 2026. The dashed point is a promise in Mistral's launch post, not something you can download yet.
In short
1Mistral Large 4 is the newest and largest model from Mistral AI, a company based in Paris. A preview (an early version that may still change) opened on 6 October 2026 on Mistral's API, the web service where Mistral runs the model for you.
2It has about 1 trillion parameters, the numbers a model learns in training. It is a mixture-of-experts model, so only about 50 billion of them do work on each token (a small piece of text, often part of a word). It reads text and images and can read 1 million tokens in one request.
3The list price is $1.36 per million input tokens and $4.18 per million output tokens, about 2.7 times Mistral Large 3. On 7 October Mistral's docs showed half that price, without saying for how long.
4Mistral says the weights, the files that hold those learned numbers, will be released by the end of October 2026, so companies can run the model on their own servers. At FP8, the likely main format, it needs a server with 8 large NVIDIA GPUs (graphics cards built for AI), because the weights alone are about 1,050 GB.
5I could not send it a prompt today: the API needs a Mistral key, which this project does not have, and Ollama's cloud copy needs a signed-in account. I ran the same code against a small Mistral model on my Mac instead. Every benchmark here is Mistral's own.
Preview released
6 October 2026
From
Mistral AI, Paris
Size
1.05T total, ~52B active
List price
$1.36 in, $4.18 out per million tokens
This page is the free part.
The course goes deeper on running and choosing large models in production
₹999 in India/$49 everywhere elseonce, for the whole course
The AI Engineering course covers running and choosing large models in production across a run of lessons, not one page. These 4 alone are about 231 minutes of step-by-step reading, every one with code you run in the browser, all with a quiz.
The same Qwen2.5 3B model stored three ways and measured on a laptop: 6.18 GB at 16 bits per weight, 3.29 GB at 8.5 and 1.93 GB at about 5. The 8-bit file kept the 16-bit file's top next word on 24 of 24 prompts and got the same 33 of 40 sums right; the 5-bit file kept 19 of 24, all 5 changes on near ties. Writing speed went from 11.1 to 18.3 to 28.4 tokens a second, while reading barely changed.
A model reads one text, not a list of messages. Measured with Ollama and qwen2.5:3b: the message 'Hi' is 1 token as raw text and 30 tokens as a chat, because the template adds markers and a default system message. Filled in by hand, the template gave the same 39 tokens and the same reply as the chat endpoint. On 20 short instructions, its own template gave 20 clean replies; another model family's template gave 0, though 16 first answers were still right.
Why GPU serving gets expensive and how to fix it: the VRAM ceiling, batching, quantization, distillation, caching, card sharing, bin-packing, parallelism, autoscaling, spot, and horizontal vs vertical scaling.
“ML4 is a 1 trillion-parameter natively multimodal model with 49 billion active parameters.”
Mistral, Introducing Mistral Large 4, 6 October 2026
On 6 October 2026 Mistral AI published "Introducing Mistral Large 4". Mistral is a French company that builds AI models and sells access to them, and it often publishes the weights of its models: the numbers a model learned in training, which anyone can then download and run. A model published that way is called an open-weight model.
This release is a preview. You can call the model today through Mistral's API, the web service that runs it for you, from Mistral Studio, its developer console. The weights are not out yet. The launch post says: "We will release the weights by the end of the month." Until then, Mistral says it is red-teaming the model, which means attacking it on purpose to find how it could be misused, with cybersecurity companies, vetted partners and state authorities.
The post calls the model "ML4" for short, and adds a joke name: "very officially: le Chonk". It says the model was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own data centres in Europe, and that the preview runs on that same hardware. Mistral also says a large share of the training data was in more than 160 languages, including every official language of the European Union.
Mistral describes ML4 as a "hybrid instruct-and-reasoning" model. In plain words, the same model can answer straight away, or think step by step before answering when a task is hard, instead of Mistral shipping two separate models for those jobs. It takes images as well as text as input, and it writes text.
2
What is different from Mistral Large 3
Mistral Large 3 came out on 2 December 2025. Its weights are on Hugging Face, the main website for sharing models, under the Apache 2.0 licence, which allows commercial use. Mistral's docs keep a model card for each version, so the two can be compared from Mistral's own pages.
Large 4 is bigger in every row. Total parameters go from 675 billion to 1.05 trillion. The context window, the amount of text the model can read in one request, grows four times, from 256K tokens to 1M. A token is a piece of text, often part of a word; 1 million tokens is roughly 750,000 English words, well over a thousand pages.
It also costs more. Large 3 is $0.50 in and $1.50 out per million tokens. Large 4's list price, in the launch post and on its model card, is $1.36 in and $4.18 out. The model card also lists $0.14 for cached input (text the API has seen recently, such as a long instruction block you send every time, which is cheaper to read again). On 7 October the model card showed those list prices crossed out, with $0.68, $2.09 and $0.07 as the current prices. It does not say why or for how long, so I plan budgets at the list price.
Two numbers on Mistral's pages do not quite match. The launch post says 1 trillion total and 49 billion active parameters. The docs card says 1.05 trillion and 52 billion, plus a 1.6 billion parameter vision encoder (the part that turns an image into something the model can read). Ollama's listing, which I checked myself, also says 1,050,000,000,000. Mistral does not explain the gap. I think the post rounds the numbers, so I give both.
Licence: the Large 4 card says "Open" but names no licence yet. Large 3 is Apache 2.0. Do not assume Large 4 will be the same until Mistral publishes it with the weights.
Large 3 against Large 4, from Mistral's own model cards
Everything grew: size, context and price. Each pair has its own scale, so compare the two bars in a row, not bars in different rows.
3
Why 1 trillion parameters does not mean 1 trillion of work
A parameter is one number inside the model. A normal, or dense, model uses every parameter for every token it reads or writes. Large 4 is a mixture-of-experts model, often shortened to MoE. A model passes each token through many layers, one after another. In Large 4, part of each layer (the feed-forward part) is split into many smaller blocks called experts, and a small piece of the layer, the router, picks a few experts for each token. The rest sit idle for that token. Other parts of the layer, such as attention, run for every token.
So for each token, Large 4 does about the arithmetic of a 52 billion parameter model, which is about 5% of its total. That is why a model this large can still answer at a usable speed.
The catch is memory. Any expert can be picked for the next token, so all of them have to be loaded on the GPUs all the time. The work per token follows the active parameters. The memory follows the total. Real serving speed also depends on how many users share the GPUs, the traffic between GPUs and how long the context is. The memory point decides what it costs to run the weights yourself, which is a later section.
About 5% of the weights work on each token, but 100% must be in memory
The five dark squares change from token to token. Which ones light up here is only an illustration; the 52B of 1.05T is from Mistral's docs card.
4
How good is it? Mistral's numbers, not ours
A benchmark is a fixed set of test tasks with a score, so models can be compared. Every number below is from Mistral's launch post. Some tests were run or designed by outside groups (Artificial Analysis, vals.ai, Surge AI, Lakera), but I found no re-run published by anyone else, and I could not run Large 4 myself.
Benchmark, and what it checks
Large 4's score, as Mistral reports it
DeepSWE v1.1: long, real software engineering tasks
61.7%
SWE-Atlas-QnA: questions about a code repository
59.4%
Terminal-Bench 4: tasks done in a command line
28.3%
Cybench: 40 exercises from security competitions
93% solved
Reproduce and patch a real vulnerability (an Artificial Analysis Cyber Index test)
82%, which Mistral says is the highest of any model
AutomationBench: 657 business workflows in apps like Gmail and Salesforce
59.9%
Dense 200: finding objects in crowded images
42%, against 41% for GPT-6-Astra
B3 (Lakera): resisting prompt injection, text that tries to trick the model into ignoring its instructions
93.3% of attacks resisted
Blind human rating of code quality (Surge AI, 1 to 5)
3.74, second of five models; Claude Opus 5 rated 4.22
Mistral's own figures, from its 6 October 2026 launch post. Not reproduced by us.
Mistral compares ML4 mostly with other open-weight models, such as DeepSeek V4 Pro, Kimi K3, Qwen3.8 Max and GLM-5.3, and says it is the strongest open-weight model built in the US or Europe. It says it is state of the art among open models for cybersecurity, finance and law.
For a sense of where that sits: Google reported 77.9% for Gemini 4 Argon on the same DeepSWE v1.1 test a week earlier. Both numbers come from the companies themselves, run in their own way, so treat the gap as a hint, not a result.
The cybersecurity claim has a twist. Mistral says Claude Opus 5.5 and GPT-6 Astra score near zero on that Artificial Analysis vulnerability test because they refuse the task, while ML4 does it. That is useful to defenders and risky in the wrong hands, and Mistral says it is red-teaming the model before the weights go out.
My hands-on: what I could reach, and the same code on a small Mistral model
First I checked where the model can be reached today. Mistral's API answered every request without a key with HTTP 401, "Invalid API Key". An API key is the secret that proves who is paying, and this project does not have a Mistral key. So the 401 only shows that the API wants a key. Mistral's docs list the model as mistral-large-4.
Then I tried Ollama, a free tool that downloads and runs models on your own computer. Asking for mistral-large-4 failed with "file does not exist" and told me it exists only as a cloud model, mistral-large-4:cloud, which means Ollama runs it on its own servers. Its listing gave 1,050,000,000,000 parameters and a context length of 1,048,576 tokens, which matches Mistral's model card. Running it needs a signed-in Ollama account, which I do not have here, so it stopped at that line.
So no prompt reached Large 4. What I could do was prepare the code. Mistral's API and Ollama both accept the same request format, a chat request sent to a /chat/completions address. I wrote one small script, ask.py, where two settings choose the server and the model. With a Mistral key it calls Large 4. With Ollama it calls a model on my laptop. I ran it against Ministral 3 3B, a small open-weight model from Mistral's previous family, a 3.0 GB download.
Two things surprised me. The short question used 574 input tokens, not about 20. Ollama wraps a message in the model's chat template, the text format a model expects, and for this model the template adds a long default system message (hidden instructions placed before your message) when you do not send one. It begins "You are Ministral-3-3B-Instruct-2512, a Large Language Model (LLM) created by Mistral AI". Locally that costs reading time on every request; with any provider, check what its template adds, because you may be billed for it.
The second is the tool call. I offered the model one tool, get_order_status, and asked about order A-1042. A tool call is when the model replies with a request to run a function, with the inputs filled in, instead of text. This small model did not call the tool. It asked which delivery company I used, and gave the same answer on a repeat run (at temperature 0 a repeat is close to a copy, not a second trial). Large 4 is far bigger and Mistral says it is strong at agent work that uses tools, but this is the kind of thing you test on your own prompts before you trust any model with it.
Its explanation of mixture of experts was also loose: it said a router picks one expert per task. In a real MoE model the router picks a few experts for every token, inside every layer, as described above. A 3B model is a fine stand-in for testing code, not for checking facts.
What we measured: 7 October 2026, from India, on a MacBook with an Apple M4 chip and 24 GB of memory. api.mistral.ai without a key: HTTP 401 in 0.24 to 0.27 s. Ollama 0.32.14 lists mistral-large-4:cloud with 1,050,000,000,000 parameters and a 1,048,576 token context; running it needs sign-in. Local ministral-3:3b: first call 11.27 s while the model loaded, then 2.37 to 3.76 s per call. Tool test: no tool call; it asked a question instead, the same on the repeat run.
$ curl -s https://api.mistral.ai/v1/models
{"detail":"Invalid API Key"}
$ ollama pull mistral-large-4
Error: pull model manifest: file does not exist
"mistral-large-4:cloud" is available as a cloud model. Try:
ollama pull mistral-large-4:cloud
$ ollama show mistral-large-4:cloud
Model
architecture mistral4
parameters 1050000000000
context length 1048576
embedding length 0
quantization
Capabilities
completion
thinking
tools
vision
$ ollama run mistral-large-4:cloud "hello"
You need to be signed in to Ollama to run Cloud models.
# ask.py: one chat request in the OpenAI-style format that Mistral's API and Ollama both accept.
# BASE_URL and MODEL pick the backend; MISTRAL_API_KEY is read from the environment, never printed.
import json, os, sys, time, urllib.request, urllib.error
BASE = os.environ["BASE_URL"]
MODEL = os.environ["MODEL"]
KEY = os.environ.get("MISTRAL_API_KEY", "")
TOOLS = [{"type": "function", "function": {
"name": "get_order_status",
"description": "Look up the delivery status of an order by its id.",
"parameters": {"type": "object", "properties": {"order_id": {"type": "string"}},
"required": ["order_id"]}}}]
def ask(prompt, tools=None):
body = {"model": MODEL, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
if tools:
body["tools"] = tools
headers = {"Content-Type": "application/json"}
if KEY:
headers["Authorization"] = f"Bearer {KEY}"
req = urllib.request.Request(f"{BASE}/chat/completions", data=json.dumps(body).encode(), headers=headers)
t0 = time.perf_counter()
try:
with urllib.request.urlopen(req, timeout=300) as r:
status, data = r.status, json.load(r)
except urllib.error.HTTPError as e:
status, data = e.code, json.loads(e.read() or b"{}")
secs = time.perf_counter() - t0
print(f" {MODEL} at {BASE}: HTTP {status} in {secs:.2f} s")
if status != 200:
print(" error body:", json.dumps(data))
return
msg = data["choices"][0]["message"]
u = data.get("usage", {})
print(f" tokens in {u.get('prompt_tokens')} out {u.get('completion_tokens')}")
if msg.get("tool_calls"):
for c in msg["tool_calls"]:
print(f" tool call: {c['function']['name']}({c['function']['arguments']})")
else:
print(" reply: " + (msg.get("content") or "").strip().replace("\n", "\n "))
print("[plain question]")
ask("In two short sentences, explain what a Mixture-of-Experts language model is to a beginner.")
print("[tool call]")
ask("Where is my order A-1042?", TOOLS)
$ BASE_URL=https://api.mistral.ai/v1 MODEL=mistral-large-4 python3 ask.py # no key set
[plain question]
mistral-large-4 at https://api.mistral.ai/v1: HTTP 401 in 0.27 s
error body: {"detail": "Invalid API Key"}
[tool call]
mistral-large-4 at https://api.mistral.ai/v1: HTTP 401 in 0.24 s
error body: {"detail": "Invalid API Key"}
$ BASE_URL=http://localhost:11434/v1 MODEL=ministral-3:3b python3 ask.py
[plain question]
ministral-3:3b at http://localhost:11434/v1: HTTP 200 in 11.27 s
tokens in 574 out 67
reply: A **Mixture-of-Experts (MoE)** language model divides tasks into specialized "experts" that handle specific types of questions or contexts, while a simpler "router" decides which expert to use. This makes the model more efficient by focusing resources only on relevant parts, improving speed and performance for complex tasks.
[tool call]
ministral-3:3b at http://localhost:11434/v1: HTTP 200 in 2.50 s
tokens in 619 out 36
reply: Could you please provide me with the delivery service or company associated with your order **A-1042**? This will help me check its current status more accurately.
$ BASE_URL=http://localhost:11434/v1 MODEL=ministral-3:3b python3 ask.py # second run, model already loaded
[plain question]
ministral-3:3b at http://localhost:11434/v1: HTTP 200 in 3.76 s
tokens in 574 out 67
[tool call]
ministral-3:3b at http://localhost:11434/v1: HTTP 200 in 2.37 s
tokens in 619 out 36
Real output from my Mac, with these changes for reading: ask.py's source is shown inline above its runs, the comments after # on the $ lines are added, the progress spinner and date lines are removed, the ollama pull lines were recorded a few minutes after the other probes and are placed with them, and in the second local run the two replies, word for word the same as the first, are cut. No API key was set or printed.
6
If you download the weights: the hardware, worked out
Mistral has not published what hardware Large 4 needs, so I did the arithmetic myself with a short script, sizing.py. GPU memory, the fast memory on a graphics card, has to hold all the weights, because of the mixture-of-experts point above. NVIDIA's H100 has 80 GB of it and the newer H200 has 141 GB; big servers hold 8 cards. Each parameter takes 2 bytes at BF16, the usual training format, 1 byte at FP8, and about half a byte at 4-bit. Storing a model in fewer bits like this is called quantization; it makes the files smaller and can change the answers slightly. Mistral shipped Large 3 in FP8, so FP8 is the likely main format; 4-bit is where quality loss is more likely.
The weights are not the only thing in memory. The KV cache stores the work the model has already done on the conversation so far, and it grows with every token of context. A 1 million token context can need a lot of it. The serving software also keeps its own working memory on every card.
Format
Weights
One 8-GPU server: 8 x H100 (640 GB) or 8 x H200 (1,128 GB)
BF16, 2 bytes
2,100 GB
fits neither
FP8, 1 byte
1,050 GB
H100: no. H200: yes, 78 GB left
4-bit (real files larger)
525 GB
H100: on paper, 115 GB left. H200: yes, 603 GB left
My arithmetic from Mistral's 1.05T figure. The serving software and its working memory take part of what is left, so the real space for the KV cache is smaller than shown.
For most teams that means one of three paths: call Mistral's API, wait for cloud providers to host the open weights, or rent a full server. For long contexts at FP8, plan on 8 B200-class cards (the newer, larger NVIDIA GPU) or two 8-GPU H100 or H200 servers working together, not one H200 server with almost no room left. A 4-bit copy is smaller, but check how much quality it loses on your own tests before you serve it.
$ python3 sizing.py
BF16, 2 bytes 2,100 GB -> 27 x H100 80 GB or 15 x H200 141 GB
FP8, 1 byte 1,050 GB -> 14 x H100 80 GB or 8 x H200 141 GB
4-bit, half a byte 525 GB -> 7 x H100 80 GB or 4 x H200 141 GB
Large 3 at FP8 675 GB
8 x H200 = 1,128 GB; at FP8 that leaves 78 GB for the KV cache and everything else
share of weights used per token: 5.0%
one request, 20,000 in / 2,000 out: Large 4 $0.0356 Large 3 $0.0130
the same, 100,000 times a month: Large 4 $3,556 Large 3 $1,300
Real output from sizing.py on my Mac. Weights only, in decimal GB (1 GB = 1,000,000,000 bytes), with the cards at NVIDIA's listed 80 GB and 141 GB. Real 4-bit files keep some parts at higher precision and come out larger, roughly 550 to 650 GB depending on the format. Prices are Mistral's list prices and assume no hidden reasoning tokens. The arithmetic is mine.
About 1 TB of weights at FP8, before any conversation is loaded
For comparison, Mistral's Hugging Face card for Large 3 says its FP8 version runs on one node of B200s or H200s; by my arithmetic that is 675 GB of weights. Large 4 at FP8 just fits an H200 node, with little room left for long conversations.
7
Should you use Mistral Large 4?
A preview is a good moment to plan, not to move production traffic (the live system your users depend on). These are the questions I would answer, in order.
1
Do you need the weights, or only the answers?
If you only need answers, the API works today with a key. Renting an 8-GPU server for a month usually costs far more than the $3,556 API bill in my example below, so self-hosting pays off only at high volume or when the data must stay with you. If your data must stay on your own servers, or you need to change the model, wait for the weights and check the licence that comes with them.
2
Where must your data be processed?
Mistral says Large 4 will be offered in several regions, including a European deployment it runs itself under European law. If you are in Europe and have data-residency rules, that is the main reason to look at it.
3
Is it good at your task?
Take 30 to 50 real examples from your product with known right answers and score Large 4 against the model you use now. This is called an eval. The scores above are Mistral's, on Mistral's chosen tests.
4
What will it cost per request?
My example request, a 20,000-token document with a 2,000-token answer, costs $0.0356 on Large 4 and $0.0130 on Large 3 at list prices. At 100,000 requests a month that is $3,556 against $1,300. If the model thinks step by step first, those hidden tokens are output too and the bill grows. Switch only if your eval shows the better answers are worth the difference.
5
Will tool calls work?
Test every tool your app uses, and check each call's inputs against a schema (a description of the inputs you expect) before running it. In my test the small 3B model never called the tool; a bigger model should do better, but measure it.
6
What is the fallback?
A preview can change or slow down. Keep your current model as a fallback, log which model answered each request, and only switch the default once Large 4 is generally available, meaning the finished release.
My lessons "Token Cost Engineering" and "Model Routing and Cascades: Running Three Models Without Chaos" go deeper on the cost and fallback questions, and "Quantization: The Same Model in Fewer Bits" measures what smaller formats change.
8
What I could and could not check
Everything about Large 4 comes from Mistral's launch post and Mistral's docs model cards, read on 7 October 2026 and linked below. The comparison with Large 3 uses Mistral's Large 3 model card and its Hugging Face card.
I did not send a prompt to Large 4. The API needs a Mistral key, which this project does not have, Ollama's cloud copy needs a signed-in account, and the weights are not out. The parameter count and context length I saw in Ollama's listing are Ollama's description of the model, which matches Mistral's card.
The hardware numbers are arithmetic for the weights alone, not a measured deployment. The local runs used a 3B model on one laptop: they show the code works and how the request looks, not how good Large 4 is. I will update this page when the weights are released.
Free PDF · 12 pages
Get the free AI Engineering Cheat Sheet
What our labs measured about tokens, prompting, RAG, evals, agents, serving, cost and fine-tuning, and the rule each result teaches.
the question most readers have next
Could my company run a model as big as Mistral Large 4 on its own servers?
Once the weights are out, yes, but it is a full server, not a GPU. At 1 byte per parameter the weights are about 1,050 GB, which just fits one 8-GPU H200 machine with very little room left for long conversations. Before you rent that, you would test a smaller 4-bit copy on your own tasks, check that tool calls come back well formed, and compare the bill with the API list price of $1.36 in and $4.18 out per million tokens.
NAVER Cloud: Offers Mistral-based manufacturing AI in Korea, such as quality anomaly detection, so factories can use it without sending their data abroad. Mistral: NAVER Cloud accelerates manufacturing AI
Build it yourself: what you would have to solve
The FP8 weights are about 1 TB and a 4-bit copy is about half that. Measure how often the smaller copy changes the answer before you serve it.
In my test a small 3B Mistral model answered with a question instead of calling the tool. Validate every tool call against its schema and handle the ones that never come.
Mistral AI's newest and largest model, released as a preview on 6 October 2026. It is a mixture-of-experts model with about 1 trillion parameters, of which about 50 billion work on each token. It reads text and images, has a 1 million token context window, and is built for coding, agents and document work.
Is Mistral Large 4 open source?
Its weights are not out yet. On 7 October 2026 it was available only through Mistral's API. Mistral says it will release the weights by the end of October 2026. Open weights means you can download and run the model; it is not the same as open source, which would also include the training code and data. Its model card says "Open" but no licence is named yet; Mistral Large 3 used Apache 2.0.
How much does Mistral Large 4 cost?
The list price is $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14. On 7 October 2026 Mistral's model card showed those crossed out with half prices ($0.68, $2.09 and $0.07) and did not say how long that lasts.
What is the API model name for Mistral Large 4?
Mistral's docs list it as mistral-large-4. You need a Mistral API key; without one, the API answers HTTP 401.
How much GPU memory does Mistral Large 4 need?
Mistral has not published requirements. By my arithmetic, the weights alone are about 2,100 GB at BF16, 1,050 GB at FP8 and 525 GB at 4-bit (real 4-bit files are a little larger). At FP8 the weights just fit one server with eight 141 GB H200 GPUs, with at most 78 GB left over, which is very little for long conversations.
Can I run Mistral Large 4 on my laptop?
No. Even a 4-bit copy is about 525 GB of weights. Ollama lists it only as a cloud model that runs on Ollama's servers. On a laptop you can run Mistral's small open models, such as Ministral 3 3B, which is a 3 GB download.
Is Mistral Large 4 better than GPT-6 Astra or Claude Opus 5?
Mistral reports it ahead of GPT-6 Astra on one image test (Dense 200, 42% against 41%) and on two finance and legal tests run by vals.ai, and behind Claude Opus 5 in a blind human rating of code (3.74 against 4.22). These are Mistral's own reported numbers, and I found no re-run published by anyone else yet.
Sources
What this explainer is based on, so you can check it.
This explainer sits on top of our AI Engineering course: 204 lessons on RAG, evals, agents, serving, security and MLOps, many built around a real experiment. 10 lessons are free to read, with no card needed.
course 2
AI Engineering
Take models from notebook to production, with labs on real models.