> For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://unsloth.ai/docs/zh/mo-xing/mistral-3.5.md).

# Mistral 3.5：如何在本地运行

用于 Mistral 3.5 模型的指南，可在你的设备上本地运行或微调

Mistral 发布了 Mistral-Medium-3.5-128B，他们新的稠密型 128B 参数、多模态、混合推理模型。它支持文本和图像输入、文本输出、256K 上下文窗口，并擅长推理、编码、长上下文、工具使用、智能体工作流以及多模态文档/图像理解。

Mistral Medium 3.5 在体积仅为其 5 倍大小模型的情况下也提供极具竞争力的性能。本地运行约需 64GB 内存。GGUF： [Mistral-Medium-3.5-128B-GGUF](https://huggingface.co/unsloth/Mistral-Medium-3.5-128B-GGUF)

{% hint style="success" %}
**2026年5月1日更新：** 我们与 Mistral 合作修复了影响某些实现方式的 Mistral Medium 3.5 推理问题，并发布了包含修复的更新版 GGUF（**与 Unsloth 无关** 或我们的量化版本）。该问题由 YaRN 解析怪癖导致，影响了多个实现，包括 `transformers` 和 `llama.cpp`。修改 `mscale_all_dim` 从 `1` 到 `0` 即可解决。我们也修复了 `mmproj` 文件未能正确生成的问题。

<mark style="background-color:$success;">**Mistral 现已将我们的修复提交到他们的官方仓库！**</mark>
{% endhint %}

### 使用指南

{% hint style="info" %}
GGUF 的视觉功能目前已受支持。其余支持会在之后提供。
{% endhint %}

表：Mistral Medium 3.5 推荐硬件需求。单位为总内存：RAM + VRAM，或统一内存。

| Mistral 3.5     | 3位    | 4位    | 8位         |
| --------------- | ----- | ----- | ---------- |
| Medium 3.5 128B | 64 GB | 80 GB | 128-170 GB |

{% hint style="info" %}
你的可用总内存应至少超过你下载的量化模型大小。如果不满足，llama.cpp 仍可在部分 RAM / 磁盘卸载的情况下运行，但生成速度会更慢。对于长上下文、更大的批次、工具繁重的智能体运行以及图像提示词，你还需要更多内存。
{% endhint %}

#### 推荐设置

使用 Mistral 推荐的推理设置：

* `reasoning_effort="none"` → 快速即时回复、聊天、信息提取和简单指令。
* `reasoning_effort="high"` → 推理模式，推荐用于复杂提示、编码、研究、数学和智能体使用。

推荐的采样默认值：

* 使用 `temperature = 0.7` 用于 `reasoning_effort="high"`.
* 使用 `temperature = 0.0` 到 `0.7` 用于 `reasoning_effort="none"`，具体取决于任务。
* 将重复和存在惩罚保持禁用，或设置为 `1.0` ，除非你看到循环输出。
* 最大上下文长度为 `262,144`

#### **推理模式**

Mistral Medium 3.5 支持即时指令模式和带有 'high' 选项的推理模式。

要在 llama.cpp / llama-server 中启用高推理：

```bash
--chat-template-kwargs '{"reasoning_effort":"high"}'
```

要禁用推理：

```bash
--chat-template-kwargs '{"reasoning_effort":"none"}'
```

如果你使用的是 Windows PowerShell，请使用：

```powershell
--chat-template-kwargs "{\"reasoning_effort\":\"none\"}"
```

## 运行 Mistral 3.5 教程

由于 Mistral Medium 3.5 是一个稠密型 128B 模型，本地推理推荐从动态 4 位 GGUF 开始。GGUF： `unsloth/Mistral-Medium-3.5-128B-GGUF`

<a href="/docs/zh/mo-xing/mistral-3.5.md#unsloth-studio-guide" class="button primary">在 Unsloth Studio 中运行</a><a href="/pages/5902f65155c9213c17d6735294471de2eb587dd1#llama.cpp-guide" class="button secondary">在 llama.cpp 中运行</a>

{% hint style="warning" %}
目前没有任何多模态/视觉 GGUF 可以在 **Ollama** 中工作，因为它们使用单独的 `mmproj` 视觉文件。请使用兼容 llama.cpp 的后端。

不要使用 **CUDA 13.2** ，否则你可能会得到乱码输出。NVIDIA 正在修复。
{% endhint %}

### 🦥 Unsloth Studio 指南

在本教程中，我们将使用 [Unsloth Studio](/docs/zh/xin/studio.md)，这是我们用于运行和训练 LLM 的新 Web UI。使用 Unsloth Studio，你可以在本地运行模型并输入 **音频**、图像和文本，适用于 **Mac、Windows**和 Linux，并且：

{% columns %}
{% column %}

* 搜索、下载、 [运行 GGUF](/docs/zh/xin/studio.md#run-models-locally) 以及 safetensor 模型
* **对比** 模型 **并排**
* [**自我修复的** 工具调用](/docs/zh/xin/studio.md#execute-code--heal-tool-calling) + **网页搜索**
* [**代码执行**](/docs/zh/xin/studio.md#run-models-locally) （Python、Bash）
* [自动推理](https://unsloth.ai/docs/desktop#feature-deep-dive) 参数调优（temp、top-p 等）
* [训练 LLM](/docs/zh/xin/studio.md#no-code-training) 快 2 倍，显存占用减少 70%
  {% endcolumn %}

{% column %}

<div data-with-frame="true"><figure><img src="https://2657992854-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FFeQ0UUlnjXkDdqhcWglh%2Fskinny%20studio%20chat.png?alt=media&amp;token=c2ee045f-c243-4024-a8e4-bb4dbe7bae79" alt=""><figcaption></figcaption></figure></div>
{% endcolumn %}
{% endcolumns %}

{% stepper %}
{% step %}

#### 安装 Unsloth

**MacOS、Linux、WSL：**

```bash
curl -fsSL https://unsloth.ai/install.sh | sh
```

**Windows PowerShell：**

```bash
irm https://unsloth.ai/install.ps1 | iex
```

{% endstep %}

{% step %}

#### 设置 Unsloth Studio（一次性）

设置会自动安装 Node.js（通过 nvm）、构建前端、安装所有 Python 依赖，并构建启用 CUDA 支持的 llama.cpp。

{% hint style="info" %}
**WSL 用户：** 系统会提示你输入 `sudo` 密码，以安装构建依赖项（`cmake`, `git`, `libcurl4-openssl-dev`).
{% endhint %}
{% endstep %}

{% step %}

#### 启动 Unsloth

**MacOS、Linux、WSL：**

```bash
source unsloth_studio/bin/activate
unsloth studio -H 0.0.0.0 -p 8888
```

**Windows Powershell：**

```bash
& .\unsloth_studio\Scripts\unsloth.exe studio -H 0.0.0.0 -p 8888
```

<div data-with-frame="true"><figure><img src="https://2657992854-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2Fd1yMMNa65Ccz50Ke0E7r%2FScreenshot%202026-03-17%20at%2012.32.38%E2%80%AFAM.png?alt=media&amp;token=9369cfe7-35b1-4955-b8cb-42f7ecb43780" alt="" width="375"><figcaption></figcaption></figure></div>

**然后打开 `http://localhost:8888` 在你的浏览器中。**
{% endstep %}

{% step %}

#### 搜索并下载 Mistral Medium 3.5

首次启动时，你需要创建一个密码来保护你的账户，并在之后重新登录。然后前往 [Unsloth Chat](/docs/zh/xin/studio/chat.md) 选项卡，在搜索栏中搜索 Mistral 3.5，并下载你想要的模型和量化版本。
{% endstep %}

{% step %}

#### 运行 Mistral 3.5

使用 Unsloth Studio 时，推理参数应自动设置，不过你仍然可以手动修改。你也可以编辑上下文长度、聊天模板和其他设置。

更多信息请查看我们的 [Unsloth Studio 推理指南](/docs/zh/xin/studio/chat.md).
{% endstep %}
{% endstepper %}

### 🦙 Llama.cpp 指南

在本指南中，我们将使用 Unsloth Dynamic 4-bit 版本运行 Mistral Medium 3.5。见： `unsloth/Mistral-Medium-3.5-128B-GGUF`.

在这些教程中，我们将使用 llama.cpp 进行快速本地推理，尤其适合你拥有 CPU 或大内存统一内存机器时。

**1. 构建 llama.cpp**

获取最新的 `llama.cpp` 在 GitHub 上。修改 `-DGGML_CUDA=ON` 到 `-DGGML_CUDA=OFF` 如果你没有 GPU，或者只想进行 CPU 推理。对于 Apple Mac / Metal 设备，请设置 `-DGGML_CUDA=OFF`；Metal 支持默认开启。

```bash
apt-get update
apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y
git clone https://github.com/ggml-org/llama.cpp
cmake llama.cpp -B llama.cpp/build \\
    -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON
cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split
cp llama.cpp/build/bin/llama-* llama.cpp
```

**2. 直接从 Hugging Face 运行**

```bash
export LLAMA_CACHE="unsloth/Mistral-Medium-3.5-128B-GGUF"

./llama.cpp/llama-cli \\
    -hf unsloth/Mistral-Medium-3.5-128B-GGUF:UD-Q4_K_XL \\
    --temp 0.7 \\
    --chat-template-kwargs '{"reasoning_effort":"none"}'
```

用于高推理模式：

```bash
./llama.cpp/llama-cli \\
    -hf unsloth/Mistral-Medium-3.5-128B-GGUF:UD-Q4_K_XL \\
    --temp 0.7 \\
    --chat-template-kwargs '{"reasoning_effort":"high"}'
```

**3. 手动下载模型**

在安装 `huggingface_hub` 和 `hf_transfer`:

```bash
pip install huggingface_hub hf_transfer

hf download unsloth/Mistral-Medium-3.5-128B-GGUF \\
    --local-dir unsloth/Mistral-Medium-3.5-128B-GGUF \\
    --include "*UD-Q4_K_XL*" \\
    --include "*mmproj*"
```

如果下载卡住，请设置：

```bash
export HF_HUB_ENABLE_HF_TRANSFER=1
```

**4. 运行本地 GGUF**

```bash
./llama.cpp/llama-cli \\
    --model unsloth/Mistral-Medium-3.5-128B-GGUF/Mistral-Medium-3.5-128B-UD-Q4_K_XL.gguf \\
    --temp 0.7 \\
    --chat-template-kwargs '{"reasoning_effort":"none"}'
```

如果包含多模态投影器 GGUF，请使用：

```bash
./llama.cpp/llama-cli \\
    --model unsloth/Mistral-Medium-3.5-128B-GGUF/Mistral-Medium-3.5-128B-UD-Q4_K_XL.gguf \\
    --mmproj unsloth/Mistral-Medium-3.5-128B-GGUF/mmproj-BF16.gguf \\
    --temp 0.7 \\
    --chat-template-kwargs '{"reasoning_effort":"none"}'
```

#### llama-server 部署

要在 llama-server 上部署 Mistral Medium 3.5，请使用：

```bash
./llama.cpp/llama-server \\
    -hf unsloth/Mistral-Medium-3.5-128B-GGUF:UD-Q4_K_XL \\
    --alias "mistral-medium-3.5" \\
    --host 0.0.0.0 \\
    --port 8001 \\
    --temp 0.7 \\
    --chat-template-kwargs '{"reasoning_effort":"none"}'
```

用于推理模式：

```bash
--chat-template-kwargs '{"reasoning_effort":"high"}'
```

如果你使用的是 Windows PowerShell，请使用：

```powershell
--chat-template-kwargs "{\"reasoning_effort\":\"high\"}"
```

你可以使用兼容 OpenAI 的请求来 ping llama-server：

```bash
curl http://localhost:8001/v1/chat/completions \\
  -H "Content-Type: application/json" \\
  -d '{
    "model": "mistral-medium-3.5",
    "messages": [
      {"role": "user", "content": "解释即时模式和推理模式之间的主要区别。"}
    ],
    "temperature": 0.7
  }'
```

### Mistral 3.5 最佳实践

#### 提示示例

**简单推理提示**

```
系统：
你是一个精确的推理助手。请谨慎求解，只给出最终答案和简短解释。

用户：
一列火车在上午 8:15 出发，并在上午 11:47 到达。旅程持续了多久？
```

使用 `reasoning_effort="high"` 适用于这种提示风格。

**OCR / 文档提示**

对于 OCR 和文档提取，请将图像放在前面并要求结构化输出。

```
[图像在前]
从这张收据中提取所有文本。以 JSON 形式返回商家、日期、项目列表和总计。
```

**多模态对比提示**

```
[图像 1]
[图像 2]
比较这两张截图，并告诉我哪一张更可能让新用户感到困惑。给出 3 个具体原因。
```

**编码智能体提示**

```
你是一个在代码仓库中工作的编码智能体。
先检查相关文件，然后提出一个最小补丁。
返回最终答案，包含：摘要、修改的文件、运行的测试和风险。
```

使用 `reasoning_effort="high"` 以及用于代码库探索的工具调用。

**JSON / 函数调用提示**

```
在需要计算或查询时，使用提供的工具。
只返回有效 JSON。不要在 JSON 对象之外包含说明文字。
```

### 基准测试

<div><figure><img src="https://2657992854-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2Ffzh7zfLQYF1Tn6YAriUx%2Frfffrrf.png?alt=media&amp;token=0247ce56-5b61-4ffb-a3dc-44ac6e4477d0" alt=""><figcaption></figcaption></figure> <figure><img src="https://2657992854-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2Fj7W8gGzonNSM3aEY2qQr%2Ffrfrfref.png?alt=media&amp;token=cebcb72b-6813-47ed-ba31-a80e6d4b43d8" alt=""><figcaption></figcaption></figure></div>


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://unsloth.ai/docs/zh/mo-xing/mistral-3.5.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
