---
title: Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored
canonical_url: "https://www.modelscope.ai/models/RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored"
md_url: "https://www.modelscope.ai/models/RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored.md"
repository: RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored
last_updated: 2026-09-15
license: apache-2.0
pipeline_tag: image-text-to-text
tasks:
  - image-text-to-text
model_type:
  - qwen3_5
architectures:
  - Qwen3_5ForConditionalGeneration
base_model:
  - orcarouter/Qwen3.8-27B-Uncensored
base_model_relation: quantized
library_name:
  - gguf
language:
  - en
  - zh
downloads: 38
stars: 0
tags:
  - qwen3
  - qwen3.8
  - uncensored
  - abliterated
  - orcarouter
  - gsq
  - rco
  - gguf
  - quantized
  - speculative-decoding
  - mtp
  - text-generation
  - long-context
  - multimodal
  - vision
  - tool-use
  - 16gb-vram
---

# Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored

> Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored - An open-source model by RentedNoodle on ModelScope. Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3XXS-Uncensored

RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored is a image-text-to-text model on ModelScope. licensed under apache-2.0. derived from orcarouter/Qwen3.8-27B-Uncensored.

- **Repository**: RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored
- **License**: apache-2.0
- **Tasks**: image-text-to-text
- **Base model**: orcarouter/Qwen3.8-27B-Uncensored
- **Tags**: qwen3, qwen3.8, uncensored, abliterated, orcarouter, gsq, rco, gguf, quantized, speculative-decoding, mtp, text-generation, long-context, multimodal, vision, tool-use, 16gb-vram
- **Downloads**: 38
- **Stars**: 0
- **Last updated**: 2026-09-15

Source: https://www.modelscope.ai/models/RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored

---

# Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored

9.75 GiB 混合精度 GSQ/RCO 量化版 OrcaRouter 去限制 Qwen3.8-27B，保留 S1 训练的 MTP 推测解码头与自定义 OrcaRouter-native imatrix，面向单张 16 GB 显卡运行。

这是 **`orcarouter/Qwen3.8-27B-Uncensored`** 的量化版本，不是新的微调。底层模型为 Qwen3.8-27B；外科式去限制来自 OrcaRouter；量化工作遵循 ISTA-DASLab 的 GSQ/RCO 方法。该检查点显著降低学习到的拒绝行为，但不保证普遍服从——部署者需自行负责合规使用。

## 快速规格

| | |
|---|---|
| 基座 | `Qwen/Qwen3.8-27B` 经 `orcarouter/Qwen3.8-27B-Uncensored` |
| 文件 | `Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.0.gguf` |
| 大小 | 10,466,439,424 字节 (9.75 GiB) |
| Trunk 量化 | IQ3_XXS，~3.06 bpw |
| MTP 头 | Q6_K（S1 训练 draft 头）；F32 norms 保留 |
| 对话模板 | `froggeric-qwen3.8-tool-use.jinja`（header `qwen3.8-froggeric-v22.5`） |
| 上下文 | 原生架构 262K；发布配置运行于 32K（needle 在 15K 词 haystack 上验证） |
| 视觉投影器 | 本仓库提供 `mmproj/`（BF16 + Q8_0），可选；纯文本使用无需加载 |

## 实测性能（RTX 5070 Ti 16 GB，build2 llama.cpp，本文件实测）

| 测试集 | 得分 | 配置 |
|---|---|---|
| Needle 检索 | 6/6 | 15K 词 haystack，深度 0.1–0.9，temp 0.0 |
| Toolcall v2（完整 14 例） | 9/14 | 失败：`tc-04` bash-vs-grep、`tc-07` noparse、`tc-08` read-vs-edit、`tcc-01/02` chain-miss |
| GPQA-Diamond | 75.25% | thinking ON，generous budget，temp 0，198 题——**harness-conditional**（见诚实局限第 3 条） |
| WikiText-2 困惑度 (test) | 6.17 | orcaim file |
| IFEval | 70.24% prompt-strict / 76.26% instruction-strict | |
| TruthfulQA | MC1 77.60% / MC2 81.98% | logprob MCQ |
| Refusal | 0% over-refusal | XSTest safe，250 prompts |

## MTP 头（S1 训练）

| 指标 | 结果 |
|---|---|
| 离线 k1，S1 vs 原生头（全新 holdout） | 0.79 vs 0.51（+0.278） |
| 服务 draft 接受率 vs ISTA-imatrix build | +12.4%（同会话 A/B） |
| 服务吞吐 vs ISTA-imatrix build | +13.6%（同会话 A/B） |

离线 k1 增益大且可泛化（在从未训练过的分布上测量）。`n-max 2` 下的服务吞吐受验证开销限制，头部起草优势表现为适度增益，并非离线增益的一对一映射。已在本文件重新验证 needle 6/6。

## 发布配置

Froggeric 模板是目标工具/推理行为所必需的。

```bash
llama-server \
  -m Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.0.gguf \
  --alias Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored \
  --jinja --chat-template-file froggeric-qwen3.8-tool-use.jinja \
  --ctx-size 32768 -b 2048 -ub 2048 -fa on -ngl 99 \
  --reasoning-budget 256 --spec-type draft-mtp --spec-draft-n-max 2
```

**`-b 2048` 是 depth-0.5 检索保真度所必需的。** 这是已记录的 chunking 现象：更小的 batch 会切分检索跨度，导致 depth-0.5 的 needle 丢失。对检索敏感的工作负载请保持 `-b 2048 -ub 2048`。

## 诚实局限（摘要）

- 激进 3-bit 量化——不要期待 BF16 表现。知识回忆弱于更高 bit 版本，模型更依赖检索/工具。
- **RCO 分配表逐字复用** ISTA-DASLab 公开的 stock-base 产物（未重新搜索），应用于去限制基座。
- **imatrix 为自定义 OrcaRouter-native**；在 edit-mask 张量上与 ISTA 不同。ISTA 公开的 imatrix 也已针对本文件验证。
- **Abliteration Tax——GPQA-Diamond 75.25%（harness-conditional）。** thinking ON、generous budget、temp 0、198 题；与 ISTA stock-aligned 88.89% 并非同 harness 对比。约 13 点的差距是去限制的已知代价——残差流的外科式旋转叠加长推理链上的 3-bit 噪声。这是去限制基座的局限，而非 GSQ-RCO 映射的问题（由 WikiText-2 困惑度 6.17 证明其干净）。
- **v1.1 Huihui 线已冻结/弃用**，仅保留用于复现。
- MTP 输出在量化目标上可能与串行生成不同；需要复现时固定 seed 与运行时配置，需要逐 token 一致时关闭推测解码。

## `Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.0-qatfa.gguf`（新文件 — FA 对齐 MTP 头）

本仓库中的第二个 GGUF，与 `v2.0` **除 MTP 草稿头区块外逐字节相同**，针对 FlashAttention KV 路径（`-fa on -ctk q8_0 -ctv q4_0`）校准。

### 文件

| | |
|---|---|
| 文件 | `Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.0-qatfa.gguf` |
| 大小 | 10,442,846,464 字节 (9.73 GiB) |
| SHA-256 | `7e181f30a2f3684e18633a9b670b56818c04d96d82f6e89ae1436d3bee0721a2` |
| bpw | 3.0579 |
| Trunk | IQ3_XXS（与 v2.0 相同） |
| MTP 头 | FA 对齐头，@3300 训练（blk.64） |
| ssm_alpha | BF16 ×48（精度修复） |
| 对话模板 | `froggeric-qwen3.8-tool-use.jinja`（未改动） |

### 相对 v2.0 的变化

1. **FA 对齐 MTP 头 @3300** — 草稿头针对 FA trunk 的 logits 微调，挽回 FA KV 路径下损失的大部分接受率。
2. **ssm_alpha BF16 修复** — ssm_alpha 张量修正为 BF16（×48）。

血缘：同一 v2.0 trunk、同一量化配置、同一模板——唯一差异是头部区块。

### 验证 — 全新 A/B（build2+ship 对照 vs build-fa+qatfa 处理组）

两腿使用相同 flags；测试期 2026-09-11 → 2026-09-13。

| 指标 | build2 + v2.0（对照） | build-fa + v2.0-qatfa | Δ | 门槛 |
|---|---|---|---|---|
| **CUM 接受率（canonical）** | 0.6957 (3093/4446) | 0.6852 (3054/4457) | **−0.0105** | ≤0.020 GREEN |
| CUM 接受率（clean-clean 敏感度） | 0.6957 (3093/4446) | 0.6761 (3085/4563) | −0.0196 | ≤0.020 GREEN（贴近边界） |
| 服务 t/s 中位数 ALL | 62.21 | 73.67 | **+18.4%** | ≥ margin |
| PPL（40×512 wikitext2） | 6.1745 ± 0.14889 | 6.1717 ± 0.14885 | −0.045% | ±3% |
| Needle 1–6 @ b2048 | 5/6 | 5/6 | parity | parity-or-better |

务必知悉（如实陈述）：
- **CUM 逐次运行噪声 ≈ 0.009。** clean-clean Δ（0.0196）仅以 0.0004 落于 0.020 阈值之内；canonical Δ（0.0105）为头条数字。两者均为 GREEN；未移动任何阈值。
- **Needle 为 5/6，而非 6/6。** needle-05（深度 0.5）在所有测量中均未命中——这是该文件血缘的稳定属性，而非两腿差异。v2.0 卡片中的 "Needle 6/6" 是另一次测量（15K 词 haystack、build2 配置）；本次 A/B 测试两腿结果均为 5/6。

### 适用区间（regime guidance）

| 上下文 | 引擎 + 文件 | 说明 |
|---|---|---|
| ≤128K | **build2 + `v2.0.gguf`**（不可变回退） | 快速路径；无需 q8_0/q4_0 FA kernel |
| >128K | **build-fa + `v2.0-qatfa.gguf`**（golden） | FA KV 路径；在该区间取代 build-fa + v2.0 |

`v2.0.gguf`（build2 + ship）保持不变并保留。`build-fa + v2.0.gguf` 在 >128K 区间由 `build-fa + v2.0-qatfa.gguf` 取代。不删除任何二进制或文件。

> 完整英文说明见下方。

---

[⬇️ 完整英文文档 / Full English Documentation Below ⬇️](#english)

<a id="english"></a>

---

# Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored

A 3-bit GGUF quant of the uncensored OrcaRouter Qwen3.8-27B, carrying an S1-trained MTP draft head and a custom OrcaRouter-native imatrix. Built to run on a single 16 GB GPU.

A 9.75 GiB, mixed-precision GSQ/RCO quant of `orcarouter/Qwen3.8-27B-Uncensored`, preserving the MTP head and validated for 16 GB GPUs.

This is a **quantization of `orcarouter/Qwen3.8-27B-Uncensored`**, not a new fine-tune. The underlying model is Qwen3.8-27B; the surgical uncensoring comes from OrcaRouter; the quantization work here follows the GSQ/RCO methodology developed by ISTA-DASLab. The checkpoint is designed to substantially reduce learned refusal behavior. This is not a guarantee of universal compliance — the deployer is responsible for compliant use.

## Quick specs

| | |
|---|---|
| Base | `Qwen/Qwen3.8-27B` via `orcarouter/Qwen3.8-27B-Uncensored` |
| File | `Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.0.gguf` |
| Size | 10,466,439,424 bytes (9.75 GiB) |
| Trunk quant | IQ3_XXS, ~3.06 bpw |
| MTP head | Q6_K (S1-trained draft head); F32 norms preserved |
| Template | `froggeric-qwen3.8-tool-use.jinja` (header `qwen3.8-froggeric-v22.5`) |
| Context | Native architecture 262K; ship profile runs at 32K (needle verified on a 15K-word haystack) |

## What I changed

Starting from the OrcaRouter uncensored checkpoint, I produced this GGUF using a GSQ/RCO-based non-uniform allocation and IQ3_XXS packing, then trained the MTP draft head (S1) and built a custom OrcaRouter-native importance matrix.

The important honesty clause: I **reproduced ISTA-DASLab's published per-tensor RCO allocation verbatim** from their stock-base artifacts. I did not independently re-run the multi-GPU budget search — same map, applied to the uncensored base.

The imatrix is **custom and OrcaRouter-native**; it differs from ISTA's on the edit-mask tensors. ISTA's published imatrix was also validated against this file. If a change makes the benchmark prettier but the model worse to actually use, it does not ship.

## Tested performance (RTX 5070 Ti 16 GB, [build2](https://github.com/RentedNoodle/den_llama.cpp) llama.cpp, this exact file)

| Suite | Score | Setup |
|---|---|---|
| Needle retrieval | 6/6 | 15K-word haystack, depths 0.1–0.9, temp 0.0 |
| Toolcall v2 (full 14-case) | 9/14 | Fails: `tc-04` bash-vs-grep, `tc-07` noparse, `tc-08` read-vs-edit, `tcc-01/02` chain-miss |
| GPQA-Diamond | 75.25% | thinking ON, generous budget, temp 0, 198 questions — **harness-conditional** (see honesty clause 3) |
| WikiText-2 perplexity (test) | 6.17 | orcaim file |
| IFEval | 70.24% prompt-strict / 76.26% instruction-strict | |
| TruthfulQA | MC1 77.60% / MC2 81.98% | logprob MCQ |
| Refusal | 0% over-refusal | XSTest safe, 250 prompts |

## MTP head (S1-trained)

| Metric | Result |
|---|---|
| Offline k1, S1 vs native head (fresh holdout) | 0.79 vs 0.51 (+0.278) |
| Serve draft acceptance vs ISTA-imatrix build | +12.4% (same-session A/B) |
| Serve throughput vs ISTA-imatrix build | +13.6% (same-session A/B) |

The offline k1 gain is large and generalizes (measured on a distribution never trained on). Serve-time throughput at `n-max 2` is bounded by verification cost, so the head's drafting advantage shows as a modest serve-time gain — not a one-to-one map of the offline gain. Needle 6/6 re-verified on this file.

## Usage

Froggeric template is required for the intended tool/reasoning behavior.

```bash
llama-server \
  -m Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.0.gguf \
  --alias Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored \
  --jinja --chat-template-file froggeric-qwen3.8-tool-use.jinja \
  --ctx-size 32768 -b 2048 -ub 2048 -fa on -ngl 99 \
  --reasoning-budget 256 --spec-type draft-mtp --spec-draft-n-max 2
```

**`-b 2048` is required for depth-0.5 retrieval fidelity.** This is a documented chunking artifact: a smaller batch splits the retrieval span and the depth-0.5 needles are lost. Keep `-b 2048 -ub 2048` for retrieval-sensitive workloads.

Vision: projector files live in this repo under `mmproj/` (BF16 or Q8_0). Start the server with one, then send a multimodal chat request (skip both flags for text-only):
```bash
llama-server -m Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.0.gguf --mmproj mmproj/mmproj-Qwen3.8-27B-BF16.gguf [same flags as above]
# then POST /v1/chat/completions with content blocks including {"type":"image_url",...}
```

Ollama (text-only; tool/reasoning behavior is uncertified — use llama-server + Froggeric for the certified path):

```dockerfile
FROM ./Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.0.gguf
PARAMETER num_ctx 32768
PARAMETER num_gpu 99
PARAMETER temperature 0.6
```

- **Dual-engine routing:** for ≤128K context use standard build2 (fast path; measured ~82 vs ~72 t/s at short context). For >128K up to 256K use the BrunoPPassini KV-ring fork — it holds 200K context and passes 2/2 needle retrieval where build2 cannot fit the KV cache.

## Honest limitations

Aggressive 3-bit quant — do not expect BF16 behavior. Knowledge recall is weaker than higher-bit variants; the model leans on retrieval/tools instead. Vision + long context is tight on 16 GB (split text/vision profiles or offload the projector).

1. **RCO allocation is reproduced verbatim** from ISTA-DASLab's stock-base artifacts (not re-searched) and applied to the uncensored base.
2. **The imatrix is custom OrcaRouter-native**; it differs on the edit-mask tensors. ISTA's published imatrix was also validated against this file.
3. **Abliteration Tax — GPQA-Diamond 75.25% (harness-conditional).** Measured with thinking ON, a generous thinking budget, and temp 0 over 198 questions; not a matched comparison to ISTA's stock-aligned 88.89%. The ~13-point gap is the documented cost of uncensoring — surgical rotation of the residual stream combined with 3-bit noise over long reasoning chains. It is a limitation of the abliterated base, not of the GSQ-RCO map, which is proven clean by WikiText-2 perplexity 6.17.
4. **The v1.1 Huihui line is frozen/deprecated** in favor of this OrcaRouter line. v1.1 is retained for reproducibility only.
5. **Quant layout:** IQ3_XXS trunk (~3.06 bpw), MTP head Q6_K, F32 norms preserved.

MTP output may differ from serial generation on quantized targets; fix the seed and runtime config for reproducibility, and disable spec when exact serial behavior matters.

## Files in this repo

| File | What |
|---|---|
| `Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.0.gguf` | The quant (SHA256 `41ad7dfb3f4397d626408a96e88a46c5964e88bd6c4240c191e1131af92ea8cd`) |
| Vision projector | `mmproj/mmproj-Qwen3.8-27B-BF16.gguf` (0.87 GB) + Q8_0 alt (0.60 GB, 27 tensors fallback-quantized) |
| `froggeric-qwen3.8-tool-use.jinja` | Required chat template |
| `config.json`, `tokenizer.json`, `tokenizer_config.json`, `preprocessor_config.json` | Source metadata; `tokenizer_config.json` embeds the Froggeric template in its `chat_template` key |
| `REF-IQ3_XXS-mtp.rco-allocation.txt` | 866-row per-tensor allocation map (authoritative over any summary) |
| `imatrix.dat` | Custom OrcaRouter-native calibration matrix as used |
| `SHA256SUMS.txt` | Hashes |
| `LICENSE` | Upstream Apache-2.0 terms apply; see Qwen / OrcaRouter repos |

GGUF embeds tokenizer metadata — no separate tokenizer needed for llama.cpp/Ollama/LM Studio/Pi. The vision projector is included under `mmproj/` (BF16 default, Q8_0 alternative); it is optional and text-only use works without it.

## Credits

- **Qwen** — architecture + pretrained weights: [Qwen3.8 repo](https://github.com/QwenLM/Qwen3.8), [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B)
- **OrcaRouter** — uncensored checkpoint: [orcarouter/Qwen3.8-27B-Uncensored](https://huggingface.co/orcarouter/Qwen3.8-27B-Uncensored)
- **GSQ** — Dadgarnia, Tabesh, Nikdan, Helcig, Kurtic, Kleinegger, Alistarh (2026): [paper](https://arxiv.org/abs/2604.18556), [code](https://github.com/IST-DASLab/GSQ)
- **RCO** — Helcig & Alistarh (2026): [paper](https://arxiv.org/abs/2605.00649), [code](https://github.com/IST-DASLab/RCO)
- **ISTA-DASLab quant release** — [Qwen3.8-27B-GSQ-RCO-GGUF](https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF) (allocation source; imatrix also validated)
- **Froggeric** — [Qwen-Fixed-Chat-Templates](https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates)
- **llama.cpp/GGML** — [runtime + format](https://github.com/ggml-org/llama.cpp); IQ3_XXS is their standard type, nothing custom here

## Reproducibility

- Base: `orcarouter/Qwen3.8-27B-Uncensored`
- Allocation: `REF-IQ3_XXS-mtp.rco-allocation.txt` (in this repo) — ISTA-DASLab's published per-tensor RCO map, reproduced verbatim, applied via `llama-quantize --tensor-type-file`
- Imatrix: `imatrix.dat` (in this repo) — custom OrcaRouter-native; ISTA's published matrix also validated
- MTP head: S1-trained draft head (Q6_K), F32 norms preserved
- Runtime: [den_llama.cpp](https://github.com/RentedNoodle/den_llama.cpp) (build2)
- GPU: RTX 5070 Ti 16 GB, Windows 11

## `Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.0-qatfa.gguf` (new file — FA-aligned MTP head)

A second GGUF in this repo, byte-identical to `v2.0` **except the MTP draft head block**, calibrated for
the FlashAttention KV path (`-fa on -ctk q8_0 -ctv q4_0`).

### File

| | |
|---|---|
| File | `Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.0-qatfa.gguf` |
| Size | 10,442,846,464 bytes (9.73 GiB) |
| SHA-256 | `7e181f30a2f3684e18633a9b670b56818c04d96d82f6e89ae1436d3bee0721a2` |
| bpw | 3.0579 |
| Trunk | IQ3_XXS (same as v2.0) |
| MTP head | FA-aligned head trained @3300 (blk.64) |
| ssm_alpha | BF16 ×48 (precision fix) |
| Template | `froggeric-qwen3.8-tool-use.jinja` (unchanged) |

### What changed vs v2.0

1. **FA-aligned MTP head @3300** — the draft head is fine-tuned against FA-trunk logits, restoring a
   large share of the acceptance lost under the FA KV path.
2. **ssm_alpha BF16 fix** — ssm_alpha tensors corrected to BF16 (×48).

Lineage: same v2.0 trunk, same quant config, same template — the only delta is the head block.

### Validation — fresh A/B (build2+ship control vs build-fa+qatfa treatment)

Same flags both legs; battery 2026-09-11 → 2026-09-13.

| Metric | build2 + v2.0 (control) | build-fa + v2.0-qatfa | Δ | Gate |
|---|---|---|---|---|
| **CUM accept (canonical)** | 0.6957 (3093/4446) | 0.6852 (3054/4457) | **−0.0105** | ≤0.020 GREEN |
| CUM accept (clean-clean sensitivity) | 0.6957 (3093/4446) | 0.6761 (3085/4563) | −0.0196 | ≤0.020 GREEN (boundary-adjacent) |
| Serve t/s median ALL | 62.21 | 73.67 | **+18.4%** | ≥ margin |
| PPL (40×512 wikitext2) | 6.1745 ± 0.14889 | 6.1717 ± 0.14885 | −0.045% | ±3% |
| Needle 1–6 @ b2048 | 5/6 | 5/6 | parity | parity-or-better |

Caveats (stated plainly):
- **CUM noise ≈ 0.009 run-to-run.** The clean-clean Δ (0.0196) sits only 0.0004 inside the 0.020 bound;
  the canonical Δ (0.0105) is the headline. Both are GREEN; no threshold was moved.
- **Needle is 5/6, not 6/6.** needle-05 (depth 0.5) misses on all measurements — a consistent property of
  the file lineage, not a leg difference. The v2.0 card line "Needle 6/6" is a different measurement
  (15K-word haystack, build2 profile); on this A/B battery the result is 5/6 both legs.

### Regime guidance

| Context | Engine + file | Notes |
|---|---|---|
| ≤128K | **build2 + `v2.0.gguf`** (immutable fallback) | fast path; no q8_0/q4_0 FA kernel needed |
| >128K | **build-fa + `v2.0-qatfa.gguf`** (golden) | FA KV path; supersedes build-fa + v2.0 in this regime |

`v2.0.gguf` (build2 + ship) is retained unchanged. `build-fa + v2.0.gguf` is superseded by
`build-fa + v2.0-qatfa.gguf` for >128K. No binary or file is deleted.

## Versions

### Changelog
* **v2.0 (current):** OrcaRouter uncensored base; S1-trained MTP head; custom OrcaRouter-native imatrix; ISTA-verbatim RCO allocation. Needle 6/6, GPQA-Diamond 75.25% (harness-conditional), WikiText-2 ppl 6.17, toolcall 9/14, IFEval 70.24/76.26, TruthfulQA 77.60/81.98, 0% over-refusal. MTP offline k1 +0.278 over native; serve +13.6% t/s / +12.4% accept vs ISTA-imatrix build.
* **v2.0-qatfa:** same trunk/quant as v2.0; FA-aligned MTP head @3300 + ssm_alpha BF16 fix. Fresh A/B:
  CUM accept Δ −0.0105 canonical / −0.0196 clean-clean (both ≤0.020 GREEN), t/s +18.4%, PPL −0.045%,
  needle 5/6 both legs. Designated >128K golden; v2.0 (build2) retained as ≤128K fallback.
* **v1.x (Huihui line) — frozen/deprecated.** Superseded on every axis by this OrcaRouter line, except the native-head toolcall 10/14 vs 9/14 on this file. Retained for reproducibility only.

### Roadmap
* Bilingual (EN/ZH) ModelScope card: deferred to upload time if requested. This staging card is English-only.

Community quant, not affiliated with Qwen, OrcaRouter, Huihui, or ISTA-DASLab. Bug reports: include model revision, llama.cpp build, ctx/KV/spec settings, sampler. "Feels worse" is fine; reproducible info makes it actionable.
