---
title: Gemma4-31B-Uncensored-HauhauCS-1M-GGUF
canonical_url: "https://www.modelscope.ai/models/satgeze/Gemma4-31B-Uncensored-HauhauCS-1M-GGUF"
md_url: "https://www.modelscope.ai/models/satgeze/Gemma4-31B-Uncensored-HauhauCS-1M-GGUF.md"
repository: satgeze/Gemma4-31B-Uncensored-HauhauCS-1M-GGUF
last_updated: 2026-07-06
license: gemma
pipeline_tag: text-generation
tasks:
  - text-generation
base_model:
  - HauhauCS/Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-MTP
base_model_relation: quantized
library_name:
  - gguf
downloads: 356
stars: 0
tags:
  - gguf
  - long-context
  - yarn
  - gemma4
  - uncensored
  - mtp
  - speculative-decoding
  - vision
  - llama.cpp
  - ollama
---

# Gemma4-31B-Uncensored-HauhauCS-1M-GGUF

> Gemma4-31B-Uncensored-HauhauCS-1M-GGUF - An open-source model by satgeze on ModelScope. Gemma4-31B Uncensored: 1M Context + MTP + Vision

satgeze/Gemma4-31B-Uncensored-HauhauCS-1M-GGUF is a text-generation model on ModelScope. licensed under gemma. derived from HauhauCS/Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-MTP.

- **Repository**: satgeze/Gemma4-31B-Uncensored-HauhauCS-1M-GGUF
- **License**: gemma
- **Tasks**: text-generation
- **Base model**: HauhauCS/Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-MTP
- **Tags**: gguf, long-context, yarn, gemma4, uncensored, mtp, speculative-decoding, vision, llama.cpp, ollama
- **Downloads**: 356
- **Stars**: 0
- **Last updated**: 2026-07-06

Source: https://www.modelscope.ai/models/satgeze/Gemma4-31B-Uncensored-HauhauCS-1M-GGUF

---

<img src="banner.jpeg" width="720"/>

# Gemma4-31B Uncensored: 1M Context + MTP + Vision

[HauhauCS/Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-MTP](https://huggingface.co/HauhauCS/Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-MTP) (31B dense, Google QAT checkpoint) with a **1,048,576-token context baked in** (4x the native 262,144), shipping with its MTP speculative-decoding draft head and vision tower. All numbers below were measured on these exact files.

<table>
<tr>
<th style="background:#1a73e8;color:#fff;padding:8px 14px;">Capability</th>
<th style="background:#1a73e8;color:#fff;padding:8px 14px;">Status</th>
</tr>
<tr><td><b>1M context</b></td><td>10/10 through 131K; 262K to 1M rungs in progress, card will update</td></tr>
<tr><td><b>MTP speculative decoding</b></td><td>69.2 to 101.0 tok/s (<b>+46%</b>), acceptance 0.658 (measured on this trunk, RTX 5090)</td></tr>
<tr><td><b>Vision</b></td><td>Verified July 6, 2026: reads image text and identifies objects</td></tr>
<tr><td><b>Uncensored</b></td><td>HauhauCS Balanced abliteration; trunk weights bit-identical to the source release</td></tr>
</table>

## Needle-in-a-haystack

<img src="niah_heatmap.png" width="640"/>

Perfect scores as far as a 32 GB card could take a dense 31B at f16 KV (131K). The 262K through 1M rungs are running on a 128 GB Mac at publish time and this card will be updated as each lands. Side note: DeepReinforce's Ornith-1.0 family description lists an unreleased 31B Dense variant built on this same Gemma 4 trunk; only their Qwen-based models have shipped.

## MTP speculative decoding

<img src="mtp_speedup.png" width="480"/>

The draft head predicts ahead and the trunk verifies every token, so output is identical to standard decoding, only faster. Measured speedup on this uncensored trunk beats the ~35 percent claimed upstream.

## Files

| File | Size | Role |
|---|---|---|
| `gemma4-31b-uncensored-1M-Q4.gguf` | 18.7 GB | Trunk, 1M baked, QAT 4-bit |
| `mtp-gemma-31b.gguf` | 280 MB | MTP draft head, pair with `-md` |
| `mmproj-gemma31b-hauhau.gguf` | 1.2 GB | Vision tower, pair with `--mmproj` |
| `niah_heatmap.png`, `mtp_speedup.png`, `results.jsonl` | small | Verification evidence |

## Every file, every mirror

Nothing was discontinued: every quant is one click away. Hugging Face carries the curated picks, ModelScope always carries everything, and Ollama serves ready-to-run tags.

| File | Size | Hugging Face | ModelScope | Ollama |
|---|---|---|---|---|
| `gemma4-31b-uncensored-1M-Q4.gguf` | 18.7 GB | [download](https://huggingface.co/satgeze/Gemma4-31B-Uncensored-HauhauCS-1M-GGUF/resolve/main/gemma4-31b-uncensored-1M-Q4.gguf) | [download](https://www.modelscope.ai/models/satgeze/Gemma4-31B-Uncensored-HauhauCS-1M-GGUF/resolve/master/gemma4-31b-uncensored-1M-Q4.gguf) | - |
| `mmproj-gemma31b-hauhau.gguf` | 1.2 GB | [download](https://huggingface.co/satgeze/Gemma4-31B-Uncensored-HauhauCS-1M-GGUF/resolve/main/mmproj-gemma31b-hauhau.gguf) | [download](https://www.modelscope.ai/models/satgeze/Gemma4-31B-Uncensored-HauhauCS-1M-GGUF/resolve/master/mmproj-gemma31b-hauhau.gguf) | - |
| `mtp-gemma-31b.gguf` | 280 MB | [download](https://huggingface.co/satgeze/Gemma4-31B-Uncensored-HauhauCS-1M-GGUF/resolve/main/mtp-gemma-31b.gguf) | [download](https://www.modelscope.ai/models/satgeze/Gemma4-31B-Uncensored-HauhauCS-1M-GGUF/resolve/master/mtp-gemma-31b.gguf) | - |

## Run it

llama.cpp, everything on:

```bash
llama-server -m gemma4-31b-uncensored-1M-Q4.gguf \
  -c 1048576 -np 1 --jinja \
  -md mtp-gemma-31b.gguf --spec-type draft-mtp --spec-draft-n-max 3 \
  --mmproj mmproj-gemma31b-hauhau.gguf
```

Ollama (1M and vision work; Ollama has no speculative decoding yet, so the MTP head adds no speed there):

```
FROM ./gemma4-31b-uncensored-1M-Q4.gguf
RENDERER gemma4
PARSER gemma4
PARAMETER num_ctx 262144
```

The RENDERER and PARSER lines avoid imported-GGUF template bugs under tool-heavy use. Raise `num_ctx` as memory allows.

## How this was built

YaRN rope-scaling metadata (factor 4.0 over native 262,144) baked into the GGUF header with gguf-py; weights are bit-identical to the HauhauCS release, no fine-tuning. Gemma 4's dual-rope design takes YaRN on its global-attention layers. Certification harness: 10 needles per rung at depths 5 to 95 percent, temperature 0, seeded prompts, f16 KV only. Method and tooling: [github.com/satindergrewal/aviary-1m](https://github.com/satindergrewal/aviary-1m).

For base capability benchmarks see Google's official Gemma 4 cards; uncensoring quality versus the official trunk has not been independently benchmarked here.

## Credits

Base model and QAT: Google (Gemma license; its terms flow down to these files). Uncensoring and packaging: [HauhauCS](https://huggingface.co/HauhauCS/Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-MTP). MTP head: Unsloth (via the HauhauCS repo). 1M YaRN extension, benchmarking, and certification: [SatGeze](https://huggingface.co/satgeze).

Sister repos: [12B](https://huggingface.co/satgeze/Gemma4-12B-Uncensored-HauhauCS-1M-GGUF) | [26B-A4B](https://huggingface.co/satgeze/Gemma4-26B-A4B-Uncensored-HauhauCS-1M-GGUF) | [31B](https://huggingface.co/satgeze/Gemma4-31B-Uncensored-HauhauCS-1M-GGUF) | [Qwen3.6-35B](https://huggingface.co/satgeze/Qwen3.6-35B-Uncensored-HauhauCS-1M-GGUF)

Mirrors: Hugging Face | [ModelScope](https://www.modelscope.ai/models/satgeze/Gemma4-31B-Uncensored-HauhauCS-1M-GGUF)
