---
title: MiniMax-H3-TrainingAdapter
canonical_url: "https://www.modelscope.ai/models/DiffSynth-Studio/MiniMax-H3-TrainingAdapter"
md_url: "https://www.modelscope.ai/models/DiffSynth-Studio/MiniMax-H3-TrainingAdapter.md"
repository: DiffSynth-Studio/MiniMax-H3-TrainingAdapter
last_updated: 2026-09-09
pipeline_tag: text-to-video-synthesis
tasks:
  - text-to-video-synthesis
base_model:
  - MiniMax/MiniMax-H3
base_model_relation: adapter
parameters: 310.1M
tensor_type:
  - BF16
library_name:
  - safetensors
downloads: 98
stars: 1
---

# MiniMax-H3-TrainingAdapter

> MiniMax-H3-TrainingAdapter - An open-source model by DiffSynth-Studio on ModelScope. MiniMax-H3 DeCFG Training Adapter

DiffSynth-Studio/MiniMax-H3-TrainingAdapter is a 310.1M-parameter text-to-video-synthesis model on ModelScope. derived from MiniMax/MiniMax-H3.

- **Repository**: DiffSynth-Studio/MiniMax-H3-TrainingAdapter
- **Tasks**: text-to-video-synthesis
- **Parameters**: 310.1M
- **Base model**: MiniMax/MiniMax-H3
- **Downloads**: 98
- **Stars**: 1
- **Last updated**: 2026-09-09

Source: https://www.modelscope.ai/models/DiffSynth-Studio/MiniMax-H3-TrainingAdapter

---

# MiniMax-H3 DeCFG Training Adapter

## Introduction

This repository provides a LoRA training adapter for [MiniMax-H3](https://www.modelscope.cn/models/MiniMax/MiniMax-H3). The base MiniMax-H3 model is CFG-distilled, which can make direct LoRA fine-tuning unstable or degrade the distilled CFG-free behavior. This adapter temporarily pulls the distilled DiT back toward its pre-distillation behavior during training, providing a better optimization landscape for new LoRAs.

The idea is inspired by:

- Differential LoRA training in [DiffSynth-Studio](https://github.com/modelscope/DiffSynth-Studio). See [Differential LoRA Training](https://diffsynth-studio-doc.readthedocs.io/en/latest/Training/Differential_LoRA.html#using-differential-lora-training-in-the-training-framework) for details.
- The training adapter concept from [ai-toolkit](https://github.com/ostris/ai-toolkit).

## Model File

FL2VA:

- `model.safetensors`: adapter weights, for the official MiniMax-H3 DiT.
- `model_for_comfy_dit.safetensors`: format converted adapter weights, for the ComfyUI-layout DiT.

Ref2VA:

- `model_ref2va.safetensors`: adapter weights, for the official MiniMax-H3 Ref2VA DiT.
- `model_ref2va_for_comfy_dit.safetensors`: format converted adapter weights, for the ComfyUI-layout Ref2VA DiT.

The `*_for_comfy_dit` files differ from their counterparts only in the row order of the `qkv_proj` `lora_B` tensors: the official attention reads qkv as `[head][role][dim]` while the ComfyUI layout reads it as `[role][head][dim]`. Every other tensor is byte-identical. Load the variant that matches your DiT layout.

## Usage

### How to Load

1. Download the adapter from ModelScope:

```bash
modelscope download --model DiffSynth-Studio/MiniMax-H3-TrainingAdapter --local_dir ./models/DiffSynth-Studio/MiniMax-H3-TrainingAdapter
```

2. Add the following arguments to your MiniMax-H3 LoRA training command in DiffSynth-Studio:

```bash
--preset_lora_path ./models/DiffSynth-Studio/MiniMax-H3-TrainingAdapter/model.safetensors \
--preset_lora_model dit
```

For Ref2VA training, use the Ref2VA variant instead:

```bash
--preset_lora_path ./models/DiffSynth-Studio/MiniMax-H3-TrainingAdapter/model_ref2va.safetensors \
--preset_lora_model dit
```

This injects the adapter into the DiT during training. After your new LoRA is trained, **discard this adapter and keep only the newly trained LoRA** for inference. A complete training script is shown below.

### Complete Training Script Example

```bash
# 1. Download this adapter
modelscope download --model DiffSynth-Studio/MiniMax-H3-TrainingAdapter --local_dir ./models/DiffSynth-Studio/MiniMax-H3-TrainingAdapter

# 2. Stage 1: data preprocessing
accelerate launch train.py \
  --dataset_base_path "" \
  --dataset_metadata_path <path_to_your_dataset>.jsonl \
  --data_file_keys "video,input_audio" \
  --extra_inputs "input_audio" \
  --max_pixels 399360 \
  --num_frames 124 \
  --dataset_repeat 1 \
  --model_id_with_origin_paths "MiniMax/MiniMax-H3:FL2VA/text_encoder/model*.safetensors,MiniMax/MiniMax-H3:FL2VA/video_vae/source/model.safetensors,MiniMax/MiniMax-H3:FL2VA/audio_vae/model.safetensors" \
  --learning_rate 1e-4 \
  --num_epochs 1 \
  --remove_prefix_in_ckpt "pipe.dit." \
  --output_path "./models/train/my_lora_cache" \
  --lora_base_model "dit" \
  --lora_target_modules "qkv_proj,out_proj,fc1,fc2" \
  --lora_rank 32 \
  --use_gradient_checkpointing \
  --task "sft:data_process"

# 3. Stage 2: train your LoRA (with adapter loaded)
accelerate launch train.py \
  --dataset_base_path ./models/train/my_lora_cache \
  --data_file_keys "video,input_audio" \
  --extra_inputs "input_audio" \
  --max_pixels 399360 \
  --num_frames 124 \
  --dataset_repeat 100 \
  --model_id_with_origin_paths "MiniMax/MiniMax-H3:FL2VA/transformer/model*.safetensors" \
  --learning_rate 1e-4 \
  --num_epochs 1000 \
  --remove_prefix_in_ckpt "pipe.dit." \
  --output_path "./models/train/my_lora" \
  --lora_base_model "dit" \
  --lora_target_modules "qkv_proj,out_proj,fc1,fc2" \
  --lora_rank 32 \
  --use_gradient_checkpointing \
  --find_unused_parameters \
  --task "sft:train" \
  --save_steps 100 \
  --preset_lora_path "./models/DiffSynth-Studio/MiniMax-H3-TrainingAdapter/model.safetensors" \
  --preset_lora_model "dit"
```

### Complete Ref2VA Training Script Example

```bash
# 1. Download this adapter
modelscope download --model DiffSynth-Studio/MiniMax-H3-TrainingAdapter --local_dir ./models/DiffSynth-Studio/MiniMax-H3-TrainingAdapter

# 2. Stage 1: data preprocessing
accelerate launch --num_processes 8 train.py \
  --dataset_base_path "" \
  --dataset_metadata_path <path_to_your_dataset>.jsonl \
  --data_file_keys "video,input_audio,references" \
  --extra_inputs "input_audio,references" \
  --max_pixels 399360 \
  --num_frames 124 \
  --dataset_repeat 1 \
  --model_id_with_origin_paths "MiniMax/MiniMax-H3:Ref2VA/text_encoder/model*.safetensors,MiniMax/MiniMax-H3:Ref2VA/video_vae/source/model.safetensors,MiniMax/MiniMax-H3:Ref2VA/audio_vae/model.safetensors" \
  --processor_path "MiniMax/MiniMax-H3:Ref2VA/processor/" \
  --learning_rate 1e-4 \
  --num_epochs 1 \
  --remove_prefix_in_ckpt "pipe.dit." \
  --output_path "./models/train/my_lora_cache" \
  --lora_base_model "dit" \
  --lora_target_modules "qkv_proj,out_proj,fc1,fc2" \
  --lora_rank 64 \
  --use_gradient_checkpointing_offload \
  --task "sft:data_process"

# 3. Stage 2: train your LoRA (with the Ref2VA adapter loaded)
accelerate launch --num_processes 8 train.py \
  --dataset_base_path ./models/train/my_lora_cache \
  --data_file_keys "video,input_audio,references" \
  --extra_inputs "input_audio,references" \
  --max_pixels 399360 \
  --num_frames 124 \
  --dataset_repeat 8 \
  --model_id_with_origin_paths "MiniMax/MiniMax-H3:Ref2VA/transformer/model*.safetensors" \
  --processor_path "MiniMax/MiniMax-H3:Ref2VA/processor/" \
  --learning_rate 1e-4 \
  --num_epochs 1 \
  --remove_prefix_in_ckpt "pipe.dit." \
  --output_path "./models/train/my_lora" \
  --lora_base_model "dit" \
  --lora_target_modules "qkv_proj,out_proj,fc1,fc2" \
  --lora_rank 64 \
  --use_gradient_checkpointing_offload \
  --find_unused_parameters \
  --task "sft:train" \
  --save_steps 100 \
  --audio_loss_weight 0 \
  --preset_lora_path "./models/DiffSynth-Studio/MiniMax-H3-TrainingAdapter/model_ref2va.safetensors" \
  --preset_lora_model "dit"
```

### Inference

This adapter is for training only and should not be loaded during inference.

## Example LoRA Models

Two toy LoRA models were trained with this adapter as the preset, as references for inference and fine-tuning:

- [MiniMax-H3-Songyu-LoRA](https://www.modelscope.cn/models/mibei0804/MiniMax-H3-Songyu-LoRA): a character identity LoRA on FL2VA, trained with the FL2VA adapter. With a trigger word in the prompt, MiniMax-H3 generates video and audio of that character.
- [MiniMax-H3-Ref2VA-FirstFrame-Lineart](https://www.modelscope.cn/models/mibei0804/MiniMax-H3-Ref2VA-FirstFrame-Lineart): a lineart first-frame control LoRA on Ref2VA, trained with the Ref2VA adapter. Given a lineart first frame and a prompt, it generates the corresponding finished animation with video and audio.

## Training Adapter Details

### FL2VA

- Resolution: 480P
- Frames: 124
- LoRA rank: `64`
- Learning rate: `1e-4`
- Training data: about **3000 self-generated MiniMax-H3 samples**
- Released checkpoint: **step-3000**.

Training command example (two-stage):

```bash
# Stage 1: data preprocessing
accelerate launch train.py \
  --dataset_base_path "" \
  --dataset_metadata_path <path_to_your_dataset>.jsonl \
  --data_file_keys "video,input_audio" \
  --extra_inputs "input_audio" \
  --max_pixels 399360 \
  --num_frames 124 \
  --dataset_repeat 1 \
  --model_id_with_origin_paths "MiniMax/MiniMax-H3:FL2VA/text_encoder/model*.safetensors,MiniMax/MiniMax-H3:FL2VA/video_vae/source/model.safetensors,MiniMax/MiniMax-H3:FL2VA/audio_vae/model.safetensors" \
  --learning_rate 1e-4 \
  --num_epochs 1 \
  --remove_prefix_in_ckpt "pipe.dit." \
  --output_path "./decfg-split-cache" \
  --lora_base_model "dit" \
  --lora_target_modules "qkv_proj,out_proj,fc1,fc2" \
  --lora_rank 64 \
  --use_gradient_checkpointing \
  --task "sft:data_process"

# Stage 2: train the adapter
accelerate launch train.py \
  --dataset_base_path ./decfg-split-cache \
  --data_file_keys "video,input_audio" \
  --extra_inputs "input_audio" \
  --max_pixels 399360 \
  --num_frames 124 \
  --dataset_repeat 100 \
  --model_id_with_origin_paths "MiniMax/MiniMax-H3:FL2VA/transformer/model*.safetensors" \
  --learning_rate 1e-4 \
  --num_epochs 1000 \
  --remove_prefix_in_ckpt "pipe.dit." \
  --output_path "./output" \
  --lora_base_model "dit" \
  --lora_target_modules "qkv_proj,out_proj,fc1,fc2" \
  --lora_rank 64 \
  --use_gradient_checkpointing \
  --find_unused_parameters \
  --task "sft:train" \
  --save_steps 500
```

### Ref2VA

- Resolution: 480P
- Frames: 124
- LoRA rank: `64`, injected into all 50 DiT blocks (`qkv_proj,out_proj,fc1,fc2`)
- Learning rate: `1e-4`
- Training data: about **2500 self-generated MiniMax-H3 Ref2VA samples**
- Released checkpoint: **step-4500**.

Training command example (two-stage):

```bash
# Stage 1: data preprocessing
accelerate launch train.py \
  --dataset_base_path "" \
  --dataset_metadata_path <path_to_your_dataset>.jsonl \
  --data_file_keys "video,input_audio,references" \
  --extra_inputs "input_audio,references" \
  --max_pixels 399360 \
  --num_frames 124 \
  --dataset_repeat 1 \
  --model_id_with_origin_paths "MiniMax/MiniMax-H3:Ref2VA/text_encoder/model*.safetensors,MiniMax/MiniMax-H3:Ref2VA/video_vae/source/model.safetensors,MiniMax/MiniMax-H3:Ref2VA/audio_vae/model.safetensors" \
  --processor_path "MiniMax/MiniMax-H3:Ref2VA/processor/" \
  --learning_rate 1e-4 \
  --num_epochs 1 \
  --remove_prefix_in_ckpt "pipe.dit." \
  --output_path "./ref2va-decfg-split-cache" \
  --lora_base_model "dit" \
  --lora_target_modules "attn.qkv_proj,attn.out_proj,mlp.fc1,mlp.fc2" \
  --lora_rank 64 \
  --use_gradient_checkpointing_offload \
  --task "sft:data_process"

# Stage 2: train the adapter
accelerate launch train.py \
  --dataset_base_path ./ref2va-decfg-split-cache \
  --data_file_keys "video,input_audio,references" \
  --extra_inputs "input_audio,references" \
  --max_pixels 399360 \
  --num_frames 124 \
  --dataset_repeat 100 \
  --model_id_with_origin_paths "MiniMax/MiniMax-H3:Ref2VA/transformer/model*.safetensors" \
  --processor_path "MiniMax/MiniMax-H3:Ref2VA/processor/" \
  --learning_rate 1e-4 \
  --num_epochs 5 \
  --remove_prefix_in_ckpt "pipe.dit." \
  --output_path "./output" \
  --lora_base_model "dit" \
  --lora_target_modules "attn.qkv_proj,attn.out_proj,mlp.fc1,mlp.fc2" \
  --lora_rank 64 \
  --use_gradient_checkpointing_offload \
  --find_unused_parameters \
  --task "sft:train" \
  --save_steps 500
```
