Fine-tuning
Adapt Muse Glimmer to your domain with supervised fine-tuning (SFT) on your own labeled examples. Low-rank adaptation (LoRA) is the recommended approach: it trains a small set of adapter weights while keeping the base model frozen, reducing GPU memory requirements by an order of magnitude.
Fine-tuning is the first stage of customization. For optimizing against a reward signal after SFT, see reinforcement learning.
Choose an approach
| Method | Relative training cost | When to use |
|---|---|---|
| LoRA | Low | Domain adaptation, style transfer, instruction tuning |
| QLoRA | Low (slower steps, less memory) | Same use cases as LoRA, with a quantized base model |
| Full fine-tune | High | Maximum quality when compute is not a constraint |
Prepare your data
Format your training data as JSONL with the chat format Muse Glimmer expects:
JSON{"messages": [{"role": "system", "content": "You are a medical assistant."}, {"role": "user", "content": "What are the symptoms of type 2 diabetes?"}, {"role": "assistant", "content": "Common symptoms include increased thirst, frequent urination, fatigue, and blurred vision."}]}{"messages": [{"role": "system", "content": "You are a medical assistant."}, {"role": "user", "content": "How is hypertension diagnosed?"}, {"role": "assistant", "content": "Hypertension is diagnosed by measuring blood pressure. A reading consistently at or above 130/80 mmHg indicates hypertension."}]}
- One conversation per line.
- Include a system prompt in every example if you want the model to follow a consistent persona.
- Aim for at least a few hundred (roughly 500+) high-quality examples for meaningful adaptation; several thousand (5,000–10,000+) for strong domain performance. Quality and consistency matter more than raw count.
LoRA fine-tuning
Pythonfrom transformers import AutoModelForCausalLM, AutoTokenizer, TrainingArgumentsfrom peft import LoraConfig, get_peft_modelfrom trl import SFTTrainermodel_name = "meta-models/Muse-Glimmer-30B"model = AutoModelForCausalLM.from_pretrained(model_name,torch_dtype="auto",device_map="auto",)tokenizer = AutoTokenizer.from_pretrained(model_name)lora_config = LoraConfig(r=16,lora_alpha=32,target_modules=["q_proj", "v_proj", "k_proj", "o_proj"],lora_dropout=0.05,task_type="CAUSAL_LM",)model = get_peft_model(model, lora_config)training_args = TrainingArguments(output_dir="./muse-glimmer-lora",num_train_epochs=3,per_device_train_batch_size=4,gradient_accumulation_steps=4,learning_rate=2e-4,bf16=True,logging_steps=10,save_strategy="epoch",)trainer = SFTTrainer(model=model,args=training_args,train_dataset=train_dataset, # your loaded datasetprocessing_class=tokenizer,)trainer.train()trainer.save_model("./muse-glimmer-lora")
QLoRA fine-tuning
QLoRA loads the base model in 4-bit precision with bitsandbytes, then trains LoRA adapters on top. This cuts GPU memory requirements substantially compared to standard LoRA.
Pythonfrom transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig, TrainingArgumentsfrom peft import LoraConfig, get_peft_model, prepare_model_for_kbit_trainingfrom trl import SFTTrainerimport torchmodel_name = "meta-models/Muse-Glimmer-30B"bnb_config = BitsAndBytesConfig(load_in_4bit=True,bnb_4bit_quant_type="nf4",bnb_4bit_compute_dtype=torch.bfloat16,bnb_4bit_use_double_quant=True,)model = AutoModelForCausalLM.from_pretrained(model_name, quantization_config=bnb_config, device_map="auto")model = prepare_model_for_kbit_training(model)tokenizer = AutoTokenizer.from_pretrained(model_name)lora_config = LoraConfig(r=16,lora_alpha=32,target_modules=["q_proj", "v_proj", "k_proj", "o_proj"],lora_dropout=0.05,task_type="CAUSAL_LM",)model = get_peft_model(model, lora_config)trainer = SFTTrainer(model=model,args=TrainingArguments(output_dir="./muse-glimmer-qlora",num_train_epochs=3,per_device_train_batch_size=1,gradient_accumulation_steps=16,learning_rate=2e-4,bf16=True,logging_steps=10,save_strategy="epoch",),train_dataset=train_dataset,processing_class=tokenizer,)trainer.train()trainer.save_model("./muse-glimmer-qlora")
Hyperparameter guidance
| Parameter | Recommended range | Notes |
|---|---|---|
| Learning rate | 1e-4 – 3e-4 (LoRA/QLoRA) | Lower for larger datasets; full fine-tunes use ~1e-5 |
LoRA rank (r) | 8 – 32 | Higher rank = more capacity, more memory |
| LoRA alpha | 2× rank | Standard scaling |
| Batch size | 4 – 16 (effective) | Use gradient accumulation to reach this |
| Epochs | 2 – 5 | Watch for overfitting on small datasets |
| Model context window | 128K by default | The model supports longer contexts; this describes model capacity, not a recommended per-example training length |
| Training sequence length | Dataset- and hardware-dependent | Start with shorter sequences to reduce memory, and increase only after validating quality and hardware headroom |
Evaluate after fine-tuning
Compare your fine-tuned model against the base model on a held-out test set before deploying. Track both your task metric and a general-capability benchmark so you catch regressions from over-fitting:
Load the base model with the adapter you just trained — that is what you have at
this point. Swap ./muse-glimmer-lora for ./muse-glimmer-qlora if you took the
QLoRA path, or point from_pretrained at the merged directory once you have run
Merge adapters below.
Pythonfrom peft import PeftModelfrom transformers import AutoModelForCausalLM, AutoTokenizerbase_model = AutoModelForCausalLM.from_pretrained("meta-models/Muse-Glimmer-30B",torch_dtype="auto",device_map="auto",)model = PeftModel.from_pretrained(base_model, "./muse-glimmer-lora")tokenizer = AutoTokenizer.from_pretrained("meta-models/Muse-Glimmer-30B")correct = 0for example in test_set: # each has .messages (prompt) and .expectedprompt = tokenizer.apply_chat_template(example.messages, tokenize=False, add_generation_prompt=True)inputs = tokenizer(prompt, return_tensors="pt").to(model.device)output = model.generate(**inputs, max_new_tokens=512, do_sample=False)answer = tokenizer.decode(output[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)correct += int(example.expected in answer)print(f"accuracy: {correct / len(test_set):.1%}")
If your fine-tune improves the task metric but drops general capability, reduce epochs, lower the learning rate, or add more diverse data.
Merge adapters
Merge LoRA adapters back into the base model for deployment without the PEFT library:
Pythonfrom peft import PeftModelfrom transformers import AutoModelForCausalLMbase_model = AutoModelForCausalLM.from_pretrained("meta-models/Muse-Glimmer-30B",torch_dtype="auto",device_map="auto",)model = PeftModel.from_pretrained(base_model, "./muse-glimmer-lora")merged = model.merge_and_unload()merged.save_pretrained("./muse-glimmer-30b-merged")
The merged model can be served with any runtime (vLLM, llama.cpp) without adapter overhead.
Next steps
Optimize your fine-tuned model further with reinforcement learning, then deploy it with vLLM, llama.cpp, or ExecuTorch. To convert it for local inference, see quantization.